An interpretable human-machine hybrid enhanced perception method based on a new generation of artificial intelligence theory
By constructing an interpretable human-machine hybrid augmented perception method based on next-generation artificial intelligence theory, integrating multimodal data streams and information streams, establishing a human-machine hybrid augmented perception model and an online human-machine fusion perception knowledge base, the uninterpretability and consistency issues of the perception layer in the human-machine co-driving system are solved, and the robustness and practicality of the system are improved.
Patent Information
- Application Number
- CN202311120447.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-01
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2043-09-01
AI Technical Summary
Existing human-machine co-driving systems suffer from human-machine conflict and uninterpretable perception issues in their perception layer design, resulting in low system robustness and difficulty in achieving high feasibility and practicality.
By constructing an interpretable human-machine hybrid augmented perception method based on next-generation artificial intelligence theory, integrating multimodal data streams and information streams, establishing a human-machine hybrid augmented perception model and an online human-machine fusion perception knowledge base, and employing extended Kalman filtering and finite state machine models, consistency and interpretability of human-machine perception are achieved.
It improves the perceptual consistency and interpretability of the human-machine co-driving system, enhances the system's robustness and practicality, possesses self-correction capabilities, and ensures a high degree of consistency between the driver and the autonomous driving system.
Smart Images

Figure CN117151218B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to an interpretable human-machine hybrid enhanced perception method, in particular to an interpretable human-machine hybrid enhanced perception method based on a new generation of artificial intelligence theory. BACKGROUND
[0002] Intelligence is one of the important parts of the "new four modernizations" of automobiles, and has attracted widespread attention due to its ability to greatly improve vehicle safety and traffic efficiency, and has been researched and applied in various aspects of vehicles. The introduction of relevant policies also marks that the intelligent technology of automobiles has become a national development strategy. According to the grading standards of relevant authoritative agencies, the current main technical breakthroughs are concentrated on the conditional automatic driving and high-level automatic driving of L3 and L4 levels. In the current development of automobile intelligence, the high-level automobile intelligent technology generally has strict limitations and low robustness, and the human-machine co-driving technology as a highly feasible and practical solution can overcome these shortcomings. The vehicle operation process of an intelligent vehicle can be roughly summarized as three processes: perception, decision-making and planning, and control. The data processing method of the perception layer as the input of the whole human-machine co-driving system directly affects the overall effect of the human-machine co-driving system.
[0003] The human-machine co-driving system integrates the functions of the driver and the automatic driving system, and fuses or switches the functions of the two through the introduction of arbitration. The existing human-machine co-driving system does not have special processing in the perception layer and the perception layer of the traditional automatic driving system, and mainly focuses on sensor accuracy, multi-sensor fusion and vehicle state estimation. The proposal of this scheme is initially to meet the needs of fully automatic driving. One of the essential differences between the fully automatic driving system and the human-machine co-driving system is that the fully automatic driving system does not need the driver's function. This difference directly leads to the inevitable phenomenon of human-machine conflict when the human-machine co-driving system adopts a similar perception system as the fully automatic driving system. In order to overcome this defect and achieve high consistency between humans and machines in the human-machine co-driving system, it is necessary to specially design the perception layer of the human-machine co-driving system, which can further improve the practicality of the human-machine co-driving system.
[0004] In addition, the perception system of modern intelligent vehicles often integrates various machine learning artificial intelligence algorithms. Traditional artificial intelligence algorithms are often based on statistics and probability theory, and have strong non-interpretable nature, which makes the intelligent vehicle a black box model. Developing an interpretable human-machine fusion perception system is human-centered - how can we explain the perception model to achieve human trust in the human-machine co-driving system and even the intelligent technology of automobiles, so that its application is safer and more reliable, and the human-machine collaborative co-driving becomes an effective and feasible method of automobile intelligence.
[0005] Chinese patent 202211129873.3 discloses a human-computer consistent human-computer co-driving environment perception method, introduces a safety officer attention direction detection mechanism, enhances perception accuracy, obtains environment perception results, and improves short-term and short-distance road condition perception accuracy to ensure driving safety. Chinese patent CN202111451350.6 discloses a human-computer co-driving intention fusion control method, which better fuses machine driving intention and human driving intention, concentrates the advantages of both, obtains the most reasonable driving intention, and maximizes the needs of both parties. Chinese patent CN202310348694.7 discloses a target perception method, computer equipment, storage medium and vehicle, which can improve the 3D perception efficiency of the target under the premise of ensuring the perception accuracy. The above three patents are representative of the existing automatic driving system or human-computer co-driving system perception layer patent, which focuses on perception accuracy and perception efficiency. In order to make the human-computer co-driving system more and more practical, we also need to solve the problem of explainability and human-computer perception consistency on the basis of these problems. SUMMARY
[0006] The purpose of the present application is to solve the problem of human-computer perception consistency in human-computer co-driving system, overcome the defect of black box unexplainable perception layer, and provide an explainable human-computer hybrid enhanced perception method based on new generation artificial intelligence theory by realizing human-computer hybrid enhanced perception model and online human-computer fusion perception knowledge base.
[0007] The explainable human-computer hybrid enhanced perception method based on new generation artificial intelligence theory provided by the present application includes the following steps:
[0008] The first step is to integrate input data stream and information stream, and the specific steps are as follows:
[0009] Step one, multi-modal human-traffic mixed situation data flow and information flow, multi-modal human-traffic mixed situation data flow and information flow are integrated through the sensor data collected by the vehicle-mounted sensor, the multi-modal human-traffic mixed situation data flow includes two parts of the driver data flow of the host vehicle and the driving environment data flow, the driver data flow of the host vehicle refers to the data flow collected by each sensor of the driver monitoring system in the intelligent driving cabin, specifically including the driver operation data flow composed of the steering wheel angle, the throttle and brake pedal opening and gear data flow, and the driver biological signal data flow composed of the driver biological electrical signal and the driver face image data flow, the driving environment data flow refers to the sensor data flow for sensing the surrounding traffic scene on the intelligent vehicle, including the inertial navigation data flow composed of IMU, GPS positioning and map data flow, the visual information data flow composed of front-view camera and surround-view camera image data flow, the radar point cloud data flow composed of laser radar point cloud, millimeter wave radar point cloud and ultrasonic wave radar data flow, and the other sensor information data flow composed of infrared sensor data and road surface monitoring sensor data flow;
[0010] The multi-modal human-traffic mixed situation data flow is preprocessed to obtain the multi-modal human-traffic mixed situation information flow, for the driver data flow of the host vehicle, the driver eye movement and head movement information are extracted and the blink frequency is calculated according to the image information of the camera in the DMS, forming the driver information flow of the host vehicle; for the driving environment data flow, the traffic participants are extracted and tracked according to the visual information, and the traffic participant information is extracted according to the radar point cloud information, forming the driving environment information flow;
[0011] Step two, internal parameter data flow and information flow of human-machine mixed augmented perception system, the human-machine mixed augmented perception system framework includes the human-machine mixed augmented perception model constructed in the second step and the interpretable online human-machine fusion perception knowledge base constructed in the third step, this step integrates the data flow and information flow between the internal modules of the human-machine mixed augmented perception system, these data flow and information flow include the online driver perception model data flow and the machine perception model data flow, and the offline perception model knowledge information flow and the perception atlas knowledge information flow;
[0012] Step three, human-machine mixed augmented perception interpretability evaluation information flow, the system interpretability is defined as the performance of the system having a clear physical meaning, as shown in the following formula:
[0013] Argmax E Q(E|M,H,D,T)(1)
[0014] In the formula, Q represents an explanatory evaluation equation; E represents a specific method of realizing explainability; M represents a model parameter to be explained; H represents a driver model parameter; D represents a result data of a driving test; and T represents a driving task performed, and formula (1) represents that the essence of explainability is to seek an explanation method: the method makes a specific driver have maximum understanding for a specific model given specific data and for a specific task;
[0015] Second step, build a human-machine hybrid enhanced perception model, the specific steps are as follows:
[0016] Step one, machine perception model based on multi-sensor fusion, this step first receives the driving environment data stream and information stream output by step one in the first step, and reconstructs the traffic participants and the road in the surrounding environment in three-dimensional space according to the size of the host vehicle and the installation position of the sensor, restores the driving environment, as follows:
[0017] This step first analyzes the traffic participants by establishing a state equation based on the driving environment data stream and information stream output by step one in the first step, then predicts the state of the traffic participants at the current time based on the traffic participant information at the last time, and updates the state of the traffic participants at the current time according to the sensor data at the current time to continuously update the state of the detection target, wherein the prediction and update formulas of the state of the traffic participants at the current time are as follows:
[0018]
[0019]
[0020] In the formula, x t represents the state vector at time t predicted by the extended Kalman filter; u t represents the control vector at time t predicted by the extended Kalman filter; F t represents the state transition matrix; B t represents the control input at time t; w t represents the state vector noise term at time t with zero mean; Q t represents the covariance matrix of the noise term at time t; x t represents the state variable at time t output by the extended Kalman filter; y t represents the state vector input by the sensor at time t; P t represents the uncertainty degree of the system state at time t predicted by the extended Kalman filter, represented by the covariance of the state vector; P t represents the uncertainty degree of the system state at time t output by the extended Kalman filter; H tHt represents a mapping matrix from the state space to the measurement space at time t; R represents an uncertainty matrix of the measurement value, which is provided by the manufacturer of the sensor;
[0021] For the traffic participants who are not perceived at the current time due to the limited perception range of the sensor, the state vector predicted by the extended Kalman filter is taken as the output value for the traffic participant information output;
[0022] Step two, driver perception model based on situational awareness theory;
[0023] Step three, human-machine perception consistency comparison model, in this step, the perception results of the driver and the human-machine co-driving system are compared to realize the consistency of human-machine perception;
[0024] Step four, human-machine fusion perception output model, the output of the human-machine hybrid enhanced perception model is formed through the calculation of this step;
[0025] Third step, build an interpretable online human-machine fusion perception knowledge base, the human-machine fusion perception knowledge base contains two parts: perception atlas knowledge and perception model knowledge. The perception atlas knowledge is the original data knowledge recorded after screening the human-machine fusion perception atlas output in the second step, and the perception model knowledge is the knowledge recording the system decision process of the human-machine hybrid enhanced perception model used in the second step. In this step, first, the system framework of the human-machine fusion perception knowledge base is built, and then the human-machine fusion perception knowledge base is built and updated online according to the human-machine perception consistency comparison results output in step three of the second step and the human-machine fusion perception model output model output in step four of the second step. The specific steps are as follows:
[0026] Step one, system framework of human-machine fusion perception knowledge base, according to the interpretability requirements and use requirements of the human-machine fusion perception knowledge base, a finite state machine FSM is used to build the system framework of the human-machine fusion perception knowledge base. FSM analyzes and processes various types of information and makes behavioral decisions in combination with experience. The human-machine fusion perception knowledge base is composed of several FSMs according to different scenarios. Each state in FSM is composed of two parts: semantic knowledge and human-machine hybrid enhanced perception model parameters. The semantic knowledge is obtained by clustering method and is used for interpretability analysis of the model. The semantic knowledge includes three parts: driver, human-machine co-driving system and driving scene. The perception state of FSM is classified by fuzzy rule method. The model parameters are the human-machine hybrid enhanced perception model parameters set by human according to the state of FSM, which are output to each step in the second step for decision operation;
[0027] In order to ensure the integrity of the human-machine fusion perception knowledge base, the perception states applicable to FSM are defined as a finite universal set U, and the i-th perception state in FSM is defined as A iIf the number of perception states is n, then the perception states must satisfy the following formula:
[0028]
[0029] Formula (7) indicates that there is no overlap between each different sensing state in the FSM, and all sensing states cover the entire sensing state space;
[0030] Step 2: Perception Model Knowledge. In this step, prior knowledge is formed based on driving experiments and then learned through imitation. This step makes the constructed online human-machine fusion perception knowledge base interpretable. The specific process is as follows:
[0031] First, through multiple driving tests, the decision trajectory sequence D = {τ1, τ2, ..., τ} of human drivers was collected. n}, where each decision trajectory τ i The sequence of human-machine fusion perception state-human-machine fusion perception knowledge base update action pairs represents a human driver's driving process. The human-machine fusion perception state is used as a feature and the human-machine fusion perception knowledge base update action is used as a label. The initial state human-machine fusion perception knowledge base is generated by using D as the training dataset through supervised learning. In order to make the generated human-machine fusion perception knowledge base universal, training dataset D is collected by selecting multiple drivers of average level to perform driving tasks.
[0032] Step 3: Perceive knowledge of the graph;
[0033] Step 4: Knowledge base update mechanism. In order to achieve consistency between human and machine perception, the human-machine fusion perception knowledge base is kept updated after the initial construction in Step 2 to improve the consistency between human and machine perception. When the human-machine fusion perception map output in Step 4 of Step 2 is received, the human-machine fusion perception knowledge base is updated.
[0034] The fourth step is to integrate the output variables, consolidate the system's output variables, and form the internal outputs for human-computer interaction and human-machine co-driving systems, as well as for interpretability testing. The specific steps are as follows:
[0035] Step 1: Integration of Human-Machine Hybrid Enhanced Perception Process Quantities. This step integrates the inputs and outputs of all modules included in Step 1, Step 2, and Step 3 in a time-aligned manner, and outputs the integrated human-machine hybrid enhanced perception process quantities to the corresponding modules of the human-machine hybrid enhanced perception model for calculation according to the input-output relationship required by each module.
[0036] Step two, online evaluation of explainability integration, in this step, the human-machine hybrid enhanced perception explainability evaluation data stream calculated in step three of the first step and the updated human-machine fusion perception knowledge base generated in the third step are integrated to form an explainability detection interface, wherein the human-machine hybrid enhanced perception explainability evaluation data stream objectively evaluates the explainability of the model, and the human-machine fusion perception knowledge base is used for subjective evaluation of the explainability of the human-machine hybrid enhanced perception model by the evaluator;
[0037] Step three, human-machine fusion perception atlas integration, in this step, the human-machine fusion perception atlas calculated in step four of the second step is recorded;
[0038] Step four, human-machine hybrid enhanced perception online output integration, in this step, the human-machine hybrid enhanced perception online output is formed.
[0039] The implementation links in the first step step three are as follows:
[0040] Link one, a method based on sensitivity analysis, this method analyzes the sensitivity of the model to different data in the test instance to determine whether it conforms to the human thinking mode or habit, first, a set of virtual test scenarios are prepared, perturbations are applied to the feature quantities in these scenarios, the degree of change in the output of each module of the human-machine hybrid enhanced perception model caused by the perturbation of different feature quantities is observed, and the process of system change caused by sensitive data is analyzed, and the human-machine hybrid enhanced perception explainability is objectively evaluated through scoring;
[0041] Link two, a method based on a proxy model, a proxy model refers to a simple model used to explain a complex model, common proxy models include linear models and decision tree models, these models have intuitive and easy-to-explain characteristics, the method based on the proxy model refers to using a number of representative typical scenarios to test each module of the human-machine hybrid enhanced perception model, constructing a proxy model through the output of each module through a set method, the proxy model provides a globally explainable layer on top of the nonlinear and non-monotonic model, and the human-machine hybrid enhanced perception explainability is analyzed by analyzing the white-box proxy model;
[0042] The explainability objective evaluation obtained according to formula (1) from the above two links is added according to the weight to obtain the human-machine hybrid enhanced perception explainability evaluation information stream.
[0043] The specific links included in the second step step two are as follows:
[0044] Step 1, driving environment perception, this step reconstructs the micro traffic flow in which the host vehicle is located, first receives the driving environment data output by step 1 of the first step, then models the driver's area of interest (ROI) according to the driver's eye movement and head movement data stream output by step 1 of the first step, and divides the perceived traffic participants into two parts: driver's attention objects in the ROI and driver's non-attention objects not in the ROI;
[0045] Step 2, driving scene understanding, this step extracts risk factors in the current driving scene from the driver's attention objects, and classifies them according to the kinematics information of the driver's attention objects output by step 1, and classifies them by numerical evaluation method, as shown in the following formula:
[0046]
[0047] In the formula, k1 and k2 are weight coefficients; TTC represents the time to contact when the host vehicle and the traffic participant travel at the current speed; THW represents the time to pass through the front vehicle's head when the host vehicle travels at the current speed;
[0048] The risk generated by the road element occurs in the turning process in the curve scene, including lane deviation caused by excessive turning and insufficient turning, which is classified by numerical evaluation method, as shown in the following formula:
[0049]
[0050] In the formula, k1, k2 and k3 are weight coefficients; CTLC represents the time to contact when the host vehicle travels at the current speed and yaw rate; STLC represents the time to contact when the host vehicle travels at the current speed and the heading angle is constant; TAD represents the time to contact when the host vehicle travels at the current speed and the heading angle is constant; TAD represents the time to contact when the host vehicle travels at the current speed and the heading angle is constant;
[0051] Step 3, driving situation reasoning, this step predicts the motion state and trajectory of the perceived risk factors in the next period of time, receives the risk factor information output by step 2, and makes predictions according to the risk factor information, which is a typical time series data prediction problem, which is realized by a long short-term memory (LSTM) network model;
[0052] Each network node unit of the LSTM network model includes a forget gate, an input gate and an output gate structure, which is used to simulate the memory processing process of human drivers in the prediction process, and the specific parameters are represented by the following formula group:
[0053]
[0054] In the formula, x t , h tand C t respectively represent input, output and transfer state, wherein the input is the kinematic state information of the target vehicle, and the output is the predicted kinematic state information of the target vehicle; f t , i t and o t respectively represent the output of the forget gate, the input gate and the output gate, for processing the past long-term memory and the newly added short-term memory; W f , W i , W C and W o respectively represent the weight matrix of the forget gate, the input gate, the transfer link and the output gate; b f , b i , b C and b o respectively represent the bias term of the forget gate, the input gate, the transfer link and the output gate;
[0055] The single-step prediction LSTM network model is extended to a multi-step LSTM network model. The method is to input the prediction value of the current time step as the input value of the next prediction into the LSTM network model through the input gate for iterative calculation until the required time point is reached. In order to ensure the effectiveness of the LSTM network model prediction, the prediction iteration cannot exceed 5 steps.
[0056] The specific steps of the third step of the second step are as follows:
[0057] Step one, driver state detection, in order to comprehensively reflect the driver state, the driver state detection includes two parts of the driver driving skill level detection and the driver real-time physiological state detection. Since the driver driving skill level is difficult to obtain directly, accurately and comprehensively, the driver driving task load detection is used instead of the driver driving skill level detection. The driver driving task load is calculated according to the host vehicle driver data stream and the host vehicle driver information stream output in step one of the first step. The driver real-time physiological state detection refers to judging the abnormal physiological state that will significantly affect the driver through the host vehicle driver data stream and the host vehicle driver information stream output in step one of the first step. Specifically, it includes distraction, fatigue, drunkenness, anxiety, sadness, anger, impulsiveness, depression, disease and drug influence state. The driver driving skill level and the driver real-time physiological state are summed in the form of weighted addition to obtain the driver state. According to the method of setting threshold, the driver state is divided into good and poor. The driver state detection is performed in a loop. According to the detection result: if the driver state is poor, the man-machine co-driving system quickly reduces the driving right of the driver to 0.1 under the premise of ensuring the stability of the vehicle, and the driving right allocation result is output to the man-machine co-driving system decision layer, and then the driver state detection is performed again in a loop. If the driver state is good, the loop is exited.
[0058] The second link is driving intention recognition. The driving intention of the driver is obtained and output to the decision layer of the man-machine co-driving system to realize the consistency of man-machine perception. The recognition of driving intention is solved by the method of hidden Markov model (HMM). The process is described as follows:
[0059] First, the driving intention of the driver is represented as Q=(Q lat lon ), Q lat ={-3,-2,-1,0,1,2}, where -3, -2, -1, 0, 1 and 2 represent the driving intention of the driver for turning, turning left, changing lane left, keeping straight, changing lane right and turning right, respectively; Q lon ={-2,-1,0,1}, where -2, -1, 0 and 1 represent the driving intention of the driver for parking, deceleration, keeping and acceleration, respectively. The eye movement and head movement signals of the driver output in step one and the information of the traffic participants output in step one in the second step are input as observation variables. After time discretization, the observation variable sequence O1, O2, …, O t is obtained. The driving intention of the driver is taken as the hidden state variable of the hidden Markov model. After time discretization, the hidden state variable sequence Q1, Q2, …, Q t is obtained. A certain amount of training data is collected through driving test, and the initial parameters of HMM are obtained according to the K-means clustering algorithm. The model is trained to improve the accuracy of the model. The trained model is used for driving intention recognition of the driver. As a kind of probability model, the method of HMM for recognizing driving intention is to calculate the probability of all driving intentions, and the one with the maximum probability is taken as the driving intention of the driver output by this link.
[0060] Step three, human-machine perception consistency comparison, in this step, the risk factors in the driver's non-attention object are analyzed to determine the human-machine perception consistency. After excluding all driver's attention objects in a period of time before the current time, if there is no risk factor in the driver's non-attention object, it proves that the driver and the human-machine co-driving system have consistent perception of the risk factors in the current driving environment, meeting the requirement of human-machine perception consistency, and outputting the human-machine fusion perception knowledge base update instruction as 0; if there is a risk factor in the driver's non-attention object, it does not meet the requirement of human-machine perception consistency, and then the risk factors in the driver's non-attention object are analyzed. The risk factors in the driver's non-attention object include three types: risk factors in the driver's perception blind area, risk factors that suddenly appear but have not been noticed by the driver, and risk factors that the driver considers safe. For the first and second types of risk factors, the driver is immediately warned, the driver's mistake is corrected, the human-machine perception consistency is achieved, and the human-machine fusion perception knowledge base update instruction is output as 0; for the third type of risk factor, the human-machine fusion perception knowledge base update instruction is output as 1, the incorrect knowledge in the human-machine fusion perception knowledge base in the third step is adjusted to adjust the risk factor judgment standard, and the human-machine perception consistency is achieved by correcting the error of the human-machine co-driving system.
[0061] The fourth step of the second step includes the following specific steps:
[0062] Step one, online output of human-machine fusion perception, the online output of human-machine hybrid enhanced perception model includes two parts: human-machine co-driving planning control input information flow composed of driving right allocation, driver's driving intention and traffic participant information for decision layer and control layer calculation, and driver's perception fusion information flow composed of risk factor warning information. In this step, the outputs of steps one, two and three in the second step are integrated according to the calculation results of step three in the second step, wherein the driver's perception fusion information flow is output to step four in the third step, and the human-machine co-driving planning control input information flow is output to the decision layer and control layer of the human-machine co-driving system for automatic driving calculation.
[0063] Step two, human-machine fusion perception atlas output, the human-machine hybrid enhanced perception model records each driving process through human-machine fusion perception atlas. The horizontal coordinate of the human-machine fusion perception atlas is the time of the driving process, and the vertical coordinate is the human-machine fusion perception state in the driving process, which specifically includes the driver's state, the driving right allocation value calculated and output by the perception layer, and the human-machine perception consistency comparison result. When each driving ends or the third type of risk factor described in step three of the third step appears, the established atlas will be output to the third step to build or update the online human-machine fusion perception knowledge base.
[0064] The third step of the third step includes the following specific steps:
[0065] Step 1: Perception Map Screening and Recording. First, the perception maps output in Step 4 of Step 2 are screened and recorded. This is used to analyze the output of the human-machine hybrid enhanced perception model constructed in Step 2 and explain the physical meaning of the human-machine hybrid enhanced perception model. The screening criterion for the perception maps in this step is: when the update instruction value of the human-machine fusion perception knowledge base output in Step 3 of Step 2 is 1, the perception map output in Step 4 of Step 2 is recorded at that moment.
[0066] Step 2: Update sequence of the human-machine fusion perception knowledge base. The update actions of the human-machine fusion perception knowledge base and the perception map are combined in chronological order to form the update sequence of the human-machine fusion perception knowledge base, which is used to analyze the interpretability of the update process of the human-machine fusion perception knowledge base.
[0067] Repeat steps one and two to obtain the update sequence L = {(m1, a...}}} 1) , (m2, a2), ..., (m n a n The perceptual map knowledge obtained in this step is m. i Represents the perceptual map recorded for the i-th time; a i This represents the update action of the human-machine fusion perception knowledge base recorded for the i-th time. In order to maintain the timeliness of the perception graph knowledge, the upper limit of the sequence length of L is set to 10. That is, when n<10, the newly recorded perception graph and the update action of the human-machine fusion perception knowledge base are directly added to the end of the perception graph knowledge L; when n=10, (m1, a1) is deleted first, and then the newly recorded perception graph and the update action of the human-machine fusion perception knowledge base are added to the end of the perception graph knowledge L.
[0068] The specific steps in the third step, step four, are as follows:
[0069] Step 1: Search for divergence points in human-machine fusion perception. This step uses a key feature search method to address the human-machine fusion perception inconsistencies that lead to these inconsistencies. First, key features for human-machine fusion perception divergences arising from each state in the FSM are established based on logical relevance. Then, feature matching is performed based on the human-machine perception inconsistency information output in Step 3 of Step 2. Successfully matched states are output as human-machine fusion perception divergence points. If there is more than one successfully matched state, each state is treated as a human-machine fusion perception divergence point. The set of all human-machine fusion perception divergence points is taken as the output of this step.
[0070] The second link is updating the human-machine fusion perception knowledge base. The updating of the human-machine fusion perception knowledge base is performed by the method of transfer training. First, according to the human-machine fusion perception divergence point output in the first link, the part of the human-machine fusion perception knowledge base that does not need to be updated is frozen. Then, the human-machine fusion perception knowledge base before updating is used as a pre-training model, the output layer is removed, and the remaining whole model is used as a feature extractor, which is applied to the human-machine fusion perception atlas output in step four in the second step. Then, the method of reinforcement learning is used to update each human-machine fusion perception divergence point by setting a reward function.
[0071] The fourth step four specifically includes the following links:
[0072] The first link is that the human-machine hybrid enhanced perception model outputs to the driver. In order to maintain the consistency of human-machine perception, the driver is assisted in perception and warned by the head-up display (HUD) in this link. First, the kinematic information of the traffic participants calculated in step one in the second step and the driver's attention object calculated in step two in the second step are integrated. According to the comparison result of the consistency of human-machine perception in step three in the second step, if the first type of risk factor and the second type of risk factor appear, the kinematic information of the driver's attention object, the first type of risk factor and the second type of risk factor is output on the HUD. If the first type of risk factor and the second type of risk factor do not appear, the kinematic information of the driver's attention object is output on the HUD. In order to play a warning role, the kinematic information of the first type of risk factor and the second type of risk factor displayed on the HUD is displayed in bright red, while the kinematic information of the driver's attention object is displayed in soothing white or blue.
[0073] The second link is that the human-machine hybrid enhanced perception model outputs to the human-machine co-driving system. This link receives the kinematic information of the traffic participants output in step one in the second step and the driver's intention information output in step three in the second step, integrates them in a time alignment manner, and filters useful information by logical judgment to form the information required by the decision layer and the control layer of the human-machine co-driving system. The decision layer and the control layer of the human-machine co-driving system are output respectively.
[0074] The beneficial effects of the present application are as follows:
[0075] This invention provides an interpretable human-machine hybrid augmented perception method based on next-generation artificial intelligence theory. By fusing sensitivity analysis and surrogate model methods, it achieves an objective quantitative evaluation of the interpretability of human-machine hybrid augmented perception. Based on contextual awareness theory, this invention constructs a driver perception model, simulating the driver's perception process of "driving environment perception - driving context understanding - driving context reasoning," ensuring the correct understanding of the driver's perception process by the human-machine co-driving system and greatly improving human-machine perception consistency. By establishing a human-machine perception consistency comparison model, this invention continuously iterates the human-machine hybrid augmented perception model, solving the problem of human-machine perception consistency. Based on a finite state machine model, this invention constructs a human-machine fusion perception knowledge base, forming a closed-loop human-machine hybrid augmented perception system framework by combining the human-machine hybrid augmented perception model and the human-machine fusion perception knowledge base. This invention achieves subjective evaluation of the interpretability of human-machine hybrid augmented perception through the human-machine fusion perception knowledge base, combining subjective and objective aspects to solve the interpretability problem of the human-machine hybrid augmented perception system. This invention has self-correction capabilities and can adjust the human-machine hybrid augmented perception system based on the human-machine perception consistency comparison results, improving the system's accuracy. Attached Figure Description
[0076] Figure 1 This is a schematic diagram of the overall steps and structure of the interpretable human-machine hybrid enhanced perception method described in this invention.
[0077] Figure 2 This is a schematic diagram of the overall architecture of the interpretable human-machine hybrid enhanced perception method described in this invention.
[0078] Figure 3 This is a schematic diagram of the overall architecture of the first step described in this invention.
[0079] Figure 4 This is a schematic diagram of the overall architecture of the second step described in this invention.
[0080] Figure 5 This is a schematic diagram of the overall architecture of the third step described in this invention.
[0081] Figure 6 This is a schematic diagram of the overall architecture of the fourth step described in this invention.
[0082] Figure 7 This is a schematic diagram of the LSTM network model for driving situation reasoning described in the second step of this invention.
[0083] Figure 8 This is an exemplary representation of the human-machine fusion perception map described in this invention.
[0084] Figure 9 This is a schematic diagram of the framework of the human-machine fusion perception knowledge base system described in the third step of this invention.
[0085] Figure 10 An example diagram of the result of the HUD display output by the human-machine hybrid enhanced perception model described in the fourth step of the present application to the driver. DETAILED DESCRIPTION
[0086] Referring to Figures 1 to 10 as shown:
[0087] The present application provides an interpretable human-machine hybrid enhanced perception method based on a new generation of artificial intelligence theory, and the application object is a human-machine co-driving system in which a driver and an automatic driving system cooperatively drive. The method is described as follows:
[0088] First step, integrate input data stream and information stream.
[0089] Second step, build a human-machine hybrid enhanced perception model.
[0090] Third step, build an interpretable online human-machine fusion perception knowledge base.
[0091] Fourth step, integrate output variables.
[0092] First step, integrate input data stream and information stream. In the first step, the input data stream and information stream of each module of the interpretable human-machine hybrid enhanced perception method based on a new generation of artificial intelligence theory described in the present patent are combed and integrated. The specific steps of the first step are described as follows:
[0093] Step one, multi-modal human-traffic mixed situation data stream and information stream. In this step, the multi-modal human-traffic mixed situation data stream and information stream are integrated through the multi-modal sensor data collected by the vehicle-mounted sensor. The multi-modal human-traffic mixed situation data stream includes two parts of driver data stream and driving environment data stream. The driver data stream refers to the data stream collected by each sensor of the driver monitoring system (DMS) located in the intelligent driving cabin, specifically including the driver manipulation data stream composed of steering wheel angle, accelerator and brake pedal opening and gear data stream, and the driver biological signal data stream composed of driver biological electrical signal and driver face image data stream. The driving environment data stream refers to the sensor data stream for sensing the surrounding traffic scene on the intelligent vehicle, including the inertial navigation data stream composed of IMU, GPS positioning and map data stream, the visual information data stream composed of forward-looking camera and surround-view camera image data stream, the radar point cloud data stream composed of laser radar point cloud, millimeter wave radar point cloud and ultrasonic wave radar data stream, and the other sensor information data stream composed of infrared sensor data and road surface monitoring sensor data stream.
[0094] Further, the multi-modal human-traffic mixed situation data stream is preprocessed to obtain a multi-modal human-traffic mixed situation information stream in this step. For the host vehicle driver data stream, driver eye movement and head movement information are extracted and blink frequency is calculated according to the image information of the camera in the DMS to form a host vehicle driver information stream. For the driving environment data stream, traffic participants are extracted and tracked according to visual information, and traffic participant information is extracted according to radar point cloud information to form a driving environment information stream.
[0095] Step two, internal data stream and information stream of the human-machine hybrid enhanced perception system. As shown in Figure 2 the human-machine hybrid enhanced perception system framework described in this patent includes the human-machine hybrid enhanced perception model constructed in the second step and the interpretable online human-machine fusion perception knowledge base constructed in the third step. This step integrates the data stream and information stream between the internal modules of the system. Specifically, these data streams and information streams include online driver perception model data stream and machine perception model data stream, and offline perception model knowledge information stream and perception atlas knowledge information stream.
[0096] Step three, interpretable evaluation information stream of the human-machine hybrid enhanced perception. This step calculates the online evaluation data stream of the model based on the new generation of artificial intelligence theory to improve the interpretability of the human-machine hybrid enhanced perception model. The system interpretability defined in this patent is the performance of the system having a clear physical meaning, as shown in the following formula:
[0097] Argmax E Q(E|M,H,D,T)(1)
[0098] In the formula, Q represents an explanatory evaluation equation; E represents a specific method of realizing interpretability; M represents a model parameter to be explained; H represents a driver model parameter; D represents the result data of a driving test; and T represents a driving task. Formula (1) indicates that the essence of interpretability is to seek an explanation method: given specific data and for a certain specific task, a specific driver has a maximum understanding of a specific model. This step is realized through the following two links:
[0099] Link one, method based on sensitivity analysis. This method analyzes whether the sensitivity of the model to different data in the test instance conforms to the human thinking mode or habit to analyze the interpretability of the model. First, prepare several groups of virtual test scenes, apply perturbations on the feature quantities in these scenes, observe the degree of change in the output of each module of the human-machine hybrid enhanced perception model caused by the perturbations of different feature quantities, and analyze the process of system change caused by sensitive data. The interpretability of the human-machine hybrid enhanced perception is objectively evaluated through scoring.
[0100] The second link is a method based on a proxy model. The proxy model refers to a simple model used to explain a complex model. Commonly used proxy models include linear models and decision tree models, which have intuitive and easy-to-explain characteristics. The method based on the proxy model refers to testing each module of the human-machine hybrid enhanced perception model using several representative typical scenarios, and constructing a proxy model through the output of each module by a certain method (such as regression and classification method). The proxy model provides a globally explainable layer on top of the nonlinear and non-monotonic model, and the human-machine hybrid enhanced perception explainability is analyzed by analyzing the white-box proxy model.
[0101] The explainability objective evaluation obtained according to formula (1) for the above two links is added by weight, respectively, to obtain the human-machine hybrid enhanced perception explainability evaluation information flow.
[0102] The second step is to construct a human-machine hybrid enhanced perception model. In this step, the human-machine hybrid enhanced perception is completed, and the human-machine perception consistency is ensured. The specific steps of the second step are described as follows:
[0103] Step one, machine perception model based on multi-sensor fusion. This step first receives the driving environment data stream and information stream output in step one in the first step, and reconstructs the traffic participants and the road in the surrounding environment in the three-dimensional space according to the size of the host vehicle and the installation position of the sensor, and restores the driving environment. The specific description of this step is as follows:
[0104] This step first establishes a state equation to analyze the traffic participants through the driving environment data stream and information stream output in step one in the first step, then predicts the state of the traffic participants at the current time based on the traffic participant information at the previous time, and updates the state of the traffic participants at the current time according to the sensor data at the current time to continuously update the state of the detection target. The prediction and update formulas of the state of the traffic participants at the current time are as follows:
[0105]
[0106]
[0107] In the formula, x t represents the state vector at time t predicted by the extended Kalman filter; u t represents the control vector at time t predicted by the extended Kalman filter; F t represents the state transition matrix; B t represents the control input at time t; w t represents the state vector noise term at time t with zero mean; Q t represents the covariance matrix of the noise term at time t; x t ’ represents the state variable at time t output by the extended Kalman filter; yt xk represents the state vector of sensor input at time t; P t P represents the uncertainty degree of system state at time t predicted by extended Kalman filter, expressed by the covariance of state vector; P t P represents the uncertainty degree of system state at time t output by extended Kalman filter; H t H represents the mapping matrix from state space to measurement space at time t; R represents the uncertainty matrix of measurement value, generally provided by the manufacturer of sensor.
[0108] For the traffic participants not perceived at the current time due to the limited perception range of the sensor, the state vector predicted by the extended Kalman filter is output as the output value of the traffic participant information.
[0109] Step two, driver perception model based on situational awareness theory. This step includes the following three links:
[0110] Link one, driving environment perception. This link reconstructs the micro traffic flow in which the host vehicle is located, first receives the driving environment data output by step one, and then models the region of interest (ROI) of the driver according to the driver eye movement and head movement data stream output by step one in the first step, divides the perceived traffic participants into two parts: the driver's attention objects in the ROI and the driver's non-attention objects not in the ROI.
[0111] Link two, driving scene understanding. This link extracts the risk factors in the current driving scene from the driver's attention objects, classifies them according to the kinematics information of the driver's attention objects output by link one, and classifies them by numerical evaluation method, for example, as shown in the following formula:
[0112]
[0113] In the formula, k1 and k2 are weight coefficients; TTC represents the time to contact when the host vehicle and the traffic participant travel at the current speed; THW represents the time to pass the cross section of the front vehicle when the host vehicle travels at the current speed.
[0114] The risk generated by road elements occurs in the turning process in the curve scene, including lane deviation caused by excessive turning and insufficient turning, which is classified by numerical evaluation method, for example, as shown in the following formula:
[0115]
[0116] In the formula, k1, k2 and k3 are weight coefficients; CTLC represents the time distance when the host vehicle maintains the current speed and the yaw rate and contacts the lane boundary line; STLC represents the time distance when the host vehicle maintains the current speed and the heading angle and contacts the lane boundary line; and TAD represents the time distance obtained by dividing the distance from the current position of the host vehicle to the lane boundary line by the speed.
[0117] Step three, driving situation reasoning. This step predicts the motion state and trajectory of the perceived risk factors in the next period of time, receives the risk factor information output by step two, and makes a prediction according to the risk factor information. This is a typical time series data prediction problem, which can be realized by a long short-term memory (LSTM) network model.
[0118] As shown in FIG. 3, it is a structure diagram of an LSTM network model that can realize driving situation reasoning. Each network node unit includes a forgetting gate, an input gate and an output gate structure, which are used to simulate the memory processing process of human drivers in the prediction process. The specific parameters can be represented by the following formula group: Figure 7
[0119]
[0120] In the formula, x t , h t and C t represent input, output and transfer state, respectively, wherein the input is the kinematic state information of the target vehicle, and the output is the predicted kinematic state information of the target vehicle; f t , i t and o t represent the outputs of the forgetting gate, the input gate and the output gate, respectively, which are used to process the past long-term memory and the newly added short-term memory; W f , W i , W C and W o represent the weight matrices of the forgetting gate, the input gate, the transfer link and the output gate, respectively; b f , b i , b C and b o represent the bias terms of the forgetting gate, the input gate, the transfer link and the output gate, respectively.
[0121] Further, since the driving situation reasoning link needs to predict multiple time steps in the future for a short time, it needs to be expanded from a single-step prediction LSTM network model to a multi-step LSTM network model. The method is to input the prediction value of the current time step as the input value of the next prediction through the input gate into the LSTM network model for iterative calculation until the required time point is reached. In order to ensure the effectiveness of the LSTM network model prediction, the prediction iteration cannot exceed 5 steps.
[0122] Step three, human-machine perception consistency comparison model. In this step, the perception consistency between the driver and the human-machine co-driving system is realized by comparing the perception results of the driver and the human-machine co-driving system. Specifically, it includes the following links:
[0123] Link one, driver state detection. In order to comprehensively reflect the state of the driver, the driver state detection described in this link includes two parts: driver driving skill level detection and real-time physiological state detection of the driver. Since it is difficult to directly, accurately and comprehensively obtain the driving skill level of the driver, the driving task load of the driver is used instead of the driving skill level of the driver. The driving task load of the driver is calculated according to the host vehicle driver data stream and the host vehicle driver information stream output in step one of the first step. The real-time physiological state detection of the driver refers to the judgment of abnormal physiological states that will significantly affect the driver through the host vehicle driver data stream and the host vehicle driver information stream output in step one of the first step, which specifically includes distraction, fatigue, drunkenness, anxiety, sadness, anger, impulsiveness, depression, illness and drug influence state. The driving skill level of the driver and the real-time physiological state of the driver are summed in the form of weighted addition to obtain the state of the driver, and the state of the driver is divided into good and poor according to the method of setting threshold. The driver state detection is performed in a loop, and according to the detection result: if the state of the driver is poor, the human-machine co-driving system quickly reduces the driving right of the driver to 0.1 under the premise of ensuring the stability of the vehicle, and the driving right allocation result is output to the decision layer of the human-machine co-driving system, and then the driver state detection is performed again. If the state of the driver is good, the loop is exited.
[0124] Link two, driving intention recognition. This link obtains the driving intention of the driver and outputs it to the decision layer of the human-machine co-driving system to realize the perception consistency between man and machine. The recognition of driving intention can be solved by the method of Hidden Markov Model (HMM), and the process is described as follows:
[0125] First, the driving intention of the driver is represented as Q=(Q lat , Q lon ), Q lat= {-3, -2, -1, 0, 1, 2}, where -3, -2, -1, 0, 1 and 2 represent the driver's lateral driving intention as U-turn, left turn, left lane change, keep straight, right lane change, right turn, respectively; Q lon = {-2, -1, 0, 1}, where -2, -1, 0 and 1 represent the driver's longitudinal driving intention as stopping, decelerating, keeping and accelerating, respectively. The driver's eye and head movement signals output in step one in the first step and the traffic participant information output in step one in the second step are input as observation variables, and after time discretization, the observation variable sequences O1, O2, …, O t . We take the driver's driving intention as the hidden state variable of the hidden Markov model, and after time discretization, the hidden state variable sequences Q1, Q2, …, Q t . A certain amount of training data is collected through driving tests, and the initial parameters of the HMM are obtained according to the K-means clustering algorithm. The model is trained to improve the accuracy of the model, and the trained model can be used for driver driving intention recognition. As a probability model, the HMM method for recognizing driving intention is to calculate the probability of all driving intentions, and the maximum probability is taken as the driver's driving intention output in this link.
[0126] Step three, human-machine perception consistency comparison. This link judges the human-machine perception consistency by analyzing the risk factors in the driver's non-attention objects. After excluding all driver's attention objects in a period of time before the current time, the risk factors are judged. If there is no risk factor in the driver's non-attention objects, it proves that the driver and the human-machine co-driving system have consistent perception results for the risk factors in the current driving environment, meeting the requirements of human-machine perception consistency, and outputting the human-machine fusion perception knowledge base update instruction as 0; if there is a risk factor in the driver's non-attention objects, it does not meet the requirements of human-machine perception consistency, and then the risk factors in the driver's non-attention objects are analyzed. The risk factors in the driver's non-attention objects include three types: risk factors in the driver's perception blind area, risk factors that suddenly appear but have not been detected by the driver, and risk factors that the driver considers safe. For the first and second types of risk factors mentioned above, the driver is immediately warned, and the human-machine perception consistency is achieved by correcting the driver's mistake, and the human-machine fusion perception knowledge base update instruction is output as 0; for the third type of risk factors mentioned above, the human-machine fusion perception knowledge base update instruction is output as 1 by the human-machine co-driving system, and the judgment standard of the risk factors is adjusted by adjusting the knowledge in the human-machine fusion perception knowledge base that leads to errors, and the human-machine perception consistency is achieved by correcting the errors of the human-machine co-driving system.
[0127] Step four, human-machine fusion perception output model. The output of the human-machine hybrid enhanced perception model is formed through the calculation of this step, and the specific steps are as follows:
[0128] Step 1: Online Output of Human-Machine Fusion Perception. The online output of the human-machine hybrid enhanced perception model includes two parts: a human-machine co-driving planning and control input information flow consisting of driving rights allocation, driver's driving intention, and traffic participant information used for calculations at the decision-making and control layers; and a driver perception fusion information flow consisting of risk factor warning information. In this step, the outputs of steps 1, 2, and 3 are integrated based on the calculation results of step 3. The driver perception fusion information flow is output to step 4 in step 3, and the human-machine co-driving planning and control input information flow is output to the decision-making and control layers of the human-machine co-driving system for autonomous driving calculations.
[0129] Step Two: Output of the Human-Machine Fusion Perception Map. The human-machine hybrid enhanced perception model described in this patent records each driving process using a human-machine fusion perception map. An exemplary representation of the human-machine fusion perception map is as follows: Figure 8 As shown, the horizontal axis represents the time of the driving process, and the vertical axis represents the human-machine fusion perception state during that driving session, specifically including the driver's state, the driving rights allocation value calculated and output by the perception layer, and the human-machine perception consistency comparison result. At the end of each driving session or when the third type of risk factor described in step three occurs, the established graph will be output to the third step for constructing or updating the online human-machine fusion perception knowledge base.
[0130] The third step is to construct an online human-machine fusion perception knowledge base. This knowledge base comprises two parts: perception graph knowledge and perception model knowledge. Perception graph knowledge is the raw data recorded after filtering the human-machine fusion perception graphs output in the second step. Perception model knowledge records the system decision-making process of the human-machine hybrid enhanced perception model used in the second step. In this third step, the framework of the human-machine fusion perception knowledge base system is first established. Then, based on the human-machine perception consistency comparison results output in step three of the second step and the output model of the human-machine fusion perception model output in step four of the second step, the human-machine fusion perception knowledge base is constructed and updated online. The specific steps of the third step are described below:
[0131] Step 1: Framework of the Human-Machine Fusion Perception Knowledge Base System. Based on the interpretability and usage requirements of the human-machine fusion perception knowledge base, a finite state machine (FSM) is used to build the system framework. The FSM analyzes and processes various types of information and combines experience to make behavioral decisions, such as... Figure 9The shown is a FSM-based human-machine fusion perception knowledge base system framework diagram. Due to the diversity of driving scenarios, the human-machine fusion perception knowledge base is composed of multiple FSMs according to different scenarios. Each state in the FSM is composed of two parts, namely semantic knowledge and human-machine hybrid enhanced perception model parameters. The semantic knowledge is obtained by clustering method, which is used for model explainability analysis. As shown in Figure 9 The semantic knowledge includes three parts of drivers, human-machine co-driving systems and driving scenarios, and the perception state of FSM is classified by fuzzy rule method. The model parameters are the human-machine hybrid enhanced perception model parameters set by the state of FSM, which are output to each step in the second step for decision operation.
[0132] In order to ensure the integrity of the human-machine fusion perception knowledge base, the perception state suitable for FSM is defined as a finite set U, the i-th perception state in FSM is defined as A i , and the number of perception states is n, which needs to satisfy the following formula:
[0133]
[0134] Formula (7) indicates that there is no intersection between each different perception state in FSM, and all perception states cover the entire perception state space.
[0135] Step two, perception model knowledge. In this step, in order to reduce the calculation amount of constructing online human-machine fusion perception knowledge base, prior knowledge is formed according to driving test and imitation learning is carried out. This step makes the constructed online human-machine fusion perception knowledge base have explainability, and the specific process is as follows:
[0136] This step first collects the decision trajectory sequence D={τ1,τ2,…,τ n} of human drivers through multiple driving tests, wherein each decision trajectory τ i represents a sequence of human-machine fusion perception state-human-machine fusion perception knowledge base update action of human driver's driving process, and the human-machine fusion perception state is taken as the feature and the human-machine fusion perception knowledge base update action is taken as the label. Through supervised learning method, D is taken as the training data set to generate the initial state human-machine fusion perception knowledge base. In order to make the generated human-machine fusion perception knowledge base have universality, multiple average level drivers are selected to collect training data set D for driving task.
[0137] Step three, perception atlas knowledge. This step includes the following two links:
[0138] Step 1: Sensing Map Screening and Recording. This step first screens and records the sensing maps output from Step 4 of Step 2. This is used to interpret the physical meaning of the human-machine hybrid enhanced perception model constructed in Step 2 by analyzing its output. The screening criterion for this step is: when the update instruction value of the human-machine fusion sensing knowledge base output from Step 3 of Step 2 is 1, the sensing map output from Step 4 of Step 2 is recorded at that moment.
[0139] Step Two: Update Sequence of the Human-Machine Fusion Perception Knowledge Base. The update actions of the human-machine fusion perception knowledge base are combined with the perception map in chronological order to form the update sequence of the human-machine fusion perception knowledge base, which is used to analyze the interpretability of the update process.
[0140] Repeat steps one and two to obtain the update sequence L = {(m1, a...}}} 1) , (m2, a2), ..., (m n a n The perceptual map knowledge obtained in this step is m. i Represents the perceptual map recorded for the i-th time; a i This represents the update action of the human-machine fusion perception knowledge base recorded for the i-th time. To maintain the timeliness of the perception graph knowledge, the upper limit of the sequence length of L is set to 10. That is, when n < 10, the newly recorded perception graph and the update action of the human-machine fusion perception knowledge base are directly added to the end of the perception graph knowledge L; when n = 10, (m1, a1) is first deleted, and then the newly recorded perception graph and the update action of the human-machine fusion perception knowledge base are added to the end of the perception graph knowledge L.
[0141] Step 4: Knowledge Base Update Mechanism. To achieve consistency in human-machine perception, the human-machine fusion perception knowledge base is continuously updated after its initial construction in Step 2 to improve this consistency. When the human-machine fusion perception map output from Step 4 in Step 2 is received, the human-machine fusion perception knowledge base is updated. This step is described in detail below:
[0142] Step 1: Search for Discrepancies in Human-Machine Fusion Perception. This step uses a key feature search method to identify discrepancies in human-machine fusion perception caused by inconsistencies. First, key features leading to discrepancies in human-machine fusion perception are established for each state in the Functional Model (FSM) based on logical relevance. Then, feature matching is performed based on the inconsistency information output in Step 3 of Step 2. Successfully matched states are identified as human-machine fusion perception discrepancies. If more than one state is successfully matched, each state is treated as a human-machine fusion perception discrepancy. The set of all human-machine fusion perception discrepancies is then taken as the output of this step.
[0143] The second link is the updating of the human-machine fusion perception knowledge base. The updating of the human-machine fusion perception knowledge base is performed by the method of transfer training. First, according to the human-machine fusion perception divergence points output in the first link, the part of the human-machine fusion perception knowledge base that does not need to be updated is frozen, and then the human-machine fusion perception knowledge base before updating is used as a pre-training model, the output layer is removed, and the remaining entire model is used as a feature extractor, which is applied to the human-machine fusion perception atlas output in step four in the second step. Then, by setting a reward function, each human-machine fusion perception divergence point is iteratively updated by the method of reinforcement learning.
[0144] The fourth step is to integrate the output variables. In the fourth step, the output variables of the system are integrated, and the internal output of the human-machine interaction and human-machine co-driving system and the explainability detection are formed. The specific steps of the fourth step are described as follows:
[0145] Step one is the integration of human-machine hybrid enhanced perception process variables. In this step, the inputs and outputs of all modules contained in the first step, the second step and the third step are integrated in a time-aligned manner, and the integrated human-machine hybrid enhanced perception process variables are output to the corresponding modules of the human-machine hybrid enhanced perception model for calculation according to the required input-output relationship of each module.
[0146] Step two is the integration of online explainability evaluation. In this step, the human-machine hybrid enhanced perception explainability evaluation data stream calculated in step three in the first step and the updated human-machine fusion perception knowledge base generated in the third step are integrated to form an explainability detection interface. The human-machine hybrid enhanced perception explainability evaluation data stream objectively evaluates the explainability of the model, and the human-machine fusion perception knowledge base is used for subjective evaluation of the explainability of the human-machine hybrid enhanced perception model by the evaluation personnel.
[0147] Step three is the integration of human-machine fusion perception atlas. In this step, the human-machine fusion perception atlas calculated in step four in the second step is integrated. An exemplary human-machine fusion perception atlas is shown in Figure 8 .
[0148] Step four is the integration of human-machine hybrid enhanced perception online output. In this step, the human-machine hybrid enhanced perception online output is formed, which specifically includes the following two links:
[0149] Link 1, the human-machine hybrid enhanced perception model outputs to the driver. In order to keep the consistency of human-machine perception, this link provides perception assistance and warning to the driver through the head-up display (HUD). First, the kinematic information of the traffic participants calculated in step 1 of the second step and the driver's attention object calculated in step 2 of the second step are integrated. According to the comparison result of the consistency of human-machine perception in step 3 of the second step, if the first and second risk factors appear, the kinematic information of the driver's attention object, the first risk factor and the second risk factor are output on the HUD; if there is no first and second risk factors, the kinematic information of the driver's attention object is output on the HUD. Further, in order to play a warning role, the kinematic information of the first and second risk factors displayed on the HUD is displayed in bright red, while the kinematic information of the driver's attention object is displayed in soothing white or blue.
[0150] Link 2, the human-machine hybrid enhanced perception model outputs to the human-machine co-driving system. This link receives the kinematic information of the traffic participants output in step 1 of the second step and the driver's intention information output in step 3 of the second step, integrates them in time alignment, and filters useful information through logical judgment, to form the information required by the decision layer and the control layer of the human-machine co-driving system, and outputs to the decision layer and the control layer of the human-machine co-driving system respectively.
Claims
1. An interpretable human-machine hybrid enhanced perception method based on a new generation of artificial intelligence theory, characterized in that: The method comprises the following steps: The first step is to integrate the input data stream and the information stream, and the specific steps are as follows: Step one, multi-modal human-traffic mixed situation data stream and information stream, multi-modal human-traffic mixed situation data stream and information stream are integrated through sensor data collected by vehicle-mounted sensors, multi-modal human-traffic mixed situation data stream includes two parts of driver data stream and driving environment data stream of host vehicle, the driver data stream of the host vehicle refers to the data stream collected by each sensor of the driver monitoring system in the intelligent driving cabin, specifically including the driver operation data stream composed of steering wheel angle, accelerator and brake pedal opening and gear data stream, and the driver biological signal data stream composed of driver biological electrical signal and driver face image data stream, the driving environment data stream refers to the sensor data stream for sensing the surrounding traffic scene on the intelligent vehicle, including the inertial navigation data stream composed of IMU, GPS positioning and map data stream, the visual information data stream composed of forward-looking camera and surround-view camera image data stream, the radar point cloud data stream composed of laser radar point cloud, millimeter wave radar point cloud and ultrasonic wave radar data stream, and the other sensor information data stream composed of infrared sensor data and road surface monitoring sensor data stream; The multi-modal human-traffic mixed situation data stream is preprocessed to obtain multi-modal human-traffic mixed situation information stream, for the driver data stream of the host vehicle, the driver eye movement and head movement information are extracted and the blink frequency is calculated according to the image information of the camera in the DMS, forming the driver information stream of the host vehicle; for the driving environment data stream, the traffic participants are extracted and tracked according to the visual information, and the traffic participant information is extracted according to the radar point cloud information, forming the driving environment information stream; Step two, human-machine mixed enhanced perception system internal parameter data stream and information stream, the human-machine mixed enhanced perception system framework includes the human-machine mixed enhanced perception model constructed in the second step and the interpretable online human-machine fusion perception knowledge base constructed in the third step, the data stream and the information stream between the internal modules of the human-machine mixed enhanced perception system are integrated, these data stream and the information stream include online driver perception model data stream and machine perception model data stream, and offline perception model knowledge information stream and perception atlas knowledge information stream; Step three, human-machine mixed enhanced perception interpretability evaluation information stream, the system interpretability is defined as the performance of the system having a clear physical meaning, as shown in the following formula: Argmax E Q(E|M, H, D, T) (1) In the formula, Q represents an explanatory evaluation equation; E represents a specific method for realizing interpretability; M represents a model parameter to be explained; H represents a driver model parameter; D represents a result data of a driving test; T represents a driving task performed, and formula (1) represents that the essence of the interpretability is to seek an explanation method: the method enables a specific driver to have a maximum understanding of a specific model for a specific task given specific data; The second step is to construct a human-machine mixed enhanced perception model, and the specific steps are as follows: Step one, machine perception model based on multi-sensor fusion, this step first receives the driving environment data stream and information stream output in step one of the first step, and reconstructs the traffic participants and the road in the surrounding environment in the three-dimensional space according to the size of the host vehicle and the installation position of the sensor, restores the driving environment, as follows: This step first establishes a state equation to analyze the traffic participants through the driving environment data stream and information stream output in step one of the first step, then predicts the state of the traffic participants at the current time based on the traffic participant information at the last time, updates the state of the traffic participants at the current time according to the sensor data at the current time to continuously update the state of the detection target, wherein the prediction and update formula of the state of the traffic participants at the current time are as follows: S t =H t P t H t T +R P t '=(1-K t H t )P t where x t represents the state vector at time t predicted by the extended Kalman filter; u t represents the control vector at time t predicted by the extended Kalman filter; F t represents the state transition matrix; B t represents the control input at time t; w t represents the noise term of the state vector at time t with zero mean; Q t represents the covariance matrix of the noise term at time t; x t represents the state variable at time t output by the extended Kalman filter; y t represents the state vector of the sensor input at time t; P t represents the degree of uncertainty of the system state at time t predicted by the extended Kalman filter, expressed by the covariance of the state vector; P t represents the degree of uncertainty of the system state at time t output by the extended Kalman filter; H t represents the mapping matrix from the state space to the measurement space at time t; R represents the uncertainty matrix of the measurement value, provided by the manufacturer of the sensor; For the traffic participants not perceived at the current time due to the limited perception range of the sensor, the state vector predicted by the extended Kalman filter is used as the output value for traffic participant information output; Step two, driver perception model based on situational awareness theory; Step three, human-machine perception consistency comparison model, in this step, the perception results of the driver and the human-machine co-driving system are compared to realize the consistency of human-machine perception; Step four, human-machine fusion perception output model, the output of the human-machine hybrid enhanced perception model is formed through the calculation of this step; Third step, build an interpretable online human-machine fusion perception knowledge base, the human-machine fusion perception knowledge base contains two parts of perception atlas knowledge and perception model knowledge, the perception atlas knowledge is the original data knowledge recorded after screening the human-machine fusion perception atlas output in the second step, and the perception model knowledge is the knowledge recording the system decision process of the human-machine hybrid enhanced perception model used in the second step, in this step, first build the human-machine fusion perception knowledge base system framework, then build and update the human-machine fusion perception knowledge base according to the human-machine perception consistency comparison results output in step three of the second step and the human-machine fusion perception model output model output in step four of the second step, the specific steps are as follows: Step one, human-machine fusion perception knowledge base system framework, according to the interpretability requirements and use requirements of the human-machine fusion perception knowledge base, a finite state machine FSM is used to build the human-machine fusion perception knowledge base system framework, FSM analyzes and processes various types of information, and makes behavior decisions combined with experience, the human-machine fusion perception knowledge base is composed of several FSMs according to different scenes, each state in FSM is composed of two parts, semantic knowledge and human-machine hybrid enhanced perception model parameters, wherein the semantic knowledge is obtained by clustering method and used for interpretability analysis of the model, the semantic knowledge includes three parts of driver, human-machine co-driving system and driving scene, the perception state of FSM is classified by fuzzy rule method, and the model parameters are the human-machine hybrid enhanced perception model parameters set by human according to the state of FSM, which are output to each step in the second step for decision operation; In order to ensure the integrity of the human-machine fusion perception knowledge base, the FSM applicable perception state is defined as a finite universal set U, the i-th perception state in the FSM is defined as A i , the number of perception states is n, and the perception state needs to satisfy the following formula: Formula (7) indicates that there is no intersection between each different perception state in FSM, and all perception states cover the entire perception state space; Step two, perception model knowledge, in this step, prior knowledge is formed according to driving test and imitation learning is carried out, this step makes the constructed online human-computer fusion perception knowledge base have explainability, the specific process is as follows: First, the decision trajectory sequence D = {τ1, τ2, …, τN} of human drivers is collected through multiple driving tests, wherein each decision trajectory τi represents a human-machine fusion perception state-human-machine fusion perception knowledge base update action pair sequence of a driving process of a human driver, and the human-machine fusion perception state is taken as a feature and the human-machine fusion perception knowledge base update action is taken as a label. n} wherein each decision trajectory τi represents a human-machine fusion perception state-human-machine fusion perception knowledge base update action pair sequence of a driving process of a human driver, and the human-machine fusion perception state is taken as a feature and the human-machine fusion perception knowledge base update action is taken as a label. i The initial state human-machine fusion perception knowledge base is generated by taking D as a training data set through a supervised learning method. In order to make the generated human-machine fusion perception knowledge base universal, multiple average-level drivers are selected to perform a driving task to collect the training data set D. Step three, perception atlas knowledge; Step four, knowledge base updating mechanism, in order to realize the consistency of human-computer perception, the human-computer fusion perception knowledge base is updated after the preliminary construction in step two to improve the consistency of human-computer perception, when receiving the human-computer fusion perception atlas output in step four in the second step, the human-computer fusion perception knowledge base is updated; Fourth step, integration of output variables, the output variables of the system are integrated, and an output for human-computer interaction, human-computer co-driving system internal output and explainability detection is formed, the specific steps are as follows: Step one, integration of human-machine hybrid enhanced perception process variables, in this step, the inputs and outputs of all modules contained in the first step, the second step and the third step are integrated in a time-aligned manner, and the integrated human-machine hybrid enhanced perception process variables are output to the corresponding modules of the human-machine hybrid enhanced perception model for calculation according to the required input-output relationship of each module; Step two, integration of explainability online evaluation, in this step, the human-machine hybrid enhanced perception explainability evaluation data stream calculated in step three in the first step and the updated human-computer fusion perception knowledge base generated in the third step are integrated to form an explainability detection interface, wherein the human-machine hybrid enhanced perception explainability evaluation data stream objectively evaluates the explainability of the model, and the human-computer fusion perception knowledge base is used for subjective evaluation of the explainability of the human-machine hybrid enhanced perception model by the evaluation personnel; Step three, integration of human-computer fusion perception atlas, this step integrates the human-computer fusion perception atlas calculated in step four in the second step; Step four, integration of human-machine hybrid enhanced perception online output, in this step, human-machine hybrid enhanced perception online output is formed.
2. The method of claim 1, wherein the method is based on a new generation of artificial intelligence theory. The implementation link in the first step step three is as follows: Link one, based on sensitivity analysis method, this method analyzes whether the sensitivity of the model to different data in the test instance conforms to the thinking way or habit of human being to carry out model explainability analysis, first, prepare an array of virtual test scenes, apply perturbation on the feature variables in these scenes, observe the change degree of the output of each module of the human-machine hybrid enhanced perception model caused by the perturbation of different feature variables, and analyze the process of system change caused by sensitive data, and then objectively evaluate the human-machine hybrid enhanced perception explainability through scoring; Link two, based on proxy model method, proxy model refers to a simple model used to explain a complex model, commonly used proxy models include linear model and decision tree model, these models have intuitive and easy-to-explain characteristics, the proxy model based method refers to using several representative typical scenes to test each module of the human-machine hybrid enhanced perception model, constructing a proxy model through the output of each module through a set method, the proxy model provides a global explainable layer on the nonlinear and non-monotonic model, and the human-machine hybrid enhanced perception explainability is analyzed by analyzing the white box proxy model; The two links are respectively added according to the formula (1) to obtain the human-machine hybrid enhanced perception explainability evaluation information flow.
3. The method of claim 1, wherein the method is based on a new generation of artificial intelligence theory. The second step two includes the following specific links: Link one, driving environment perception, this link reconstructs the micro traffic flow where the host vehicle is located, first receives the driving environment data output by step one of the first step, then performs ROI modeling according to the driver's eye movement and head movement data stream output by step one of the first step, and divides the perceived traffic participants into two parts: driver's attention objects in the ROI and driver's non-attention objects not in the ROI; Link two, driving scene understanding, this link extracts risk factors in the current driving scene from the driver's attention objects, divides the driver's attention objects according to the kinematics information output by link one, and divides them by a numerical evaluation method, as shown in the following formula: In the formula, k1 and k2 are weight coefficients; TTC represents the time to contact when the host vehicle and the traffic participant maintain the current speed; THW represents the time to contact when the host vehicle passes through the front vehicle's head section; The risk caused by road elements occurs in the turning process in the curve scene, including lane deviation caused by excessive turning and insufficient turning, which is divided by a numerical evaluation method, as shown in the following formula: In the formula, k1, k2 and k3 are weight coefficients; CTLC represents the time to contact when the host vehicle maintains the current speed and yaw rate; STLC represents the time to contact when the host vehicle maintains the current speed and heading angle; TAD represents the time to contact when the host vehicle maintains the current speed and heading angle; Link three, driving situation reasoning, this link predicts the motion state and trajectory of the perceived risk factors in the next period of time, receives the risk factor information output by link two, and predicts according to the risk factor information, which is a typical time series data prediction problem, which is realized by a long short-term memory (LSTM) network model; Each network node unit of the LSTM network model includes a forget gate, an input gate and an output gate structure, which is used to simulate the memory processing process of human drivers in the prediction process, and the specific parameters are represented by the following formula group: The single-step prediction LSTM network model is extended to a multi-step LSTM network model, which is achieved by inputting the prediction value of the current time step as the input value of the next prediction into the LSTM network model through the input gate for iterative calculation until the required time point is reached, in order to ensure the effectiveness of the LSTM network model prediction, the prediction iteration cannot exceed 5 steps. The second step three includes the following specific links: In the formula, x t , h t and C t respectively represent an input quantity, an output quantity and a transfer state quantity, wherein the input quantity is kinematic state information of a target vehicle, and the output quantity is predicted kinematic state information of the target vehicle; f t , i t and o t respectively represent a forget gate, an input gate and an output gate of an output, used for processing past long-term memory and newly added short-term memory; W f , W i , W C and W o respectively represent weight matrices of the forget gate, the input gate, the transfer link and the output gate; b f , b i , b C and b o respectively represent bias terms of the forget gate, the input gate, the transfer link and the output gate. 4. The method of claim 1, wherein the method is based on a new generation of artificial intelligence theory. Step one, driver state detection, in order to comprehensively reflect the state of the driver, the driver state detection includes two parts of the driver driving skill level detection and the driver real-time physiological state detection, since the driver driving skill level is difficult to obtain directly, accurately and comprehensively, the driver driving skill level detection is replaced by the driver driving task load detection, the driver driving task load is calculated according to the host vehicle driver data stream and the host vehicle driver information stream output in step one in the first step, the driver real-time physiological state detection refers to judging the abnormal physiological state of the driver which will significantly affect the driver through the host vehicle driver data stream and the host vehicle driver information stream output in step one in the first step, specifically including distraction, fatigue, drunkenness, anxiety, sadness, anger, impulsivity, depression, disease and drug influence state, the driver driving skill level and the driver real-time physiological state are summed up in the form of weighted addition to obtain the driver state, and the driver state is divided into good and poor according to the threshold setting method, the driver state detection is circularly carried out, and according to the detection result: if the driver state is poor, the man-machine co-driving system quickly reduces the driving right of the driver to 0.1 under the premise of ensuring the stability of the vehicle, and the driving right allocation result is output to the decision layer of the man-machine co-driving system, and then the driver state detection is circularly carried out again; if the driver state is good, the cycle is exited; Step two, driving intention recognition, this step obtains the driving intention of the driver and outputs it to the decision layer of the man-machine co-driving system to realize the consistency of man-machine perception, the recognition of the driving intention is solved by the method of hidden Markov model HMM, and the process is described as follows: First, let the driver's driving intention be represented as Q = (Q lat Q lon ), Q lat = {-3, -2, -1, 0, 1, 2}, where -3, -2, -1, 0, 1, and 2 represent the driver's lateral driving intentions as U-turn, left turn, left lane change, maintaining straight ahead, right lane change, and right turn, respectively; Q lon = {-2, -1, 0, 1}, where -2, -1, 0, and 1 represent the driver's longitudinal driving intention as stopping, decelerating, maintaining, and accelerating, respectively. The driver's eye and head movement signals output from step one of the first step and the traffic participant information output from step one of the second step are used as input observation variables. After time discretization, the observation variable sequence O1, O2, ..., O t The driver's driving intention is used as the hidden state variable in the Hidden Markov Model. After time discretization, the sequence of hidden state variables Q1, Q2, ..., Q is obtained. t A certain amount of training data is collected through driving tests. Then, the data is initialized and grouped according to the K-means clustering algorithm to obtain the initial parameters of the HMM. The model accuracy is improved through model training. The trained model is used to identify the driver's driving intention. As a probabilistic model, the HMM identifies the driving intention by calculating the probability of all driving intentions and taking the one with the highest probability as the driver's driving intention output in this step. Step three, man-machine perception consistency comparison, this step judges the man-machine perception consistency by analyzing the risk factors in the objects not concerned by the driver, all driver concerned objects in a period of time before the current time are excluded, and then the risk factors are judged, if there is no risk factor in the objects not concerned by the driver, it is proved that the driver and the man-machine co-driving system have consistent perception results for the risk factors in the current driving environment, which meets the requirement of man-machine perception consistency, and the man-machine fusion perception knowledge base updating instruction is 0; if there is a risk factor in the objects not concerned by the driver, it does not meet the requirement of man-machine perception consistency, and then the risk factors in the objects not concerned by the driver are analyzed, the risk factors in the objects not concerned by the driver include three types of risk factors in the driver's perception blind area, risk factors that suddenly appear but have not been perceived by the driver and risk factors that the driver thinks are safe, for the first and second types of risk factors, the driver is immediately warned, the driver's mistake is corrected, the man-machine perception consistency is achieved, and the man-machine fusion perception knowledge base updating instruction is 0; for the third type of risk factor, the man-machine co-driving system outputs the man-machine fusion perception knowledge base updating instruction as 1, the judgment standard of the risk factor is adjusted by adjusting the wrong knowledge in the man-machine fusion perception knowledge base in step three, the mistake of the man-machine co-driving system is corrected, and the man-machine perception consistency is achieved.
5. The method of claim 1, wherein the method is based on a new generation of artificial intelligence theory. The specific steps of step four of the second step include the following: Step one, online output of human-machine fusion perception, the online output of human-machine mixed enhanced perception model includes two parts, one is human-machine co-driving planning control input information stream composed of driving right allocation, driving intention of driver and traffic participant information for decision layer and control layer calculation, the other is driver perception fusion information stream composed of risk factor warning information, in this step, the outputs of step one, step two and step three in the second step are integrated according to the calculation results of step three in the second step, wherein the driver perception fusion information stream is output to step four in the third step, and the human-machine co-driving planning control input information stream is output to the decision layer and control layer of the human-machine co-driving system for automatic driving calculation; Step two, human-machine fusion perception atlas output, the human-machine mixed enhanced perception model records each driving process through the human-machine fusion perception atlas, the horizontal coordinate of the human-machine fusion perception atlas is the time of the driving process, and the vertical coordinate is the human-machine fusion perception state in the driving process, specifically including the state of the driver, the driving right allocation value calculated and output by the perception layer, and the human-machine perception consistency comparison result, when each driving ends or the third type of risk factor described in step three of the third step occurs, the established atlas is output to the third step for building or updating the online human-machine fusion perception knowledge base.
6. The interpretable human-machine hybrid augmented perception method based on a new generation artificial intelligence theory according to claim 1, characterized in that: The specific steps of step three of the third step include: Step one, perception atlas screening and recording, first, the perception atlas output in step four of the second step is screened and recorded, which is used to realize the physical meaning of the human-machine mixed enhanced perception model by analyzing the output of the human-machine mixed enhanced perception model built in the second step, the screening standard of the perception atlas in this step is that when the human-machine fusion perception knowledge base update instruction value output in step three of the second step is 1, the perception atlas output in step four of the second step is recorded; Step two, update time sequence of human-machine fusion perception knowledge base, the update action of the human-machine fusion perception knowledge base and the perception atlas are combined in time sequence to form the update time sequence of the human-machine fusion perception knowledge base, which is used to analyze the explainability of the update process of the human-machine fusion perception knowledge base; Repeat steps one and two to obtain the update sequence L = {(m1, a...}}} 1) , (m2, a2), ..., (m n a n The perceptual map knowledge obtained in this step is m. i Represents the perceptual map recorded for the i-th time; a i This represents the update action of the human-machine fusion perception knowledge base recorded for the i-th time. In order to maintain the timeliness of the perception graph knowledge, the upper limit of the sequence length of L is set to 10. That is, when n<10, the newly recorded perception graph and the update action of the human-machine fusion perception knowledge base are directly added to the end of the perception graph knowledge L; when n=10, (m1, a1) is deleted first, and then the newly recorded perception graph and the update action of the human-machine fusion perception knowledge base are added to the end of the perception graph knowledge L.
7. The interpretable human-machine hybrid augmented perception method based on a new generation artificial intelligence theory according to claim 1, characterized in that: The specific steps of step four of the third step include: Step one, human-machine fusion perception divergence point search, this step searches the human-machine fusion perception leading to human-machine perception inconsistency by the method of key feature search, first, the key features of human-machine fusion perception divergence generated by each state in the FSM according to the logical correlation are established, then the feature matching is performed according to the human-machine perception inconsistency information output in step three of the second step, the state with successful matching is output as the human-machine fusion perception divergence point, if there is more than one state with successful matching, each state is regarded as a human-machine fusion perception divergence point, and the set composed of all human-machine fusion perception divergence points is output as the output of this step; The second link is updating the human-machine fusion perception knowledge base. The human-machine fusion perception knowledge base is updated by the method of transfer training. First, according to the human-machine fusion perception divergence point output in the first link, the part of the human-machine fusion perception knowledge base that does not need to be updated is frozen. Then, the human-machine fusion perception knowledge base before updating is used as a pre-training model. The output layer is removed, and the remaining whole model is used as a feature extractor, which is applied to the human-machine fusion perception atlas output in step four of the second step. Then, the method of reinforcement learning is used to update each human-machine fusion perception divergence point by setting a reward function. 8.The method of claim 1, wherein the method is based on a new generation of artificial intelligence theory. The fourth step of the fourth step specifically includes the following links: The first link is that the human-machine hybrid enhanced perception model outputs to the driver. In order to maintain the consistency of human-machine perception, this link provides perception assistance and warning to the driver through the heads-up display (HUD). First, the kinematic information of the traffic participants calculated in step one of the second step and the driver's attention object calculated in step two of the second step are integrated. According to the comparison result of the consistency of human-machine perception in step three of the second step, if the first and second risk factors appear, the kinematic information of the driver's attention object, the first risk factor and the second risk factor are output on the HUD. If the first and second risk factors do not appear, the kinematic information of the driver's attention object is output on the HUD. In order to play a warning role, the kinematic information of the first and second risk factors displayed on the HUD is displayed in bright red, while the kinematic information of the driver's attention object is displayed in soothing white or blue. The second link is that the human-machine hybrid enhanced perception model outputs to the human-machine co-driving system. This link receives the kinematic information of the traffic participants output in step one of the second step and the driver's intention information output in step three of the second step. The information is integrated in time alignment and filtered by logical judgment to form the information required by the decision layer and the control layer of the human-machine co-driving system. The decision layer and the control layer of the human-machine co-driving system are output respectively.
Citation Information
Patent Citations
Man-machine co-driving intention fusion control method, device and equipment and storage medium
CN116198534A
Target sensing method, computer equipment, storage medium and vehicle
CN116429113A
Man-machine fusion sensing method for highly automatic driving
CN115027484A
Man-machine co-driving environment sensing method based on man-machine consistency
CN115524716A