A high-consistency human-machine hybrid decision-making method based on hybrid enhanced intelligence

By integrating multimodal data streams and constructing a highly consistent human-machine hybrid decision-making model based on a hybrid augmented intelligence approach, the adaptability and safety issues of autonomous driving systems in dynamic environments are solved, driver acceptability is improved, and a safe and reliable intelligent driving mode is achieved.

CN115564029BActive Publication Date: 2026-02-17JILIN UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211418353.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-14
Publication Date
2026-02-17
Estimated Expiration
2042-11-14

AI Technical Summary

Technical Problem

Existing autonomous driving systems are poorly adaptable and have weak safety in dynamic, open real traffic environments. The human-machine hybrid decision-making mode leads to low driver acceptability and weak safety, and lacks highly consistent human-machine co-driving system performance.

Method used

A hybrid augmented intelligence approach is adopted to integrate multimodal data streams and construct a human-machine hybrid augmented decision-making model. Through driver reasoning mechanisms, brain-like computing, and decision graphs, a highly consistent human-machine hybrid decision-making model is achieved, integrating driving rights assessment and online assessment results to optimize driving rights allocation.

Benefits of technology

It improves the safety and adaptability of autonomous driving systems in dynamic environments, enhances driver acceptability, and realizes a safe and reliable intelligent driving mode.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115564029B_ABST
    Figure CN115564029B_ABST
Patent Text Reader

Abstract

The application discloses a kind of high consistency man-machine hybrid decision-making methods based on hybrid enhanced intelligence, its method is: first, integrate input data stream and information flow;Second, build man-machine hybrid enhanced decision model;Third, build online man-machine decision knowledge base;Fourth, integrate output variable;Beneficial effect: greatly improve the safety and credibility of man-machine co-pilot system, and improve the acceptability of driver, realize safe and reliable man-machine hybrid decision mode;Realize comprehensive, reliable and rich decision information source set;Realize the decision effect with high driver acceptability and super brain mode;Greatly improve the adaptability of system to real traffic environment;Ensure that the independent process of module internal self-checking and self-optimization process of decision logic;Product has the function of self-checking and self-optimization in unknown driving situation and real traffic environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a man-machine hybrid decision method, in particular to a high-consistency man-machine hybrid decision method based on hybrid enhanced intelligence. BACKGROUND

[0002] From 2010 to 2022, automatic driving systems have gradually developed from principle prototypes to mature products, greatly reducing the driving load of drivers and improving the safety of driving tasks. At present, the automatic driving levels L1 and L2 have formed relatively mature products, and the L3 and L4 levels have formed relatively mature principle prototypes and gradually moved towards productization. The productization of the advanced driver-assisted driving mode of the automatic driving system into highly automatic driving and full unmanned driving lays a technical foundation for the advanced driver-assisted driving mode of the automatic driving system, and explores a relatively mature intelligent traffic system paradigm. With the continuous improvement of automatic driving systems and related intelligent traffic infrastructure, dynamic and open real traffic environments also pose higher challenges to the safety and adaptability of automatic driving systems. As a typical representative and constructive product mode of the transitional pseudo-unmanned driving system, the highly automatic driving system needs to have safe and reliable system performance. Therefore, the development of safe and reliable performance of the man-machine co-driving type automatic driving system and the design of man-machine consistency strategy have become the research frontier and key technology of the current automatic driving system.

[0003] At present, in the field of man-machine co-driving type automatic driving systems, realizing a man-machine hybrid decision mode with high driver acceptability and high safety is a key technology to realize a safe, reliable and comfortable intelligent driving mode. The man-machine hybrid decision mode represented by the previous man-machine game and simple man-machine arbitration has low driver acceptability and weak safety, which affects the safety and reliability of the highly automatic driving system. The man-machine hybrid decision mode with high man-machine consistency can play the advantages of man-machine through man-machine enhanced decision logic, and improve the adaptability and safety of the system. In the related research and principle prototype of high man-machine consistency, the man-centered decision logic has become one of the key technologies of man-machine decision, and at present, there is still a lack of decision data set supporting the man-centered decision and man-machine enhanced decision logic surpassing the man-machine decision capability, and the man-machine enhanced decision logic urgently needs performance improvement and "neck-stiffening" technology breakthrough.

[0004] There are few patents on high-consistency human-machine hybrid decision-making methods in automatic driving systems. Chinese patent CN201710084368.4 discloses a highway overtaking behavior decision-making method applied to an automatic driving vehicle, which realizes an overtaking decision-making method with human driving habits by establishing a driver's operation intention model. Chinese patent CN201710201086.8 discloses a generation method and device of a decision network model for vehicle automatic driving, which establishes a decision model based on deep learning and constructs a decision network based on data training. Chinese patent CN201711299043.4 discloses an automatic driving vehicle human-machine control right transfer method and system, which realizes the transfer logic of human-machine driving right and the decision logic of automatic driving vehicle by establishing a typical human-machine control right allocation mechanism. The above three patents can realize decision-making tasks for specific scenarios or specific data sets of high automatic driving systems, but cannot realize a human-machine co-driving system with high safety and driving person acceptability. SUMMARY

[0005] The main purpose of the present application is to solve the problems of poor adaptability and weak safety of unmanned vehicles in dynamic and open real traffic environments.

[0006] Another purpose of the present application is to overcome the low driving person acceptability and weak safety caused by human-machine hybrid decision-making in traditional human-machine game and simple human-machine arbitration mode.

[0007] Another purpose of the present application is to provide a safe and reliable intelligent driving mode.

[0008] The present application provides a high-consistency human-machine hybrid decision-making method based on hybrid enhanced intelligence to achieve the above purposes and solve the above problems.

[0009] The high-consistency human-machine hybrid decision-making method based on hybrid enhanced intelligence provided by the present application includes the following steps:

[0010] The first step is to integrate input data streams and information streams, and the specific process is as follows:

[0011] Step one, integrate multi-modal "human-traffic" mixed situation data stream;

[0012] Step two, integrate human-machine hybrid enhanced decision-making model internal parameter data stream. This step arranges and integrates the model internal parameters corresponding to each step of the human-machine hybrid enhanced decision-making model in the second step in time and event order, and uses the integrated model internal parameters as the data input for the knowledge construction of the second decision-making model in the third step. The input signal of this step is the model internal parameters corresponding to each step of the human-machine hybrid enhanced decision-making model constructed in the second step; the output signal is the human-machine hybrid enhanced decision-making model internal parameter data stream;

[0013] Step three, integrate human-machine hybrid enhanced decision-making driving right evaluation target information flow;

[0014] Step four, integrate human-machine hybrid enhanced decision-making online evaluation result information flow;

[0015] Second step, build human-machine hybrid enhanced decision-making model, the specific process is as follows:

[0016] Step one, build driver reasoning mechanism model;

[0017] Step two, build advanced "me" decision-making model based on brain-like computing;

[0018] Step three, build human-machine decision-making consistency comparison model;

[0019] Step four, build driving right subdivision model;

[0020] Third step, build online human-machine decision-making knowledge base, the specific process is as follows:

[0021] Step one, establish online human-machine decision-making knowledge base system framework, the specific content includes: human-machine decision-making knowledge mode, knowledge data structure between modes and data interaction logic between modes, human-machine decision-making knowledge mode contains step two decision-making model knowledge in step two, step three decision-making atlas knowledge, step four decision-making reasoning knowledge and step five new knowledge synthesis mechanism model, the knowledge data structure of step two decision-making model knowledge is object-oriented semantic mapping structure; the knowledge data structure of step three decision-making atlas knowledge is classification tree atlas structure; the knowledge data structure of step four decision-making reasoning knowledge is atlas structure based on data sequence; step five new knowledge synthesis mechanism model judges the new knowledge mode, and adopts the knowledge data structure of the corresponding knowledge mode, step three decision-making atlas knowledge and step four decision-making reasoning knowledge receive the knowledge content of step two decision-making model knowledge, and take it as the input of step three decision-making atlas knowledge and step four decision-making reasoning knowledge; step five new knowledge synthesis mechanism model receives the knowledge content of step two decision-making model knowledge, step three decision-making atlas knowledge and step four decision-making reasoning knowledge, and takes it as the basis for judging new knowledge mode, and outputs new knowledge to the corresponding step of the same new knowledge mode;

[0022] Step two, establish decision-making model knowledge;

[0023] Step three, establish decision-making atlas knowledge;

[0024] Step four, establish decision-making reasoning knowledge;

[0025] Step five, build new knowledge synthesis mechanism model;

[0026] Fourth step, integrate output variables, the specific process is as follows:

[0027] Step one, integrate advanced "class me" decision-making process quantity, this step integrates the output signals of each link included in step one, step two and step three in the second step in a time-aligned manner, for output to the decision-making online verification module in the highly automated driving system;

[0028] Step two, integrate knowledge base knowledge for online evaluation;

[0029] Step three, integrate human-machine hybrid decision-making driving right subdivision weight quantity, this step integrates the driving right allocation coefficient τ * output by link three corresponding to step four in the second step, the integration includes storing historical data of τ * , for output to the step five new knowledge synthesis mechanism model in the third step;

[0030] Step four, integrate human-machine hybrid decision-making expected control quantity.

[0031] The links included in each step of the first step are as follows:

[0032] The specific links of step one in the first step are as follows:

[0033] Link one, human situation assessment, this link assesses the human situation at the current time, including the extraction of the region of interest in the current scene by the driver and the driving intention, through the driving action of the driver at the current time and the bioelectric signal of the driver, the driving action of the driver includes the opening degree of the accelerator pedal, the brake master cylinder pressure caused by the opening degree of the brake pedal, the steering wheel angle and angular velocity, the eye movement of the driver and the head movement of the driver, the bioelectric signal of the driver includes the electrocardiogram, electroencephalogram, electromyogram and skin electricity signal of the driver, therefore, the input signal of this link is the driving signal of the driver and the bioelectric signal of the driver; the output signal is the human situation assessment result H ms ;

[0034] Link two, traffic situation assessment, this link assesses the traffic situation, including the dynamic participants in the dynamic scene and the driving rules and road conditions in the static scene, through the current dynamic scene signal and static scene signal, the dynamic scene signal includes dynamic traffic vehicle signal and dynamic pedestrian signal, the static scene signal includes lane line, traffic sign and curbstone signal, therefore, the input signal of this link is the dynamic scene signal and the static scene signal; the output signal is the traffic situation assessment result T fs ;

[0035] Link three, mixed situation fusion, this link fuses the output signals of link one and link two above, through H ms and T fsTime alignment and spatial coordinate conversion are performed to realize the hybrid situation fusion with multiple data modalities and scene elements. The input signal of this link is the human situation assessment result H output by the above link one ms and the traffic situation assessment result T output by the link two fs . The output signal is the hybrid situation fusion result M us .

[0036] The specific link of step three in the first step is as follows:

[0037] Link one, driver driving ability assessment. This link calculates the comprehensive control ability of the driver to the vehicle at the current time, i.e., assesses the driving ability of the driver at the current time, through the state and coupling of the driver-vehicle-road-environment at the current time. A typical system identification model is used as the driving ability assessment model. The input signal of this link is the human situation assessment result H ms , vehicle state signal, and vehicle-road coupling state signal. The output signal is the driver driving ability assessment result.

[0038] Link two, self-driving system driving ability assessment. This link calculates the comprehensive control ability of the automatic driving system to the vehicle at the current time, i.e., the self-driving system driving ability, through the traffic situation and all model parameters of the perception and decision levels in the automatic driving system at the current time. The input data of this link is assigned to the corresponding weight value, and a linear function with weight is used for calculation, thereby realizing the assessment of the self-driving system driving ability. The input signal of this link is the traffic situation assessment result T fs and the automatic driving system parameters. The output signal is the self-driving system driving ability assessment result.

[0039] Link three, driving right planning. The comprehensive control effects of the driver and the self-driving system to the vehicle at the current time are quantitatively assessed through the driver driving ability assessment result output by the above link one and the self-driving system driving ability assessment result output by the link two. Therefore, according to the driver driving ability assessment result and the self-driving system driving ability assessment result, the driving right distribution coefficient τ between the driver and the self-driving system at the current time can be calculated through the normalization calculation method. The input signal of this link is the driver driving ability assessment result and the automatic driving driving ability assessment result. The output signal is the driving right distribution coefficient τ.

[0040] The specific link of step four in the first step is as follows:

[0041] Link one, short-time domain online verification result, the link integrates the decision online verification module in the highly automated driving system which updates quickly in the short-time domain to obtain the short-time domain online verification result. The short-time domain verification signal output by the module is time-aligned and threshold-detected to obtain a clear and reasonable short-time domain online verification result. The input signal of the link is the short-time domain verification module signal; and the output signal is the short-time domain verification result.

[0042] Link two, long-time domain online optimization result, the link integrates the decision online verification module in the highly automated driving system which updates quickly in the long-time domain to obtain the long-time domain online verification result. The long-time domain verification signal output by the module is time-aligned and threshold-detected to obtain a clear and reasonable long-time domain online verification result. The input signal of the link is the long-time domain verification module signal; and the output signal is the long-time domain verification result.

[0043] Link three, online evaluation information flow integration, the link integrates the short-time domain verification result output by link one and the long-time domain verification result output by link two to obtain the online evaluation result. The integration process mainly includes time domain category labeling and redundant time domain elimination. The input signal of the link is the short-time domain verification result and the long-time domain verification result; and the output signal is the online evaluation result. er .

[0044] The steps included in the second step are as follows:

[0045] The specific links of step one in the second step are as follows:

[0046] Link one, multi-target learning framework, a multi-learning target based on the safety, comfort, functionality and maneuverability of human-machine hybrid decision is established. Local subgraphs are extracted in the global graph of the driver reasoning mechanism model, and a learning framework for random walk according to relationship clustering and coupling results between local subgraphs is established, which specifically includes three parts: reasoning mechanism model definition, entity relationship standard formulation and entity relationship clustering.

[0047] Link two, global graph random walk, a global graph of the driver reasoning mechanism is established, and the reachability between each entity pair in the global graph is judged to solve the reasoning result corresponding to the global graph, which specifically includes three parts: global graph establishment, entity reachability calculation and global reasoning result calculation.

[0048] Link three, local subgraph random walk, a specific local relationship subgraph of the driver reasoning mechanism is extracted from the global graph of the driver reasoning mechanism to realize random walk, which specifically includes three parts: local subgraph establishment, entity transition probability matrix calculation and local reasoning result calculation.

[0049] Step four, fusion reasoning, the global reasoning result and the local reasoning result obtained in the above step two and step three are uniformly distributed area matching, the reasoning result is fused by using nonlinear mapping logic, specifically including two parts: reasoning result normalization calculation and fusion reasoning result calculation;

[0050] Human situation assessment result H m And human-machine decision consistency comparison result C hm The driver reasoning online data flow F lf , that is, F lf ={H ms ,C hm};The decision reasoning knowledge base K df And the decision graph knowledge base K dm The driver reasoning offline knowledge flow F fk , that is, F fk ={K df ,K dm}Therefore, the calculation formula of the driver reasoning mechanism model is shown in the following formula (1):

[0051]

[0052] In the formula, G m And L m represent the global graph and the local subgraph respectively, f(G m ,L m ) is the function of DIMM, f(F lf ) represents the subset function of F fk , therefore, the DIMM model is a reasoning model based on a multi-objective learning framework and a random walk pattern, in a specific driving scene at a specific time, the independent relationship γ is compared through entity correlation, multiple iterations and clustering of the cluster |R γ | are realized, and the shared feature value set C γ containing all feature values c γ corresponding to the new cluster formed by updating is updated, the calculation formula of the cluster similarity function sim(C γ,m ,C γ,n ) is shown in the following formula (2):

[0053] (2)

[0054] In the formula, the operator Π represents the product of each element in the set C γ , and the subscripts m and n of C γ represent two different C γ with numbers m and n, on the basis of the similarity between |R γ |, a joint learning classification model is established to couple and constitute the path R of γ in each |R γ |r , the classifier structure function f cl (R r ) and the corresponding joint relationship learning model are calculated as shown in the following formula (3):

[0055]

[0056] In the formula, μ1 and μ2 are regularization coefficients, ω k and ω0 are weight coefficients and their reference values, b k and b0 are classification structure bias coefficients and their reference values, d k is a weight vector bias coefficient, and the function L(R ri,p , R ri,q ) is the training loss function of f cl (R r ), in which the subscripts p and q respectively represent two different R r s numbered p and q, the subscript i represents the number of R r , N k is the number of R r s, the subscript k represents the number of d k , and K is the number of d k s. The implementation cluster |R γ | and the corresponding path R r obtained by clustering the entity relationship are used as constraints to calculate the above-mentioned link two global graph random walk, which extracts each relationship r el and the corresponding c γ in G m , establishes a global relationship feature model, and G m is defined as G m,i ={g gm,i ={h rg,i , R gm,i , ra m,i}, i = 1, 2,..., s}, wherein g m is a subgraph of G m,i , the subscripts i and s respectively represent the subgraph number and the total number of subgraphs, h gm,i , R rg,i and ra gm,i in g m respectively represent the head entity, path and tail entity of the effective subgraph, the reachability p gm,i of h rg,i to ra gm,i through R re in G gm,i is calculated as shown in the following formula (4):

[0057]

[0058] In the formula, sl={h gm,i ∪ra gm,i ,i=1,2,...,s},R ra,i To be with ra gm,i The directly corresponding R rg,i The tail relation elements in the text, sl and ra gm,i The similarity function sim(sl,ra) gm,i The calculation method of ) and the calculation formula (2) for C γ,m and C γ,n The similarity function between them is calculated in the same way, and α represents R. rg,i The corresponding weight matrix allows the global graph random walk model to be expressed as: f(G m )=α·p re The logistic regression algorithm is used to analyze the model f(G). m The parameters are trained using the given parameters, and the sigmoid function is chosen as the normalization function for the results. The normalized global inference result p g The calculation formula is shown in equation (5) below:

[0059] (5)

[0060] L m For G m A subset of is defined as L m ={l m,i ={h lm,i ,R rl,i ,ra lm,i}, i = 1, 2, ..., z}, where l m,i For L m A subgraph, where the subscripts i and z represent the subgraph number and the total number of subgraphs, respectively, l m,i hl in m,i R rl,i and ra lm,i These represent the head entity, path, and tail entity of a valid subgraph, respectively. Therefore, for L... m When performing random walk computation, the space complexity of the computation is reduced, and L can be directly computed. m The transition probability matrix T between different entities M This leads to the corresponding local inference results, T M The calculation formula is shown in equation (6) below:

[0061]

[0062] In the formula, N hl,i and N ral,i According to hl m,i and ra lm,iDiagonal matrix constructed, sp is T M Number of transition steps, M l,i L at the sp step m Corresponding adjacency matrix, T M Corresponding element T of the a row b column M [a,b] represents that the random walk starts from hl m,i After sp steps, jump to ra lm,i Probability, using p l Indicates the local reasoning result evaluation result of R rl,i The calculation formula of p l The calculation formula is shown in the following formula (7):

[0063]

[0064] Fusion and normalization of p g And p l The fusion reasoning result p f The calculation formula is shown in the following formula (8):

[0065] (8)

[0066] In the formula, δ represents the fusion reasoning stability coefficient, which is used to balance the contribution proportion of p g And p l , and thus the step one driver reasoning mechanism model in the second step outputs the fusion reasoning result;

[0067] The specific link of step two in the second step is as follows:

[0068] Link one, neuron group model, establish neuron group model, provide personalized classification basis for the following link four corresponding "I" decision model, specifically including: neuron group model definition, feature extraction, classification based on stimulation and highlight matrix four parts;

[0069] Link two, deep convolutional network, through deep learning network, provide policy-oriented fitting basis for the following link three corresponding reinforcement learning model, specifically including: behavior network structure definition and evaluation network structure definition four parts;

[0070] Link three, reinforcement learning model, through the reinforcement learning model, realize the complex decision mode of the automatic driving system, find the optimal strategy, calculate the optimal action in the reinforcement learning model, specifically including: state definition, reward function and policy gradient three parts;

[0071] Step 4, “I-like” Decision Model: This step uses the personalized classification criteria obtained in Step 1 above, and combines the deep reinforcement learning processes in Steps 2 and 3 above, to integrate “I-like” decision data at the data level. By judging whether the iterative effect reaches the preset threshold, the final online decision result of the autonomous driving system is output. Specifically, it includes two parts: “I-like” data fusion and threshold judgment.

[0072] T fs and C hm The online classification data stream consisting of neuron groups F cf That is, F cf ={T fs C hm};K dd and K dm Composing the knowledge flow F for offline training of autonomous driving ak That is, F ak ={K dd ,K dm The computational formula for the neuron group model is shown in equation (9) below:

[0073] NGM=f(F cf (9)

[0074] According to F ak Specific features are extracted and learned using incremental learning rules based on Heblin and Haberlian learning, with a predefined N. G The category labels of the dimension are used, and the extracted specific features are used as conditional stimuli to build a personalized neuron group model belonging to a specific classification dimension. The incremental learning rule calculation formula is shown in the following formula (10):

[0075]

[0076] In the formula, ΔCHL, ΔHeb, and Δβ represent the Heb learning rule, the Heb learning rule, and the incremental learning rule, respectively, β represents the synaptic matrix corresponding to the incremental learning rule, and g j For presynaptic activation, h i For postsynaptic activation, the subscripts i and j correspond to F ak The specific feature element extracted from the i-th row and j-th column, where + and - represent the home and de-home phases respectively, ζ is the weight coefficient, and κ is the learning rate, uses the kWTA function to obtain the sparse distribution representation in the feature matrix, and extracts the first r activation units, corresponding to the inhibition function f. r The calculation formula is shown in equation (11) below:

[0077]

[0078] In the formula, χ is the suppression threshold, which ensures that the first r activation units are in the activation mode, and the activated hi By calculating the i-th row of β and g j The standardized dot product is obtained, and the corresponding calculation formula is shown in equation (12) below:

[0079]

[0080] Traverse and calculate h i Then, h i The element h with the largest value in the middle i,max Defined as a maximally reactive neuron if and only if r = h i,max When, the corresponding β ij Based on the learned feature elements, N is established. G Dimension M c ={m c,1 ,m c,2 ,...,m c,NG}, where m c For N G M after dimensional classification c Sub-model;

[0081] Will pass through M c K under the corresponding sub-model classification dd and K dm Subset K dd,NG and K dm,NG As the training data for the deep convolutional network in stage two and the reinforcement learning model in stage three, the deep convolutional network and the reinforcement learning model together constitute the deep reinforcement learning model DRLM. The calculation formula corresponding to the definition of the DRLM model is shown in the following formula (13):

[0082] DRLM={S DR A DR ,P DR ,R DR ,λ} (13)

[0083] In the formula, S DR For the vehicle state space of the autonomous driving system, A DR For the control action space of autonomous driving, P DR Let R be the state transition probability distribution of the autonomous driving system. DR Let λ be the reward function, and λ = [λ1, λ2, λ3, λ4] be the discount factor. DRIM's R... DR The calculation formula is shown in equation (14) below:

[0084]

[0085] In the formula, R safety R goal R law and Rcomft Let S represent the safety reward function, time reward function, traffic regulation reward function, and comfort reward function, respectively. These correspond to the four driving objectives of the autonomous driving system for the driving task: safety, maneuverability, traffic regulations, and driver comfort. This is to achieve parameterized estimation of continuous data sequences in autonomous driving decision-making, i.e., in S... DR When P is in continuous data sequence mode DR For accurate estimation, a deep convolutional network based on policy gradient is used to perform π(a) estimation. RDL ,s RDL The search for θ in deep convolutional networks mainly includes behavior networks L(θ). Q ) and rating networks▽ θμ L(θ Q The corresponding network definitions and calculation formulas are shown in equations (15) and (16) below:

[0086]

[0087]

[0088] In the formula, operator E is the expectation function, subscripts t and t+1 represent the current time step and the next time step, respectively, and θ Q and θ μ These represent two nonlinear estimators for two neural network structures, where Q is the action-value function, representing the sum of the states S and S. DR Take action A DR The expected discount return obtained; μ is the value that makes a RDL =μ(s) RDL ;θ μ The mapping function constructed to satisfy the condition ▽ θμ L(θ Q The gradient update method is used to update π(a) RDL ,s RDL Optimization of )

[0089] The calculated π(a) RDL ,s RDL Data fusion and updating are performed in the memory pool to ultimately obtain π. * (a RDL ,s RDL The calculation formula for the update mode is shown in equation (17):

[0090]

[0091] In the formula, ; positive numbers;

[0092] Therefore, step two in the second step, based on the advanced "self-like" decision model of brain-like computing, outputs the "self-like" decision results of the autonomous driving system.

[0093] The specific steps in step three of the second step are as follows:

[0094] Step 1: Driver Decision Map. The input signal for this step is the fusion inference result p output from Step 4 of Step 1 in Step 2. f And the human situation assessment result H corresponding to the first step of step one in the first step. ms And according to the decision graph knowledge base K corresponding to step three in step three. dm The knowledge representation rules are ultimately expressed in the form of a classification tree graph structure, which represents p. f H ms And integrate driver control signals into a driver decision map;

[0095] Step 2, Autonomous Driving Decision Map: The input signal for this step is the optimal decision strategy π output from Step 2, which corresponds to Step 3 in Step 2. * (a RDL ,s RDL And the traffic situation assessment result T corresponding to the first step of step one in the first step. fs Similarly, according to the decision graph knowledge base K corresponding to step three in step three, dm The knowledge representation rules are ultimately expressed in the form of a classification tree graph structure, which represents π. * (a RDL ,s RDL ), T fs And the control signals of the autonomous driving system are integrated into the autonomous driving system decision map;

[0096] Step 3, Human-Machine Graph Comparison and Prediction: The input signals for this step are the hybrid situation fusion result output from Step 3 of Step 1 (corresponding to Step 1 in Step 1), the driver decision graph output from Step 1 of Step 2 (corresponding to Step 3 in Step 2), and the autonomous driving system decision graph output from Step 2; the output signal is the human-machine decision consistency rate C. DK This step first predicts the spatiotemporal evolution of the vehicle's state and trajectory within a short time domain (0 to 10 seconds from the current time) under driver-only driving conditions using the driver's decision graph. Simultaneously, it predicts the same spatiotemporal evolution under autonomous driving system-only driving conditions using the autonomous driving system's decision graph. Next, it calculates the similarity between the vehicle's state and trajectory under both driver-only and autonomous driving system-only driving conditions within the same short time domain, and uses this similarity result as the output signal C for this step. DK ;

[0097] The specific steps in step four of the second step are as follows:

[0098] Step 1: Subdividing the Criterion Library. This step stores the subdivided rules for human-machine driving rights corresponding to the human-machine hybrid decision-making layer of highly automated driving systems in an offline storage manner. The subdivided criterion library mainly includes: traffic rules library, mobility rules library, safety rules library, and comfort rules library. The traffic rules library stores the set of traffic rules corresponding to urban traffic; the mobility rules library stores the set of rules constructed to ensure vehicle driving efficiency; the safety rules library stores the set of rules to ensure that the vehicle has safe longitudinal and lateral driving performance in emergency situations; and the comfort rules library stores the set of rules to ensure that the driver and passengers are in a comfortable state during vehicle operation.

[0099] Step 2, Constraint Rules: This step stores the constraint rules used to constrain the configuration and value range of the knowledge and model intrinsic parameters of the human-machine hybrid decision-making layer in an offline storage manner. These mainly include: constraint rules on the value range of the decision model knowledge in Step 2 of Step 3, threshold constraint rules on the vehicle state when only the driver is driving, and threshold constraint rules on the vehicle state when only the autonomous driving system is driving.

[0100] Step 3, Driving Rights Optimization Algorithm: The input signals for Step 3 in Step 4 of Step 2 are the driving rights allocation coefficient τ from Step 3 of Step 1, and C from Step 3 of Step 2. DK The second step involves the subdivision criteria corresponding to the first stage of step four, the constraint rules corresponding to the second stage of step four, the driver's control signals at the current moment, and the control signals of the automated driving system at the current moment. The output signals are the driver's driving weight value and the automated driving system's driving weight value for the highly automated driving system. This stage first establishes C. DK The two-dimensional linear programming plane corresponding to the output τ of stage three in step one is used. Then, the subdivision criteria of stage one output corresponding to step four in step two and the constraint rules of stage two output are used to regulate and constrain the driver's control signals and the control signals of the autonomous driving system at the current moment, and finally the optimized driving rights allocation coefficient τ is obtained. * .

[0101] The steps in the third step are as follows:

[0102] The specific steps in step two of the third step are as follows:

[0103] Step 2, the ontology design of decision model, this step adopts the meta ontology development mode from the bottom to the top, taking the main class of meta ontology as the core, defining three top classes of top model class, top data class and logic expression class respectively, and realizing the design of decision model ontology, the input signal of this step is the human-machine hybrid enhanced decision model internal parameter data stream corresponding to step 2 in the first step, and the output signal is the primary decision model knowledge base with the top classes of meta ontology;

[0104] Step 2, the ontology design of decision model, this step adopts the meta ontology development mode from the bottom to the top, taking the main class of meta ontology as the core, defining three top classes of top model class, top data class and logic expression class respectively, and realizing the design of decision model ontology, the input signal of this step is the human-machine hybrid enhanced decision model internal parameter data stream corresponding to step 2 in the first step, and the output signal is the primary decision model knowledge base with the top classes of meta ontology;

[0105] Step 3, the rule design of model knowledge, this step reasons and queries the concept model of meta ontology, and constructs a semantic rule base through two ways of knowledge path management rule and knowledge action operation rule, the input signal of this step is the primary decision model knowledge base output by step 2 corresponding to step 2 in the third step, and the output signal is the updated decision model knowledge base with semantic rules;

[0106] Step 4, the architecture design of model knowledge, this step establishes a semantic web framework as a knowledge engine to integrate dynamic knowledge meta ontology and static knowledge meta ontology, in addition, this step establishes a semantic middleware to process the multi-modal human-machine hybrid enhanced decision model internal parameter data stream into a unified meta ontology structure through semantic mapping, the input signal of this step is the human-machine hybrid enhanced decision model internal parameter data stream corresponding to step 2 in the first step and the updated decision model knowledge base output by step 3 corresponding to step 3 in the first step, and the output signal is the decision model knowledge K dd ;

[0107] The specific four steps of step 3 in the third step are as follows:

[0108] Step 1, the ontology construction of knowledge graph, this step establishes a knowledge graph ontology for decision graph knowledge, and synthesizes an effective knowledge ontology for human-machine hybrid decision through field ontology mode, the input signal of this step is the human situation assessment result H ms output by step 1 corresponding to step 1 in the first step, the traffic situation assessment result T fs output by step 2 corresponding to step 1 in the first step, and the new mode decision reasoning knowledge output by step 5 in the third step, and the output signal is the effective knowledge ontology set, considering H ms and T fsThe time-continuous data sequence is present in the domain, and the effective knowledge ontology synthesized by the domain ontology model is divided into five elements of concept, relation, function, axiom and individual;

[0109] The second link is the semantic rule making of the atlas. The effective knowledge ontology is regulated at the semantic level by specifying semantic rules. The basic semantic rules are established by combining abstract syntax and concrete syntax, and four reasoning functions of consistency checking, classification, identification and prediction are added to the basic semantic rules to realize the semantic rules with reasoning mechanism. The second link has no input signal, and the output signal is the semantic rules;

[0110] The third link is the semantic mapping of the atlas. The effective knowledge ontology set and the semantic rules are mapped in the form of a relation table to form a decision atlas case. Different interacting effective knowledge ontologies are identified by unique target addresses. The reasonable mapping of multiple ontology relations is realized by mapping the relation classes composed of the effective knowledge ontologies and their corresponding semantics. Finally, a decision atlas case library composed of multiple relation classes, semantic rules and their mutual mapping relations is formed. The input signal of the third link is the effective knowledge ontology set output by the first link corresponding to step three in the third step and the semantic rules output by the second link, and the output signal is the decision atlas case library;

[0111] The fourth link is the case knowledge calling. The case access mechanism is developed according to the mode of searching the effective knowledge ontology in the decision atlas case library first, then searching the similar relation, and finally searching the similar semantics to realize the calling of specific cases in the decision atlas case library. The case access mechanism developed by the fourth link and the decision atlas case library together constitute the decision atlas knowledge base. The input signal of the fourth link is the decision atlas case library output by the third link corresponding to step three in the third step, and the output signal is the decision atlas knowledge base K dm ;

[0112] The specific links of step four in the third step are as follows:

[0113] The first link is the knowledge rule oriented model. The knowledge rule oriented model is established to convert the pre-collected driver offline database into a driver offline knowledge base constrained by knowledge rules. The input signal of the first link is the driver offline database information and the new mode decision reasoning knowledge output by step five in the third step. The output signal is the knowledge rule oriented function and the driver offline knowledge base;

[0114] Step two, knowledge reasoning vector modeling, this step establishes a knowledge reasoning vector based on the vehicle dynamics and the vehicle-road coupling characteristics, solves the reasoning and prediction results of the vehicle state change corresponding to the driver's offline knowledge base in the form of knowledge data structure, and integrates it into a knowledge reasoning vector set. The input signal of this step is the knowledge rule oriented function output by step one in step four of the third step, and the driver's offline knowledge base; the output signal is the knowledge reasoning vector set;

[0115] Step three, hierarchical logic reasoning classification, this step establishes a hierarchical logic reasoning classification method based on the danger level of the scene and the driving mode, and divides the knowledge reasoning vector set into a typical knowledge reasoning vector subset under this classification method. The input signal of this step is the knowledge reasoning vector set output by step two in step four of the third step; the output signal is the classified knowledge reasoning vector subset;

[0116] Step four, reasoning knowledge architecture generation, this step arranges the knowledge reasoning vector subset according to the priority, and sets the search logic algorithm for the knowledge reasoning vector subset. The search logic algorithm specifies the reasoning knowledge architecture, integrates the search logic algorithm with the knowledge vector subset, and forms the decision reasoning knowledge base K df . The input signal of this step is the knowledge reasoning vector subset output by step three in step four of the third step; the output signal is the decision reasoning knowledge base K df .

[0117] The specific steps of step five in the third step are as follows:

[0118] Step one, data cleaning, this step cleans the redundant data in the historical data curves of the human situation assessment result H ms , the traffic situation assessment result T fs , the hybrid situation fusion result M us , the online assessment result O er output by step three in step four of the first step, and the historical data curve τ * output by step three in the fourth step, which includes two parts: redundant data cleaning and basic data cleaning.

[0119] Step two, feature extraction, this step extracts the effective features of the historical data curves of the human situation assessment result H ms , the traffic situation assessment result T fs , the hybrid situation fusion result M us , the online assessment result O er and τ * output by step one in step five of the third step, and forms an effective feature set, which includes three parts: "human-traffic" situation feature extraction, human-machine hybrid decision consistency feature extraction, and feature fusion.

[0120] Step three, similarity comparison and rating, the step three respectively compares the effective feature set output by step one in step five of the third step with the effective features of the decision model knowledge corresponding to step two, the decision graph knowledge corresponding to step three and the decision reasoning knowledge corresponding to step four, and calculates the similarity rating of the effective feature set output by step one in step five of the third step, which includes four parts: decision model knowledge similarity comparison, decision reasoning knowledge similarity comparison, decision graph knowledge similarity comparison and similarity rating determination;

[0121] Step four, new knowledge synthesis, the step four first judges the human situation assessment result H ms , traffic situation assessment result T fs , hybrid situation fusion result M us and the effective features of the online assessment result data corresponding to the knowledge type, and then synthesizes new knowledge according to the knowledge data structure of the corresponding knowledge type, which includes four parts: knowledge type classification, decision model knowledge synthesis, decision reasoning knowledge synthesis and decision graph knowledge synthesis;

[0122] To establish a new knowledge synthesis mechanism model, first, the input data of step one in step five of the third step is composed into a new knowledge synthesis data set N ck ={H ms ,T fs ,M us ,O er}, the similar data comparison algorithm is used to clean the configuration redundancy data and physical relationship redundancy data existing in N ck , considering the consistency of each element data on the time axis, the calculation formula of the data comparison number S num and the similarity repetition rate D rt of the data in the comparison window in the similar data comparison algorithm is as follows:

[0123]

[0124] In the formula, Tm is the time stamp of the comparison window, Da is the matrix formed by the comparison data corresponding to Tm, subscript t0 is the first comparison time, f Tm is the sampling frequency of N ck , Δt is the window length, d num is the number of similar repeated records in the window. On this basis, the data corresponding to each element in the cleaned data set N' ck is filtered to obtain the clear data set N' ck ;

[0125] To realize the filtering of N' ckpattern determination, the typical features in N ck need to be extracted, and the data is classified according to similarity, the typical features F ck in N T include "human-traffic" situation features M ST and human-computer hybrid decision consistency features C HT , the "human-traffic" situation features are used to represent the statistical features of the "human-traffic" situation data at the current time, and the human-computer hybrid decision consistency features are used to represent the index features of the human-computer hybrid decision effect at the current time, F T The calculation formula of F

[0126] F T ={M ST ,C HT}={f ST ,e ST ,m ST ,C DK} (19)

[0127] In the formula, f ST , e ST and m ST respectively represent the frequency domain features, extreme value features and mean value features corresponding to M ST , C DK is the consistency rate of human-computer decision, according to F T , N ck is respectively compared with each knowledge element in the decision model knowledge K dd , the decision graph knowledge base K dm and the decision reasoning knowledge base K df in terms of similarity, and the similarity degree D s is evaluated, the similarity comparison is divided into data transformation, semantic parameterization and similarity calculation, the data transformation part transforms N ck into the knowledge data structure in K dd , K dm and K df respectively, and the minimum unit constituting the knowledge in K dd , K dm and K df is defined as a semantic primitive, the parameterization process of the semantic primitive is the process of setting the value ψ of the semantic primitive in a specific knowledge base, when there is data modality in N ck that cannot be transformed into the knowledge data structure in the corresponding knowledge base, the corresponding D s is set to 0; when the data in N ck can be transformed into the knowledge data structure in the corresponding knowledge base, the calculation formula of D s is shown in the following formula (20):

[0128]

[0129] In the formula, d is K represents dd K dm and K df Any knowledge in N' ck The distance of knowledge after data transformation, respectively N' ck Transformed knowledge and K dd K dm and K df Perform similarity calculations based on D s and C DK The similarity level is determined if and only if D s The value is within the D specified in the corresponding knowledge base. s Within the threshold range, while C DK When N' is below the specified threshold, it is considered that N' ck It can synthesize a new piece of knowledge, N' ck The new knowledge is represented according to the knowledge data structure of the corresponding knowledge base and merged into the corresponding knowledge base to complete the synthesis of new knowledge.

[0130] The steps in the fourth step are as follows:

[0131] The specific steps in step two of the fourth step are as follows:

[0132] Step 1: Integration of the knowledge base framework for online evaluation. This step integrates the modeling process of the online evaluation result data in Step 5 of the new knowledge synthesis mechanism model in Step 3, and merges them into the knowledge base framework for online evaluation, which is then output to the decision online verification module in the highly automated driving system.

[0133] Step 2: Online evaluation knowledge base update. This step integrates the online evaluation knowledge output from Step 4 (corresponding to Step 5 in Step 3) according to the knowledge data structure corresponding to the synthesized new knowledge, and then outputs it to the decision online verification module in the highly automated driving system.

[0134] The specific steps in step four of the fourth step are as follows:

[0135] Step 1: Inverse Vehicle Dynamics Solution. This step involves performing inverse dynamics solution on the vehicle state variables at the current moment to obtain the ideal control variable information corresponding to the current moment. The input signal of this step is the vehicle state variables at the current moment; the output signal is the ideal control variable information corresponding to the current moment.

[0136] Section 2, Control Rule Framework: This section establishes a control algorithm framework for calculating the desired control quantity. This control algorithm framework is implemented using typical control algorithms from cybernetics and is used to calculate the desired control quantity.

[0137] Step three, expected control quantity calculation, combines the ideal control quantity information corresponding to the output of step one of step four in the fourth step and the control algorithm framework established by step two to calculate the final expected control quantity, the input signal of this step being the ideal control quantity information under the control rule framework, and the output signal being the expected control quantity.

[0138] The beneficial effects of the present application are as follows:

[0139] 1) The high-consistency man-machine hybrid decision-making method based on hybrid enhanced intelligence provided by the present application can greatly improve the safety and reliability of the man-machine co-driving system and improve the acceptability of the driver by integrating the input data stream and information stream, constructing the man-machine hybrid enhanced reasoning decision-making model, constructing the online man-machine decision-making knowledge base, and integrating the output variables, so as to realize a safe and reliable man-machine hybrid decision-making mode.

[0140] 2) The high-consistency man-machine hybrid decision-making method based on hybrid enhanced intelligence provided by the present application obtains comprehensive and reliable dynamic driving situation data representation based on the "man-traffic" mixed situation evaluation logic, and realizes a comprehensive, reliable, and rich decision-making information source set through the man-machine hybrid enhanced internal model and the corresponding multi-modal data stream of the online evaluation logic.

[0141] 3) The high-consistency man-machine hybrid decision-making method based on hybrid enhanced intelligence provided by the present application eliminates the man-machine decision-making difference under the machine computing attribute based on the reasoning mechanism of the driver and the brain-like computing mode represented by the new generation of artificial intelligence, judges and optimizes the man-machine decision-making result through the man-machine decision-making consistency comparison model, improves the acceptability of the driver, and establishes the man-machine hybrid enhanced reasoning decision-making model to obtain fine and reasonable dynamic driving right allocation results, so as to realize the decision-making effect with high driver acceptability and super-brain mode.

[0142] 4) The high-consistency man-machine hybrid decision-making method based on hybrid enhanced intelligence provided by the present application realizes the construction of the knowledge-level man-machine hybrid decision-making knowledge base by constructing the online man-machine decision-making knowledge base system framework, establishes the knowledge representation rules in the man-machine co-driving system for the hybrid enhanced intelligence theory represented by the new generation of artificial intelligence, establishes comprehensive and reasonable decision-making knowledge content through the decision-making model knowledge, decision-making atlas knowledge, and decision-making reasoning knowledge, realizes the knowledge-level data form of the automatic driving database, improves the system performance of the automatic driving decision-making layer, and greatly improves the adaptability of the system to the real traffic environment through the new knowledge synthesis mechanism model to judge and synthesize the new knowledge collected by the man-machine co-driving system.

[0143] 5) The high-consistency human-machine hybrid decision method based on hybrid enhanced intelligence provided by the application realizes the aggregation and fusion of decision layer signal flow through the integration of output variables, improves the portability of human-machine hybrid decision logic, reduces the complexity of data interaction and model coupling between decision logic and other sub-modules of the automatic driving system, and ensures the independent implementation of the module internal self-checking and self-optimization process of the decision logic.

[0144] 6) The high-consistency human-machine hybrid decision method based on hybrid enhanced intelligence provided by the application has good code consistency, generalization ability and reliability, can realize online decision under normal working conditions and real-time decision performance under emergency working conditions, and after productization, can have a long maintenance and optimization period, and the product has the functions of self-checking and self-optimization in unknown driving situations and real traffic environments. BRIEF DESCRIPTION OF DRAWINGS

[0145] Figure 1 The overall step schematic diagram of the high-consistency human-machine hybrid decision method described in the application.

[0146] Figure 2 The overall architecture schematic diagram of the high-consistency human-machine hybrid decision method described in the application.

[0147] Figure 3 The overall architecture schematic diagram of the first step described in the application.

[0148] Figure 4 The overall architecture schematic diagram of the second step described in the application.

[0149] Figure 5 The overall architecture schematic diagram of the third step described in the application.

[0150] Figure 6 The overall architecture schematic diagram of the fourth step described in the application.

[0151] Figure 7 The algorithm flowchart of step one in the second step described in the application.

[0152] Figure 8 The algorithm flowchart of step two in the second step described in the application.

[0153] Figure 9 The algorithm flowchart of step five in the third step described in the application.

[0154] Figure 10 The calculation result example diagram of step four in the second step described in the application.

[0155] Figure 11 The knowledge base partial example diagram of step three in the third step described in the application. DETAILED DESCRIPTION

[0156] Please see Figures 1 to 11 As shown:

[0157] The highly consistent human-machine hybrid decision-making method based on hybrid augmented intelligence provided by this invention is applied to highly automated driving systems composed of a driver and an automated driving system. The method is as follows:

[0158] The first step is to integrate the input data stream and the information stream;

[0159] The second step is to construct a human-machine hybrid augmented decision-making model.

[0160] The third step is to build an online human-machine decision-making knowledge base.

[0161] Step 4: Integrate the output variables.

[0162] The process of integrating the input data stream and information stream in the first step is as follows:

[0163] Step 1: Integrate multimodal "human-transportation" mixed situational data streams. Step 1 is completed in three stages.

[0164] Step 1: Human Situation Assessment. This step assesses the driver's current situation, including their extraction of the region of interest in the current scene and their driving intentions, based on their current maneuvers and bioelectrical signals. Driver maneuvers include accelerator pedal opening, brake pedal opening causing brake master cylinder pressure, steering wheel angle and angular velocity, eye movements, and head movements. Driver bioelectrical signals include electrocardiogram (ECG), electroencephalogram (EEG), electromyogram (EMG), and electrodermal signaling (EDS). Therefore, the input signals for this step are the driver's maneuver signals and bioelectrical signals; the output signal is the human situation assessment result H. ms .

[0165] Step Two: Traffic Situation Assessment. This step assesses the traffic situation, including dynamic participants in the dynamic scene and driving rules and road conditions in the static scene, using current dynamic and static scene signals. Dynamic scene signals include dynamic vehicle signals and dynamic pedestrian signals, while static scene signals include lane markings, traffic signs, and curb signals. Therefore, the input signals for this step are the dynamic and static scene signals; the output signal is the traffic situation assessment result T. fs .

[0166] Step 3: Hybrid Situation Fusion. This step fuses the output signals from Step 1 (corresponding to Step 1) and Step 2, by analyzing the H... ms and T fsTime alignment and spatial coordinate conversion are performed to realize the hybrid situation fusion with multiple data modalities and scene elements. The input signal of this link is the human situation assessment result H ms output by step one of the first step fs , and the traffic situation assessment result T us output by step two; and the output signal is the hybrid situation fusion result M ms .

[0167] Step two, integrate the model internal parameters of the human-machine hybrid enhanced decision model. In this link, the model internal parameters corresponding to each step of the human-machine hybrid enhanced decision model in the second step are arranged and integrated in time and event order. The integrated model internal parameters are used as the data input for the decision model knowledge construction in step two of the third step. The input signal of this link is the model internal parameters corresponding to each step of the human-machine hybrid enhanced decision model constructed in the second step; and the output signal is the model internal parameter data stream of the human-machine hybrid enhanced decision model.

[0168] Step three, integrate the driving right evaluation target information stream of the human-machine hybrid enhanced decision model. Step three is completed by three links.

[0169] Link one, driving ability assessment of the driver. In this link, the comprehensive control ability of the driver to the vehicle at the current time is calculated based on the state and coupling of the driver-vehicle-road-environment at the current time, i.e., the driving ability of the driver at the current time is assessed. A typical system identification model is used as the driving ability assessment model. The input signal of this link is input into the driving ability assessment model, and then the quantitative driving ability assessment result is obtained. The input signal of this link is the human situation assessment result H ms , vehicle state signal, and vehicle-road coupling state signal; and the output signal is the driving ability assessment result of the driver.

[0170] Link two, driving ability assessment of the self-driving system. In this link, the comprehensive control ability of the self-driving system to the vehicle at the current time is calculated based on the traffic situation and all model internal parameters of the perception and decision levels in the automatic driving system at the current time, i.e., the driving ability of the self-driving system is assessed. The state variables corresponding to the input data of this link are assigned with corresponding weight values, and a linear function with weights is used for calculation, thereby realizing the assessment of the driving ability of the self-driving system. The input signal of this link is the traffic situation assessment result T fs and the internal parameters of the automatic driving system; and the output signal is the driving ability assessment result of the self-driving system.

[0171] Step three, driving right planning. Through the driving ability evaluation result of the driver output by step three of the first step and the driving ability evaluation result of the self-driving system output by step two, the comprehensive control effect of the driver and the self-driving system on the vehicle at the current time can be quantitatively evaluated. Therefore, according to the driving ability evaluation result of the driver and the driving ability evaluation result of the self-driving system, the driving right distribution coefficient τ between the driver and the self-driving system at the current time can be calculated by normalization calculation. The input signal of this step is the driving ability evaluation result of the driver and the driving ability evaluation result of the self-driving system, and the output signal is the driving right distribution coefficient τ.

[0172] Step four, integrate human-machine hybrid enhanced decision online evaluation result information flow. Step four is completed by three steps.

[0173] Step one, short-term online verification result. This step obtains the short-term online verification result by integrating the short-term rapid updating decision online verification module in the highly automated driving system. The decision online verification module is a module in the highly automated driving system for evaluating the effect of human-machine hybrid decision. By time alignment and threshold detection of the short-term verification signal output by the module, a clear and reasonable short-term online verification result can be obtained. The input signal of this step is the short-term verification module signal; the output signal is the short-term verification result.

[0174] Step two, long-term online optimization result. This step obtains the long-term online verification result by integrating the long-term rapid updating decision online verification module in the highly automated driving system. By time alignment and threshold detection of the long-term verification signal output by the module, a clear and reasonable long-term online verification result can be obtained. The input signal of this step is the long-term verification module signal; the output signal is the long-term verification result.

[0175] Step three, online evaluation information flow integration. This step integrates the short-term verification result output by step one and the long-term verification result output by step two of step four in the first step to obtain the online evaluation result. The integration process mainly includes time domain category labeling and redundant time domain elimination. The input signal of this step is the short-term verification result and the long-term verification result, and the output signal is the online evaluation result O er .

[0176] In Figure 3 an exemplary embodiment of the first step is shown. The first step finally outputs the multi-modal "human-traffic" mixed situation data flow, the human-machine hybrid enhanced decision internal parameter data flow, the human-machine hybrid enhanced decision driving right evaluation target information flow, and the human-machine hybrid enhanced decision online evaluation result information flow.

[0177] The process of building the human-machine hybrid enhanced decision model in the second step is as follows:

[0178] Step one, constructing the driver reasoning mechanism model. Step one is completed by four links.

[0179] Link one, multi-objective learning framework. Establish multi-learning objectives based on safety, comfort, functionality and maneuverability of human-machine hybrid decision-making, extract local sub-graphs in the global graph of the driver reasoning mechanism model, and establish a learning framework for random walk based on relationship clustering and coupling results between local sub-graphs. Specifically, it includes three parts: reasoning mechanism model definition, entity relationship standard formulation, and entity relationship clustering.

[0180] Link two, global graph random walk. Establish the global graph of the driver reasoning mechanism, and judge the reachability between each entity pair in the global graph, and then solve the reasoning result corresponding to the global graph. Specifically, it includes three parts: global graph establishment, entity reachability calculation, and global reasoning result calculation.

[0181] Link three, local sub-graph random walk. Extract specific local relationship sub-graphs of the driver reasoning mechanism from the global graph of the driver reasoning mechanism, and realize random walk. Specifically, it includes three parts: local sub-graph establishment, entity transition probability matrix calculation, and local reasoning result calculation.

[0182] Link four, fusion reasoning. The global reasoning result and the local reasoning result obtained by link two and link three in step one of the second step are uniformly matched in the distribution area, and the reasoning result is fused by using nonlinear mapping logic. Specifically, it includes two parts: reasoning result normalization calculation and fusion reasoning result calculation.

[0183] In Figure 7 , an exemplary embodiment of step one of the second step is shown. The input signal corresponding to link one of step one of the second step is the human situation assessment result H ms corresponding to link one of step one of the first step in the second step hm , the human-machine decision consistency comparison result C dm corresponding to step three of the second step df , the decision graph knowledge base K ms corresponding to step three of the third step hm , and the decision reasoning knowledge base K lf corresponding to step four of the third step; the output signal is the entity relationship clustering result. The input signal corresponding to link two of step one of the second step is the entity relationship clustering result; the output signal is the global reasoning result. The input signal corresponding to link three of step one of the second step is the global graph; the output signal is the local reasoning result. The input signal corresponding to link four of step one of the second step is the global reasoning result and the local reasoning result; the output signal is the fusion reasoning result.

[0184] H ms and C hm compose the driver reasoning online data stream F lf , that is, Flf = {H ms , hm} ; K df and K dm comprise the driver inference offline knowledge flow F fk , i.e. F fk = {K df , K dm}. Therefore, the calculation formula of the driver inference mechanism model (DIMM) is shown in the following (1):

[0185]

[0186] In the formula, G m and L m represent the global graph and the local sub-graph respectively, f(G m , L m ) is the function of DIMM, and f(F lf ) represents the subset function of F fk . Therefore, the DIMM model is an inference model based on the multi-objective learning framework and the random walk pattern. In a specific driving scene at a certain time, the multiple iterations and clustering of clusters |R γ | are realized through the comparison of entity correlation between independent relationships γ, and the shared feature value set C γ containing all feature values c γ of the new cluster formed by updating is updated. The calculation formula of the inter-cluster similarity function sim(C γ,m , C γ,n ) is shown in the following formula (2):

[0187] (2)

[0188] In the formula, the operator Π represents the multiplication of each element in the set C γ , and the subscripts m and n of C γ represent two different C γ with numbers m and n. On the basis of the similarity between |R γ |, a joint learning classification model is established to couple and form the path R r of γ in each |R γ |. The calculation formula of the classifier structure function f cl (R r ) and the corresponding joint relationship learning model is shown in the following formula (3):

[0189]

[0190] In the formula, μ1 and μ2 are the regularization coefficients, and ω kω0 and ω0 are the weighting coefficients and their benchmark values, respectively, and b k b0 and b0 are the classification structure deviation coefficient and its benchmark value, respectively, and d k This represents the weight vector bias coefficient. The function L(R) ri,p ,R ri,q ) is f cl (R r The training loss function of R, where the subscripts p and q represent two distinct R values ​​numbered p and q, respectively. r The subscript i indicates R. r The number, N k For R r The quantity, the subscript k represents d k The number, K is d k The number of implementation clusters |R obtained by clustering entity relationships. γ |and its corresponding path R r As a constraint, the calculation of the global graph random walk in step one of the second step is implemented. The global graph random walk in step one is achieved by extracting G... m The various relations r el and its corresponding c γ Establish a global relational feature model. m It can be defined as G m ={g m,i ={h gm,i ,R rg,i ,ra gm,i}, i=1,2,...,s}, where g m,i For G m A subgraph, where the subscripts i and s represent the subgraph number and the total number of subgraphs, respectively, g m,i h gm,i R rg,i and ra gm,i These represent the head entity, path, and tail entity of a valid subgraph, respectively. G m Zhongyouh gm,i Through R rg,i Arrive at ra gm,i Accessibility p re The calculation formula is shown in equation (4) below:

[0191]

[0192] In the formula, sl={h gm,i ∪ra gm,i ,i=1,2,...,s},R ra,i To be with ra gm,i The directly corresponding R rg,i The tail relation elements in the text, sl and ra gm,i The similarity function sim(sl,ra)gm,i The calculation method of ) and the calculation formula (2) for C γ,m and C γ,n The similarity function between them is calculated in the same way. Let α represent R. rg,i The corresponding weight matrix allows the global graph random walk model to be expressed as: f(G m )=α·p re The logistic regression algorithm is used to analyze the model f(G). m The parameters are trained using the given parameters, and the sigmoid function is chosen as the normalization function for the results. The normalized global inference result p g The calculation formula is shown in equation (5) below:

[0193] (5)

[0194] L m For G m A subset of can be defined as L m ={l m,i ={h lm,i ,R rl,i ,ra lm,i}, i = 1, 2, ..., z}, where l m,i For L m A subgraph, where the subscripts i and z represent the subgraph number and the total number of subgraphs, respectively, l m,i hl in m,i R rl,i and ra lm,i These represent the head entity, path, and tail entity of the valid subgraph, respectively. Therefore, for L... m When performing random walk computation, the space complexity is reduced, so direct computation of L can be used. m The transition probability matrix T between different entities M This leads to the corresponding local reasoning results. M The calculation formula is shown in equation (6) below:

[0195]

[0196] In the formula, N hl,i and N ral,i According to hl m,i and ra lm,i The constructed diagonal matrix, sp is T M The number of transition steps, M l,i For the sp-th step L m The corresponding adjacency matrix. T M The element T corresponding to row a and column b M [a,b] indicates that hl m,iStarting from point A, perform a random walk, and after SP steps, jump to point A. lm,i The probability of p. l Indicates R rl,i The evaluation results of the local inference results yielded p l The calculation formula is shown in equation (7) below:

[0197]

[0198] p g and p l After fusion and normalization, the fusion inference result p was obtained. f The calculation formula is shown in equation (8) below:

[0199] (8)

[0200] In the formula, δ represents the fusion inference stability coefficient, used to balance p g and p l The contribution ratio. Therefore, the driver reasoning mechanism model in step one of the second step outputs the fusion reasoning result.

[0201] Step 2: Construct an advanced "self-like" decision-making model based on brain-like computing. Step 2 consists of four steps.

[0202] Step 1: Neuron Group Model. Establishing a neuron group model provides a personalized classification basis for the "self-like" decision-making model corresponding to Step 4 of Step 2. Specifically, it includes four parts: neuron group model definition, feature extraction, stimulus-based classification, and salience matrix.

[0203] Step Two: Deep Convolutional Networks. Deep learning networks provide the policy-oriented fitting basis for the reinforcement learning model corresponding to Step Three in Step Two. Specifically, this includes four parts: defining the behavior network structure and defining the evaluation network structure.

[0204] Step 3: Reinforcement Learning Model. This step utilizes a reinforcement learning model to realize the complex decision-making patterns of an autonomous driving system. It involves finding the optimal policy and calculating the optimal action within the reinforcement learning model. Specifically, this includes three parts: state definition, reward function, and policy gradient.

[0205] Step Four: The "Similarity" Decision-Making Model. This step utilizes the personalized classification criteria obtained in Step Two, Step One of Step Two, and combines them with the deep reinforcement learning processes in Step Two, Step Two, and Step Three of Step Two. At the data level, it integrates "Similarity" decision-making data. By judging whether the iterative effect reaches a preset threshold, it ultimately outputs the online decision-making result of the autonomous driving system. Specifically, it includes two parts: "Similarity" data fusion and threshold judgment.

[0206] existFigure 8 The diagram illustrates an exemplary implementation of step two in the second step. The input signal corresponding to the first stage of step two in the second step is the traffic situation assessment result T corresponding to the first stage of step one in the first step. fs C corresponding to step three in the second step hm In the third step, step two of the decision model knowledge K dd And the decision graph knowledge base K corresponding to step three in the third step. dm The output signal is the personalized classification model M. c In the second step, the input signal corresponding to the second stage of step two is the same as the input signal corresponding to the first stage of step one in the first step. fs C corresponding to step three in the second step hm After M c K corresponding to the classification dd and K dm The output signal is the expected discounted return Q. π And the policy gradient optimization function ▽ θ In the second step, the input signal corresponding to stage three in step two is the same as the input signal corresponding to stage one in step one in the first step. fs C corresponding to step three in the second step hm Q π And ▽ θ The output signal is the decision strategy π(a) RDL ,s RDL In the second step, the input signal corresponding to stage four in step two is π(a). RDL ,s RDL ) and M c The output signal is the optimal decision strategy π. * (a RDL ,s RDL ).

[0207] T fs and C hm The online classification data stream consisting of neuron groups F cf That is, F cf ={T fs C hm};K dd and K dm Composing the knowledge flow F for offline training of autonomous driving ak That is, F ak ={K dd ,K dm Therefore, the computational formula for the Neuron Group Model (NGM) is shown in equation (9) below:

[0208] NGM=f(F cf (9)

[0209] According to F ak Specific features are extracted and learned using incremental learning rules based on Haibo learning and Heb learning. A preset N is used. G The category labels of the dimension are determined, and the extracted specific features are used as conditional stimuli to build a personalized neuron group model belonging to a specific classification dimension. The incremental learning rule calculation formula is shown in the following formula (10):

[0210]

[0211] In the formula, ΔCHL, ΔHeb, and Δβ represent the Heb learning rule, the Heb learning rule, and the incremental learning rule, respectively, β represents the synaptic matrix corresponding to the incremental learning rule, and g j For presynaptic activation, h i For postsynaptic activation, the subscripts i and j correspond to F ak The extracted specific feature element in the i-th row and j-th column, where + and - represent the home and de-home phases respectively, ζ is the weight coefficient, and κ is the learning rate. The kWTA function is used to obtain the sparse distribution representation in the feature matrix, and the first r activation units are extracted, along with their corresponding inhibition functions f. r The calculation formula is shown in equation (11) below:

[0212]

[0213] In the formula, χ is the suppression threshold, ensuring that the first r activation units are in active mode. Therefore, the activated h i This can be achieved by calculating the i-th row of β and g. j The standardized dot product is obtained, and the corresponding calculation formula is shown in equation (12) below:

[0214]

[0215] Traverse and calculate h i Then, h i The element h with the largest value in the middle i,max Defined as a maximally reactive neuron if and only if r = h i,max When, the corresponding β ij These are the feature elements that have already been learned. In summary, N is established. G Dimension M c ={m c,1 ,m c,2 ,...,m c,NG}, where m c For N G M after dimensional classification c Sub-models.

[0216] Will pass through M c K under the corresponding sub-model classification ddand K dm sub-data set K dd,NG and K dm,NG , as the second step in step two of the second part of the deep convolutional network and reinforcement learning model corresponding to the model training data. Deep convolutional network and reinforcement learning model together constitute a deep reinforcement learning model (DRLM), DRLM model definition corresponding to the calculation formula as shown in equation (13):

[0217] DRLM = {S DR , A DR , P DR , R DR , λ} (13)

[0218] In the formula, S DR is the vehicle state space of the autonomous driving system, A DR is the control action space of autonomous driving, P DR is the state transition probability distribution of the autonomous driving system, R DR is the reward function, and λ = [λ1, λ2, λ3, λ4] is the discount factor. DRIM R DR The calculation formula is shown in equation (14):

[0219]

[0220] In the formula, R safety , R goal , R law and R comft respectively represent the safety reward function, the time reward function, the traffic regulation reward function and the comfort reward function, corresponding to the safety, maneuverability, traffic rules and driver comfort of the autonomous driving system. Four aspects of driving goals in driving tasks. To achieve parameterized estimation of continuous data sequence in autonomous driving decision-making, that is, to accurately estimate P DR when S DR is a continuous data sequence pattern, a deep convolutional network based on policy gradient is used to search π(a RDL , s RDL ). The deep convolutional network mainly includes behavior network L(θ Q ) and evaluation network ▽ θμ L(θ Q ), which can respectively obtain the definition calculation formula of the corresponding network as shown in equations (15) and (16):

[0221]

[0222]

[0223] In the formula, operator E is the expectation function, subscripts t and t+1 represent the current time step and the next time step, respectively, and θ Q and θ μ These represent two nonlinear estimators for different neural network structures. Q is the action-value function, representing the action value in state S. DR Take action A DR The expected discount return obtained; μ is the value that makes a RDL =μ(s) RDL ;θ μ A mapping function constructed to satisfy the condition ▽ θμ L(θ Q The gradient update method is used to update π(a) RDL ,s RDL ) optimization.

[0224] The calculated π(a) RDL ,s RDL Data fusion and updating are performed in the memory pool to ultimately obtain π. * (a RDL ,s RDL The calculation formula for the update mode is shown in equation (17) below:

[0225]

[0226] In the formula, . a normal number.

[0227] Therefore, step two in the second step, based on the advanced "self-like" decision model of brain-like computing, outputs the "self-like" decision results of the autonomous driving system.

[0228] Step 3: Construct a human-machine decision consistency comparison model. Step 3 consists of three parts.

[0229] Step 1: Driver Decision Map. The input signal for this step is the fusion inference result p output from Step 4 of Step 1 in Step 2. f And the human situation assessment result H corresponding to the first step of step one in the first step. ms And according to the decision graph knowledge base K corresponding to step three in step three. dm The knowledge representation rules are ultimately expressed in the form of a classification tree graph structure, which represents p. f H ms And the driver's control signals are integrated into a driver decision map.

[0230] Step Two: Autonomous Driving Decision Map. The input signal for this step is the optimal decision strategy π output from Step Two, which corresponds to Step Three in Step Two. * (a RDL ,s RDL) and the traffic situation assessment result T corresponding to step one of the first step fs , and according to the knowledge expression rules of the decision graph knowledge base K corresponding to step three of the third step dm , finally, the integration of p * (a RDL , s RDL ), T fs and the control signal of the autonomous driving system into the self-driving system decision graph is in the form of a classification tree graph structure.

[0231] Step three, human-machine graph comparison and prediction. The input signals of this step are the mixed situation fusion result of step one of the first step, the driver decision graph output of step three of the second step and the self-driving system decision graph output of step two; the output signal is the human-machine decision consistency rate C DK . First, the spatiotemporal evolution law of the vehicle state and trajectory in the short time domain of 0-10 seconds from the current time is predicted under the condition of only driver driving through the driver decision graph; at the same time, the spatiotemporal evolution law of the vehicle state and trajectory in the short time domain of 0-10 seconds from the current time is predicted under the condition of only autonomous driving system driving through the self-driving system decision graph. Next, the similarity of the vehicle state and trajectory under the conditions of only driver driving and only autonomous driving system driving in the same short time domain is solved, and the similarity result is taken as the output signal C DK of this step.

[0232] Step four, building a driving right subdivision model. Step four is completed by three steps.

[0233] Step one, subdivision criterion library. This step stores the subdivision rules of the human-machine driving right corresponding to the human-machine mixed decision layer of the highly automated driving system in an offline storage manner. The subdivision criterion library mainly includes: traffic rule library, maneuverability criterion library, safety criterion library and comfort criterion library. The traffic rule library stores the set of traffic rules corresponding to urban traffic; the maneuverability criterion library stores the set of criteria constructed to ensure vehicle driving efficiency; the safety criterion library stores the set of criteria to ensure the safety of longitudinal and lateral driving performance of the vehicle in emergency conditions; the comfort criterion library stores the set of criteria to ensure that the driver and passengers of the vehicle are in a comfortable state during vehicle driving.

[0234] Step two, constraint rules. This step stores the constraint rules for constraining the configuration and value range of the knowledge and model parameters of the human-machine mixed decision layer in an offline storage manner. Mainly including: constraint rules for the value range of the decision model knowledge of step two of the third step, threshold constraint rules for the vehicle state when only the driver drives, threshold constraint rules for the vehicle state when only the autonomous driving system drives.

[0235] Step 4, the third link of step 4 in the second step, corresponds to the input signal of the third link of step 3 in the first step, the driving right allocation coefficient τ output by the third link of step 3 in the first step, the subdivision criterion output by the first link of step 4 in the second step, the constraint rule output by the second link of step 4 in the second step, the driving signal of the driver at the current time and the driving signal of the automatic driving system at the current time; the output signal is the driving right weight value of the driver and the driving right weight value of the automatic driving system. This link first establishes a two-dimensional linear programming plane of τ output by the third link of step 3 in the first step. Then, through the subdivision criterion output by the first link of step 4 in the second step and the constraint rule output by the second link of step 4 in the second step, the driving signal of the driver and the driving signal of the automatic driving system at the current time are standardized and constrained, and finally the optimized driving right allocation coefficient τ is obtained. DK DK * .

[0236] An exemplary operation result of step 4 in the second step is shown in Figure 10 .

[0237] An exemplary embodiment of the second step is shown in Figure 4 . The second step finally outputs the driving right subdivision result, including the driving right weight value of the driver and the driving right weight value of the automatic driving system.

[0238] The process of constructing the online human-machine decision knowledge base in the third step is as follows:

[0239] Step 1, establish the system framework of the online human-machine decision knowledge base. The specific content includes: the knowledge mode for human-machine decision, the knowledge data structure between modes, and the data interaction logic between modes. The knowledge mode for human-machine decision includes the decision model knowledge in step 2, the decision graph knowledge in step 3, the decision reasoning knowledge in step 4, and the new knowledge synthesis mechanism model in step 5. The knowledge data structure of the decision model knowledge in step 2 is an object-oriented semantic mapping structure; the knowledge data structure of the decision graph knowledge in step 3 is a classification tree graph structure; the knowledge data structure of the decision reasoning knowledge in step 4 is a graph structure based on data sequence; the new knowledge synthesis mechanism model in step 5 judges the new knowledge mode and adopts the corresponding knowledge data structure. The decision graph knowledge in step 3 and the decision reasoning knowledge in step 4 receive the knowledge content of the decision model knowledge in step 2, and take it as the input of the decision graph knowledge in step 3 and the decision reasoning knowledge in step 4; the new knowledge synthesis mechanism model in step 5 receives the knowledge content of the decision model knowledge in step 2, the decision graph knowledge in step 3 and the decision reasoning knowledge in step 4, and takes it as the basis for judging the new knowledge mode, and outputs the new knowledge to the corresponding step in the same new knowledge mode.​​

[0240] Step two, establish decision model knowledge. Step two is completed by four links.

[0241] Link one, knowledge element ontology modeling. This link standardizes the model structure and definition of the decision model knowledge in step two in the third step, and determines the knowledge source of the decision model knowledge. Then the meta ontology model of the decision model knowledge is established by establishing process model and concept model respectively. The input signal of this link is the human-machine mixed enhanced decision model internal parameter data stream corresponding to step two in the first step, and the output signal is the meta ontology model.

[0242] Link two, decision model ontology design. This link adopts the meta ontology development mode from the bottom up, takes the main class of the meta ontology as the core, defines three top-level classes such as top-level model class, top-level data class and logic expression class, and realizes the design of the decision model ontology. The input signal of this link is the human-machine mixed enhanced decision model internal parameter data stream corresponding to step two in the first step, and the output signal is the primary decision model knowledge base with the top-level classes of the meta ontology.

[0243] Link three, model knowledge rule design. This link reasons and queries the concept model of the meta ontology, and constructs the semantic rule base through two ways of knowledge path management rule and knowledge action operation rule. The input signal of this link is the primary decision model knowledge base output by link two corresponding to step two in the third step, and the output signal is the updated decision model knowledge base with semantic rules.

[0244] Link four, model knowledge architecture design. This link establishes a semantic web framework as a knowledge engine to integrate dynamic knowledge meta ontology and static knowledge meta ontology. In addition, this link establishes a semantic middleware to process the multi-modal human-machine mixed enhanced decision model internal parameter data stream into a unified meta ontology structure in a semantic mapping manner. The input signal of this link is the human-machine mixed enhanced decision model internal parameter data stream corresponding to step two in the first step and the updated decision model knowledge base output by link three corresponding to step two, and the output signal is the decision model knowledge K dd .

[0245] Step three, establish decision graph knowledge. Step three is completed by four links.

[0246] Link one, knowledge graph ontology construction. This link establishes a knowledge graph ontology for decision graph knowledge, and synthesizes an effective knowledge ontology for human-machine mixed decision by adopting a domain ontology mode. The input signal of this link is the human situation assessment result H ms output by link one corresponding to step one in the first step, the traffic situation assessment result T fsand the new mode decision reasoning knowledge outputted in step five of the third step, the output signal is the effective knowledge ontology set. Considering H ms and T fs In the time-continuous data sequence, the effective knowledge ontology synthesized by the domain ontology mode is divided into five elements, i.e. concept, relation, function, axiom and individual.

[0247] Link two, atlas semantic rule making. This link specifies semantic rules to standardize the effective knowledge ontology at the semantic level. The basic semantic rules are established by combining abstract syntax and concrete syntax, and four reasoning functions, i.e. consistency checking, classification, identification and prediction, are added to the basic semantic rules to realize semantic rules with reasoning mechanism. This link has no input signal, and the output signal is semantic rules.

[0248] Link three, atlas semantic mapping. This link maps the effective knowledge ontology set and semantic rules in the form of a relation table to form a decision atlas case. Different effective knowledge ontologies are identified by unique target addresses, and the reasonable mapping of multiple ontology relations is realized by mapping the relation classes composed of effective knowledge ontologies and their corresponding semantics, finally forming a decision atlas case library composed of multiple relation classes, semantic rules and their mutual mapping relations. The input signal of this link is the effective knowledge ontology set outputted by the corresponding link one in step three of the third step and the semantic rules outputted by link two, and the output signal is the decision atlas case library.

[0249] Link four, case knowledge calling. This link formulates a case access mechanism to search for effective knowledge ontologies in the decision atlas case library first, then search for similar relations, and finally search for similar semantics to realize the calling of specific cases in the decision atlas case library. The case access mechanism established by this link and the decision atlas case library together form a decision atlas knowledge base. The input signal of this link is the decision atlas case library outputted by the corresponding link three in step three of the third step, and the output signal is the decision atlas knowledge base K dm .

[0250] An exemplary operation result of step three in the third step is shown in Table Figure 11

[0251] Step four, establishing decision reasoning knowledge. Step four is completed by four links.

[0252] Link one, knowledge rule oriented model. This link establishes a knowledge rule oriented model to convert the pre-collected driver offline database into a driver offline knowledge base with knowledge rules as constraints. The input signal of this link is the driver offline database information and the new mode decision reasoning knowledge outputted in step five of the third step, and the output signal is the knowledge rule oriented function and the driver offline knowledge base. ​

[0253] Step two, knowledge reasoning vector modeling. This step establishes a knowledge reasoning vector based on the vehicle dynamics and the vehicle-road coupling characteristics, solves the reasoning and prediction results of the vehicle state changes corresponding to the driver's offline knowledge base in the form of knowledge data structure, and integrates them into a knowledge reasoning vector set. The input signal of this step is the knowledge rule oriented function output by step one in step four of the third step, and the driver's offline knowledge base; the output signal is the knowledge reasoning vector set.

[0254] Step three, hierarchical logic reasoning classification. This step establishes a hierarchical logic reasoning classification method based on the danger level and driving mode of the scene, and divides the knowledge reasoning vector set into a typical knowledge reasoning vector subset under this classification method. The input signal of this step is the knowledge reasoning vector set output by step two in step four of the third step; the output signal is the classified knowledge reasoning vector subset.

[0255] Step four, reasoning knowledge architecture generation. This step arranges the knowledge reasoning vector subset according to the priority, and sets the search logic algorithm for the knowledge reasoning vector subset, which specifies the reasoning knowledge architecture. The search logic algorithm is integrated with the knowledge vector subset to form a decision reasoning knowledge base K df . The input signal of this step is the knowledge reasoning vector subset output by step three in step four of the third step; the output signal is the decision reasoning knowledge base K df .

[0256] Step five, build a new knowledge synthesis mechanism model. Step five is completed by four steps.

[0257] Step one, data cleaning. This step cleans the redundant data in the historical data curves of the human situation assessment result H ms , the traffic situation assessment result T fs , the hybrid situation fusion result M us , the online assessment result O er output by step three in step four of the first step, and the output τ * of step three in the fourth step. Specifically, it includes two parts: redundant data cleaning and basic data cleaning.

[0258] Step two, feature extraction. This step extracts the cleaned human situation assessment result H ms , the traffic situation assessment result T fs , the hybrid situation fusion result M us , the online assessment result O er , and τ *The effective features of historical data curves are extracted and a set of effective features is formed. Specifically, it includes three parts: extraction of "human-traffic" situational features, extraction of human-machine hybrid decision-making consistency features, and feature fusion.

[0259] Step 3: Similarity Comparison and Rating. This step compares the effective feature set output from Step 1 (corresponding to Step 5 in Step 3) with the effective features of the decision model knowledge (corresponding to Step 2 in Step 3), the decision graph knowledge (corresponding to Step 3), and the decision reasoning knowledge (corresponding to Step 4 in Step 3), and calculates the similarity level corresponding to the effective feature set output from Step 1 (corresponding to Step 5 in Step 3). Specifically, it includes four parts: similarity comparison of decision model knowledge, similarity comparison of decision reasoning knowledge, similarity comparison of decision graph knowledge, and similarity level determination.

[0260] Step 4: New Knowledge Synthesis. This step first determines the human posture assessment result H corresponding to the valid feature set with a similarity level below a threshold. ms Traffic situation assessment results T fs Hybrid situation fusion result M us The system also considers the knowledge types corresponding to the effective features of the online evaluation results data, and then synthesizes new knowledge based on the knowledge data structure of the corresponding knowledge types. Specifically, this includes four parts: knowledge type classification, decision model knowledge synthesis, decision reasoning knowledge synthesis, and decision graph knowledge synthesis.

[0261] exist Figure 9 The document illustrates an exemplary implementation of step five in step three. To establish a new knowledge synthesis mechanism model, the input data from stage one corresponding to step five in step three are first combined to form a new knowledge synthesis dataset N. ck ={H ms ,T fs M us O er}, using a similar data comparison algorithm, N ck Redundant data in terms of configuration and physical relationships is cleaned. Considering the consistency of each element's data over time, the number of data entries S compared in the similar data comparison algorithm is [not specified]. num And the similarity repetition rate D of the data within the comparison window rt The calculation formula is shown in equation (18):

[0262]

[0263] In the formula, Tm is the timestamp of the comparison window, Da is the matrix formed by the comparison data corresponding to Tm, the subscript t0 is the first comparison time, and f Tm For N ck The sampling frequency, Δt is the window length, and d numThe number of similar repeated records in the window. On this basis, the data corresponding to each element in the cleaned data set is filtered to obtain the clear data set N' ck .

[0264] To realize the mode determination of N' ck , the typical features in N' ck need to be extracted, and the data is classified according to similarity. The typical features F ck in N' T include the "person-traffic" situation feature M ST and the consistency feature C HT of human-machine mixed decision. The "person-traffic" situation feature is used to represent the statistical characteristics of the "person-traffic" situation data at the current time, and the consistency feature of human-machine mixed decision is used to represent the index characteristics of the effect of human-machine mixed decision at the current time. Therefore, the calculation formula of F T is shown in the following formula (19):

[0265] F T ={M ST ,C HT}={f ST ,e ST ,m ST ,C DK} (19)

[0266] In the formula, f ST , e ST and m ST represent the frequency domain feature, extreme value feature and mean value feature corresponding to M ST , respectively, and C DK is the consistency rate of human-machine decision. According to F T , N' ck is compared with each knowledge element in the decision model knowledge K dd , the decision graph knowledge base K dm and the decision reasoning knowledge base K df for similarity, and the similarity degree D s is evaluated. The similarity comparison includes data transformation, semantic parameterization and similarity calculation. The data transformation part transforms N' ck into the knowledge data structure in K dd , K dm and K df , respectively. The smallest unit constituting the knowledge in K dd , K dm and K df is defined as a semantic primitive, and the parameterization process of the semantic primitive is the process of setting the value ψ of the semantic primitive in a specific knowledge base. For N' ckWhen there is no data modality in K s is set to 0; N ck When there is data modality in K s The calculation formula of D is shown in the following formula (20):

[0267]

[0268] In the formula, d is represents the distance between any knowledge in K dd , K dm and K df and N ck data converted knowledge. The similarity calculation is respectively performed on the converted knowledge N ck and K dd , K dm and K df . The similarity level is determined according to D s and C DK . The determination logic is that when and only when D s is within the D s threshold value range specified by the corresponding knowledge base, and C DK is lower than the specified threshold value, it is considered that N ck can synthesize a new knowledge. N ck is represented according to the knowledge data structure of the corresponding knowledge base, and is merged into the corresponding knowledge base to complete the synthesis of new knowledge.

[0269] An exemplary embodiment of the third step is shown in Figure 5 .

[0270] The process of integrating the output variables in the fourth step is as follows:

[0271] Step one, integrate the high-level “class me” decision process variables. This step integrates the output signals corresponding to each link included in steps one, two and three in the second step in a time-aligned manner, which is used for output to the decision online verification module in the highly automated driving system.

[0272] Step two, integrate the knowledge base knowledge for online evaluation. Step two is completed by two links.

[0273] Link one, integrate the knowledge base framework for online evaluation. This link integrates the modeling process of online evaluation result data in the new knowledge synthesis mechanism model in step five in the third step, and merges it into the knowledge base framework for online evaluation, and then outputs to the decision online verification module in the highly automated driving system.

[0274] Step two, online evaluation knowledge base updating. This step integrates the online evaluation knowledge outputted by step five in the third step according to the knowledge data structure corresponding to the new knowledge after synthesis, and then outputs to the decision online verification module in the highly automated driving system.

[0275] Step three, integrating human-machine hybrid decision driving right sub-weight. This step integrates the driving right distribution coefficient τ * outputted by step four in the second step according to the historical data storage of τ * for output to the step five new knowledge synthesis mechanism model in the third step.

[0276] Step four, integrating human-machine hybrid decision expected control amount. Step four is completed by three steps.

[0277] Step one, vehicle dynamics inverse solution. This step solves the ideal control amount information corresponding to the current time by performing dynamics inverse solution on the vehicle state amount corresponding to the current time. The input signal of this step is the vehicle state amount corresponding to the current time; the output signal is the ideal control amount information corresponding to the current time.

[0278] Step two, control rule framework. This step establishes a control algorithm framework for calculating the expected control amount. The control algorithm framework is realized by using a typical control algorithm in control theory, which is used to calculate the expected control amount.

[0279] Step three, expected control amount calculation. This step calculates the final expected control amount by combining the ideal control amount information corresponding to the current time outputted by step one in step four in the fourth step and the control algorithm framework established by step two. The input signal of this step is the ideal control amount information under the control rule framework; the output signal is the expected control amount.

[0280] An exemplary embodiment of the fourth step is shown in Figure 6 .

Claims

1. A hybrid augmented intelligence based highly consistent human-machine hybrid decision making method, characterized in that: The method comprises the following steps: The first step is to integrate the input data stream and the information stream, and the specific process is as follows: Step one, integrate multi-modal "human-traffic" mixed situation data stream; Step two, integrate human-computer mixed enhanced decision model internal parameter data stream, which arranges and integrates the model internal parameters corresponding to each step of the human-computer mixed enhanced decision model in step two according to time and event order, and uses the integrated model internal parameters as the data input of the decision model knowledge construction in step two in step three; the input signal of this step is the model internal parameters corresponding to each step of the human-computer mixed enhanced decision model constructed in step two; the output signal is the human-computer mixed enhanced decision model internal parameter data stream; Step three, integrate human-computer mixed enhanced decision driving right evaluation target information stream; Step four, integrate human-computer mixed enhanced decision online evaluation result information stream; The second step is to build a human-computer mixed enhanced decision model, and the specific process is as follows: Step one, build a driver reasoning mechanism model; Step two, build a high-level "similar to me" decision model based on brain-like computing; Step three, build a human-computer decision consistency comparison model; Step four, build a driving right subdivision model; The third step is to build an online human-computer decision knowledge base, and the specific process is as follows: Step one, establish the system framework of the online human-computer decision knowledge base, which includes the knowledge mode for human-computer decision, the knowledge data structure between modes, and the data interaction logic between modes. The knowledge mode for human-computer decision includes the decision model knowledge in step two, the decision graph knowledge in step three, the decision reasoning knowledge in step four, and the new knowledge synthesis mechanism model in step five. The knowledge data structure of the decision model knowledge in step two is an object-oriented semantic mapping structure. The knowledge data structure of the decision graph knowledge in step three is a classification tree graph structure. The knowledge data structure of the decision reasoning knowledge in step four is a graph structure based on data sequence. The new knowledge synthesis mechanism model in step five judges the new knowledge mode and adopts the corresponding knowledge data structure. The decision graph knowledge in step three and the decision reasoning knowledge in step four receive the knowledge content of the decision model knowledge in step two and use it as the input of the decision graph knowledge in step three and the decision reasoning knowledge in step four. The new knowledge synthesis mechanism model in step five receives the knowledge content of the decision model knowledge in step two, the decision graph knowledge in step three, and the decision reasoning knowledge in step four, and uses it as the basis for judging the new knowledge mode. The new knowledge is output to the corresponding step of the new knowledge mode; Step two, establish the decision model knowledge; Step three, establish the decision graph knowledge; Step four, establish the decision reasoning knowledge; Step five, build a new knowledge synthesis mechanism model; The fourth step is to integrate the output variables, and the specific process is as follows: Step one, integrate the high-level "similar to me" decision process quantity, which integrates the output signals corresponding to each link in steps one, two and three in step two in a time alignment manner, and is used as the input of the decision online verification module in the highly automatic driving system; Step two, integrate the knowledge base knowledge quantity for online evaluation; Step three, integrating the human-machine mixed decision driving right sub-weight, this step will output the driving right distribution coefficient τ corresponding to step four of the third step in the second step * The integration includes the historical data storage of τ * for output to the new knowledge synthesis mechanism model in step five in the third step; Step four, integrate the human-computer mixed decision expected control quantity.

2. The method of claim 1, wherein: The links included in each step of the first step are as follows: The specific steps of step one in the first step are as follows: The first link is the human situation assessment. The human situation assessment includes the extraction of the region of interest in the current scene and the driving intention of the driver. The driving action of the driver includes the accelerator pedal opening, the brake pedal opening, the brake master cylinder pressure, the steering wheel angle and angular velocity, the driver's eye movement and head movement. The bioelectric signals of the driver include the electrocardiogram, electroencephalogram, electromyogram and skin electricity signals of the driver. Therefore, the input signals of the link are the driver's control signals and the bioelectric signals of the driver. The output signal is the human situation assessment result H ms ; The second link is traffic situation assessment. The traffic situation assessment is performed based on the dynamic scene signals and the static scene signals. The dynamic scene signals include dynamic vehicle signals and dynamic pedestrian signals. The static scene signals include lane line signals, traffic sign signals and curb stone signals. The input signals of the traffic situation assessment are the dynamic scene signals and the static scene signals. The output signal of the traffic situation assessment is the traffic situation assessment result T fs . Step three, mixed situation fusion, the output signals of the above step one and step two are fused, through time alignment and spatial coordinate conversion of H ms and T fs , the mixed situation fusion with multiple data modalities and scene elements is realized, the input signal of this step is the human situation assessment result H ms output by the above step one and the traffic situation assessment result T fs output by step two; the output signal is the mixed situation fusion result M us ; The specific steps of step three in the first step are as follows: Link one, driving ability evaluation of the driver, the link calculates the comprehensive control ability of the driver to the vehicle at the current time, that is, evaluates the driving ability of the driver at the current time, through the state and coupling of the driver-vehicle-road-environment at the current time, adopts a typical system identification model as a driving ability evaluation model, inputs the input signals of the link into the driving ability evaluation model, and then obtains a quantitative driving ability evaluation result, the input signals of the link are the human situation evaluation result H ms , vehicle state signals, and vehicle-road coupling state signals; the output signal is the driving ability evaluation result of the driver; The second link is self-driving system driving ability evaluation. The comprehensive control ability of the vehicle by the self-driving system at the current time is calculated, that is, the self-driving system driving ability, through the traffic situation at the current time and all model parameters of the perception and decision-making two levels in the automatic driving system. The state variables corresponding to the input data of this link are given corresponding weight values, and a linear function with weight is used for calculation, so as to realize the evaluation of the self-driving system driving ability. The input signal of this link is the traffic situation evaluation result T fs and the automatic driving system parameters; the output signal is the self-driving system driving ability evaluation result. Step three, driving right planning, through the driving ability evaluation results of the driver output by step one and the driving ability evaluation results of the self-driving system output by step two, the comprehensive control effect of the driver and the self-driving system on the vehicle at the current time is quantitatively evaluated, therefore, according to the driving ability evaluation results of the driver and the driving ability evaluation results of the self-driving system, the driving right distribution coefficient τ between the driver and the self-driving system at the current time can be calculated by normalization calculation, the input signal of this step is the driving ability evaluation results of the driver and the driving ability evaluation results of the self-driving system, and the output signal is the driving right distribution coefficient τ; The specific steps of step four in the first step are as follows: Step one, short-term online verification result, this step obtains the short-term online verification result by integrating the short-term fast updating decision online verification module in the highly automatic driving system, the decision online verification module is a module for evaluating the effect of human-machine hybrid decision in the highly automatic driving system, the short-term online verification result can be obtained by time alignment and threshold detection of the short-term verification signal output by the module, the input signal of this step is the short-term verification module signal; The output signal is the short-term verification result; Step two, long-term online optimization result, this step obtains the long-term online verification result by integrating the long-term fast updating decision online verification module in the highly automatic driving system, the long-term online verification result can be obtained by time alignment and threshold detection of the long-term verification signal output by the module, the input signal of this step is the long-term verification module signal; The output signal is the long-term verification result; Step three, online evaluation information flow integration, the short-term domain verification result output from step one and the long-term domain verification result output from step two are integrated to obtain an online evaluation result, the integration process includes time domain category labeling and redundant time domain elimination, the input signal of this step is the short-term domain verification result and the long-term domain verification result, and the output signal is the online evaluation result O er .

3. The method of claim 1, wherein: The steps in the second step include the following steps: The specific steps of step one in the second step are as follows: Step one, multi-target learning framework, establish multi-learning targets based on the safety, comfort, functionality and maneuverability of human-machine hybrid decision, extract local subgraphs in the global graph of the driver's reasoning mechanism model, and establish a learning framework for random walk according to the coupling results between relationship clustering and local subgraphs, which specifically includes three parts: reasoning mechanism model definition, entity relationship standard formulation and entity relationship clustering; Step two, global graph random walk, establish the global graph of the driver's reasoning mechanism, and judge the reachability between each entity pair in the global graph, and then solve the reasoning result corresponding to the global graph, which specifically includes three parts: global graph establishment, entity reachability calculation and global reasoning result calculation; Step three, local subgraph random walk, extract specific local relationship subgraphs of the driver's reasoning mechanism from the global graph of the driver's reasoning mechanism, and realize random walk, which specifically includes three parts: local subgraph establishment, entity transition probability matrix calculation and local reasoning result calculation; Step four, fusion reasoning, uniformly match the global reasoning result and the local reasoning result obtained by the above steps two and three, and fuse the reasoning results by using nonlinear mapping logic, which specifically includes two parts: reasoning result normalization calculation and fusion reasoning result calculation; Human situation assessment result H m And human-machine decision consistency comparison result C hm The composition driver reasoning online data flow F lf , that is, F lf ={H ms , C hm};Decision reasoning knowledge base K df And decision graph knowledge base K dm The composition driver reasoning offline knowledge flow F fk , that is, F fk ={K df , K dm}Therefore, the calculation formula of the driver reasoning mechanism model is shown in the following (1): (1) In the formula, G m and L m represent global graph and local sub-graph respectively, f(G m , L m ) is a function of DIMM, f(F lf ) represents a subset function of F fk , therefore, the DIMM model is an inference model based on a multi-objective learning framework and a random walk pattern, in a specific driving scene at a certain time, the multiple iterations and clustering of the cluster |R γ | are realized through the comparison of entity correlation between independent relations γ, and the shared feature value set C γ containing all feature values c γ corresponding to the new cluster formed by the update is realized, and the calculation formula of the inter-cluster similarity function sim(C γ,m , C γ,n ) is shown in the following formula (2): (2) In the formula, operator Π represents the expression for set C. γ Multiply each element in C. γ The subscripts m and n represent two distinct Cs numbered m and n. γ In obtaining |R γ Based on the similarity between |R|, a joint learning classification model is established to couple and constitute each |R|. γ |Path R of inner γ r Classifier structure function f cl (R r The calculation formula for the corresponding joint relation learning model is shown in equation (3) below: (3) In the formula, μ1 and μ2 are the regularization coefficients, respectively, and ω k ω0 and ω0 are the weighting coefficients and their benchmark values, respectively, and b k b0 and b0 are the classification structure deviation coefficient and its benchmark value, respectively, and d k The weight vector bias coefficient, function L(R) ri,p , R ri,q ) is f cl (R r The training loss function of R, where the subscripts p and q represent two distinct R values ​​numbered p and q, respectively. r The subscript i indicates R r The number, N k For R r The quantity, the subscript k represents d k The number, K is d k The number of implementation clusters |R obtained by clustering entity relationships γ |and its corresponding path R r As a constraint, the calculation of the global graph random walk in the second stage above is implemented. The global graph random walk in the second stage above is achieved by extracting G. m The various relations r el and its corresponding c γ Establish a global relational feature model, G m Defined as G m ={g m,i ={h gm,i , R rg,i ra gm,i }, i=1, 2, ..., s}, where g m,i For G m A subgraph, where the subscripts i and s represent the subgraph number and the total number of subgraphs, respectively, g m,i h gm,i R rg,i and ra gm,i G represents the head entity, path, and tail entity of a valid subgraph, respectively. m Zhongyouh gm,i Through R rg,i Arrive at ra gm,i Accessibility p re The calculation formula is shown in equation (4) below: (4) In the formula, sl={h gm,i ∪ra gm,i Let R = {i=1, 2, ..., s}, i=1, 2, ..., s}. ra,i To be with ra gm,i The directly corresponding R rg,i The tail relation elements in the text, sl and ra gm,i The similarity function sim(sl, ra) gm,i The calculation method of ) and the calculation formula (2) for C γ,m and C γ,n The similarity function between them is calculated in the same way, and α represents R. rg,i The corresponding weight matrix allows the global graph random walk model to be expressed as: f(G m )=α·p re The logistic regression algorithm is used to analyze the model f(G). m The parameters are trained using the given parameters, and the sigmoid function is chosen as the normalization function for the results. The normalized global inference result p g The calculation formula is shown in equation (5) below: (5) L m For G m A subset of is defined as L m ={l m,i ={h lm,i , R rl,i ra lm,i }, i=1, 2, ..., z}, where l m,i For L m A subgraph, where the subscripts i and z represent the subgraph number and the total number of subgraphs, respectively, l m,i hl in m,i R rl,i and ra lm,i These represent the head entity, path, and tail entity of a valid subgraph, respectively. Therefore, for L... m When performing random walk computation, the space complexity of the computation is reduced, and L can be directly computed. m The transition probability matrix T between different entities M This leads to the corresponding local inference results, T M The calculation formula is shown in equation (6) below: (6) In the formula, N hl,i and N ral,i According to hl m,i and ra lm,i The constructed diagonal matrix, sp, is T. M The number of transition steps, M l,i For the sp-th step L m The corresponding adjacency matrix, T M The element T corresponding to row a and column b M [a,b] indicates that hl m,i Starting from point A, perform a random walk, and after SP steps, jump to point A. lm,i The probability is expressed using p. l Indicates R rl,i The evaluation results of the local inference results yielded p l The calculation formula is shown in equation (7) below: (7) p g and p l fusion and normalization, the calculation formula of the fusion inference result p f is shown in the following formula (8): (8) In the formula, δ represents a fusion reasoning stability coefficient, used to balance the contribution proportion of p g and p l , and thus the driver reasoning mechanism model in step one in the second step outputs a fusion reasoning result. The specific links of step two in the second step are as follows: Link one, neuron group model, the neuron group model is established to provide personalized classification basis for the "self-like" decision model corresponding to link four described below, which specifically includes four parts of neuron group model definition, feature extraction, classification based on stimulation, and highlight matrix; Link two, deep convolutional network, through the deep learning network, fitting basis for the policy-oriented reinforcement learning model corresponding to link three described below is provided, which specifically includes four parts of behavior network structure definition and evaluation network structure definition; Link three, reinforcement learning model, through the reinforcement learning model, the complex decision mode of the automatic driving system is realized, the optimal action in the reinforcement learning model is calculated by searching for the optimal strategy, and the specific content includes three parts of state definition, reward function, and policy gradient; Link four, "self-like" decision model, the personalized classification basis obtained through link one described above is combined with the deep reinforcement learning process of link two and link three described above, and "self-like" decision data is fused at the data level, and the online decision result of the automatic driving system is finally output by judging whether the iteration effect reaches the preset threshold, and the specific content includes two parts of "self-like" data fusion and threshold judgment; T fs and C hm comprise the neuron group online classification data stream F cf , i.e. F cf ={T fs , C hm};K dd and K dm comprise the automatic driving offline training knowledge stream F ak , i.e. F ak ={K dd , K dm}, and the calculation formula of the neuron group model is shown in the following formula (9): (9) According to F ak extracting specific features and learning using an incremental learning rule based on Hebbian learning and Hebb learning, presetting an N G dimension category label, and taking the extracted specific features as a conditioned stimulus, establishing a personalized neuron group model belonging to a specific classification dimension, and the incremental learning rule calculation formula is shown in the following formula (10): (10) where ΔCHL, ΔHeb, and Δβ represent the Hebbian learning rule, the hetero-associative learning rule, and the incremental learning rule, respectively, β represents the synaptic matrix corresponding to the incremental learning rule, g j is the presynaptic activation, h i is the postsynaptic activation, the subscripts i, j correspond to the i-th row and the j-th column of the specific feature element extracted by F ak , + and - represent the addition stage and the subtraction stage, respectively, ζ is a weight coefficient, κ is a learning rate, a kWTA function is used to obtain a sparse distribution representation in the feature matrix, and the first r activated units are extracted, and the corresponding inhibition function f r is calculated as shown in the following equation (11): (11) where χ is a suppression threshold, the first r activated units are in the active mode, and the remaining h activated units are in the inactive mode. i The corresponding computation is given by the normalized dot product of the ith row of β with g j The corresponding computation is given by the normalized dot product of the ith row of β with g (12) h i is calculated i h i,max is calculated i,max h ij is calculated G h c is calculated c,1 h c,2 is calculated c,NG h c is calculated G h c is calculated will be classified by M c K dd and K dm sub-data set K dd,NG and K dm,NG , as the above-mentioned link two deep convolutional network and link three reinforcement learning model corresponding model training data, deep convolutional network and reinforcement learning model jointly constitute a deep reinforcement learning model DRLM, DRLM model definition corresponding to the formula as shown in equation (13): DRLM = {S DR , A DR , P DR , R DR , λ} (13) where S DR is the state space of the vehicle for the autonomous driving system, A DR is the action space of the maneuver for the autonomous driving, P DR is the state transition probability distribution of the autonomous driving system, R DR is the reward function, λ = [λ1, λ2, λ3, λ4] is the discount factor, R DR is calculated as shown in the following equation (14): (14) In the formula, R safety , R goal , R law and R comft respectively represent a safety reward function, a time reward function, a traffic regulation reward function and a comfort reward function, corresponding to the safety, maneuverability, traffic rules and driver comfort of the automatic driving system for the driving task, for realizing the parameterized estimation of the continuous data sequence in the automatic driving decision, that is, the accurate estimation of P DR when the continuous data sequence mode is S DR , the search of π(a RDL , s RDL ) is carried out by using a deep convolutional network based on a policy gradient, and the deep convolutional network includes a behavior network L(θ Q ) and an evaluation network ▽ θμ L(θ Q ), and the definition calculation formulas of the corresponding networks are shown in the following formulas (15) and (16). (15) (16) where E is an operator for expectation function, subscripts t and t+1 represent the current time step and the next time step, respectively, θ Q and θ μ represent the nonlinear estimators of two neural network structures, Q is an action value function, representing the expected discounted return obtained by taking action A DR in state S DR ; μ is a mapping function constructed so that a RDL = μ(s RDL ; θ μ ) is true, and ∇ θμ L(θ Q ) is optimized for π(a RDL , s RDL ) through a gradient update method. The calculated π(a) RDL , s RDL Data fusion and updating are performed in the memory pool to ultimately obtain π. * (a RDL , s RDL The calculation formula for the update mode is shown in equation (17): (17) where φ = 10 -4 a positive constant; Therefore, the high-level "self-like" decision model based on brain-like computing in step two of the second step outputs the "self-like" decision result of the automatic driving system; The specific links of step three in the second step are as follows: Link one, driver decision graph, the input signal corresponding to this link is the fusion reasoning result p output by link four corresponding to step one in the second step f and the human situation assessment result H corresponding to link one of step one in the first step ms , and according to the knowledge expression rules of the decision graph knowledge base K corresponding to step three in the third step dm , finally integrate p f , H ms and the driver control signal into the driver decision graph in the form of a classification tree graph structure; Step 2, Autonomous Driving Decision Map: The input signal for this step is the optimal decision strategy π output from Step 2, which corresponds to Step 3 in Step 2. * (a RDL , s RDL And the traffic situation assessment result T corresponding to the first step of step one in the first step. fs Similarly, according to the decision graph knowledge base K corresponding to step three in step three, dm The knowledge representation rules are ultimately expressed in the form of a classification tree graph structure, which represents π. * (a RDL , s RDL ), T fs And the control signals of the autonomous driving system are integrated into the autonomous driving system decision map; Step three, man-machine atlas comparison and prediction, the input signal corresponding to this step is the hybrid situation fusion result of step one in the first step, the driver decision atlas output by step three in the second step and the self-driving system decision atlas output by step two; the output signal is the man-machine decision consistency rate C DK Firstly, this step predicts the vehicle state and the spatio-temporal evolution rule of the driving track of the intelligent vehicle in the short time domain of 0-10 seconds from the current time under the condition of driving only by the driver through the driver decision atlas; at the same time, it predicts the vehicle state and the spatio-temporal evolution rule of the driving track of the intelligent vehicle in the short time domain of 0-10 seconds from the current time under the condition of driving only by the automatic driving system through the self-driving system decision atlas, then solves the similarity of the vehicle state and the driving track in the same short time domain under the condition of driving only by the driver and driving only by the automatic driving system, and takes the similarity result as the output signal C DK of this step; The specific links of step four in the second step are as follows: Link one, subdivision criterion library, the subdivision criterion library stores the subdivision rules of the human-machine driving right corresponding to the human-machine mixed decision layer of the highly automatic driving system in an offline storage manner, and the subdivision criterion library includes a traffic rule library, a maneuverability criterion library, a safety criterion library, and a comfort criterion library; the traffic rule library stores a set of traffic rules corresponding to urban traffic; the maneuverability criterion library stores a set of rules constructed to ensure vehicle driving efficiency; the safety criterion library stores a set of rules to ensure that the vehicle has safe longitudinal and lateral driving performance in emergency conditions; and the comfort criterion library stores a set of rules to ensure that the driver and passengers of the vehicle are in a comfortable state during vehicle driving; Link two, constraint rule, the constraint rule stores the configuration and value range of the constraint rule for constraining the knowledge and model parameters of the human-machine mixed decision layer in an offline storage manner, including the constraint rule for the value range of the knowledge of the decision model in step two in the third step, the threshold constraint rule for the vehicle state when only the driver drives, and the threshold constraint rule for the vehicle state when only the automatic driving system drives; Step 3, Driving Rights Optimization Algorithm: The input signals for Step 3 in Step 4 of Step 2 are the driving rights allocation coefficient τ from Step 3 of Step 1, and C from Step 3 of Step 2. DK The second step involves the subdivision criteria corresponding to the first stage of step four, the constraint rules corresponding to the second stage of step four, the driver's control signals at the current moment, and the control signals of the automated driving system at the current moment. The output signals are the driver's driving weight value and the automated driving system's driving weight value for the highly automated driving system. This stage first establishes C. DK The two-dimensional linear programming plane corresponding to the output τ of stage three in step one is used. Then, the subdivision criteria of stage one output corresponding to step four in step two and the constraint rules of stage two output are used to regulate and constrain the driver's control signals and the control signals of the autonomous driving system at the current moment, and finally the optimized driving rights allocation coefficient τ is obtained. * .

4. The hybrid augmented intelligence based highly consistent human-machine hybrid decision making method of claim 1, wherein: The links included in each step of the third step are as follows: The specific links of step two in the third step are as follows: Link one, knowledge element ontology modeling, the model structure and definition of the knowledge of the decision model in step two in the third step are specified, and the knowledge source of the knowledge of the decision model is determined, and then the meta ontology model of the knowledge of the decision model is established by establishing a process model and a concept model, the input signal of this link is the human-machine mixed enhanced decision model parameter data stream corresponding to step two in the first step, and the output signal is the meta ontology model; Step two, the decision model ontology design, the step adopts the meta ontology development mode from the self orientation, taking the main trunk class of the meta ontology as the core, defining three top classes of the top model class, the top data class and the logic expression class respectively, realizing the design of the decision model ontology, the input signal of the step is the human-machine mixed enhanced decision model internal parameter data flow corresponding to step two in the first step, and the output signal is the primary decision model knowledge base with the top class of the meta ontology; Step three, the model knowledge rule design, the step carries out reasoning and query on the concept model of the meta ontology, constructs the semantic rule base through the knowledge path management rule and the knowledge action operation rule, the input signal of the step is the primary decision model knowledge base output by step two corresponding to step two in the third step, and the output signal is the updated decision model knowledge base with the semantic rule; Step four, model knowledge architecture design, this step establishes a semantic web framework as a knowledge engine, integrates dynamic knowledge element ontology with static knowledge element ontology, in addition, this step establishes a semantic middleware, processes multi-modal man-machine hybrid augmented decision model internal parameter data flow into a unified meta ontology structure in a semantic mapping manner, the input signal of this step is the man-machine hybrid augmented decision model internal parameter data flow corresponding to step two in the first step and the updated decision model knowledge base output by step two corresponding to step three, and the output signal is decision model knowledge K dd ; The specific four steps of step three in the third step are as follows: Step one, the construction of knowledge graph ontology, this step establishes the knowledge graph ontology for decision graph knowledge, adopts the field ontology mode to synthesize the effective knowledge ontology for human-machine hybrid decision, the input signal of this step is the human situation assessment result H output by step one of the first step ms , the traffic situation assessment result T output by step two of the first step fs , and the new mode decision reasoning knowledge output by step five of the third step, the output signal is the effective knowledge ontology set, considering that there are time continuous data sequences in H ms and T fs , the effective knowledge ontology synthesized by the field ontology mode is divided into five elements of concept, relation, function, axiom and individual; Step two, the atlas semantic rule making, the step standardizes the effective knowledge ontology at the semantic level through the specified semantic rule, establishes the basic semantic rule in the combination mode of the abstract syntax and the concrete syntax, adds the consistency check, the classification, the identification and the prediction four reasoning functions in the basic semantic rule, realizes the semantic rule with the reasoning mechanism, the step has no input signal, and the output signal is the semantic rule; Step three, the atlas semantic mapping, the step maps the effective knowledge ontology set and the semantic rule in the form of the relation table, forms the decision atlas case, adopts the unique target address to identify different effective knowledge ontologies, realizes the reasonable mapping of the multi-element ontology relation through the relation class composed of the effective knowledge ontology and the corresponding semantic, and finally forms the decision atlas case library composed of the multiple relation classes, the semantic rule and the mutual mapping relation; the input signal of the step is the effective knowledge ontology set output by step one corresponding to step three in the third step and the semantic rule output by step two, and the output signal is the decision atlas case library; Step four, case knowledge calling, this step formulates a case access mechanism, according to the mode of searching the effective knowledge ontology in the decision graph case base first, then searching the similar relationship, and finally searching the similar semantics, to realize the calling of specific cases in the decision graph case base. The case access mechanism established in this step and the decision graph case base together constitute the decision graph knowledge base. The input signal of this step is the decision graph case base output by step three in the third step, and the output signal is the decision graph knowledge base K dm ; The specific steps of step four in the third step are as follows: Step one, the knowledge rule oriented model, the step establishes the knowledge rule oriented model, and then converts the pre-collected driver offline database into the driver offline knowledge base with the knowledge rule as the constraint; the input signal of the step is the driver offline database information and the new mode decision reasoning knowledge output by step five in the third step; the output signal is the knowledge rule oriented function and the driver offline knowledge base; Step two, the knowledge reasoning vector modeling, the step establishes the knowledge reasoning vector based on the vehicle dynamics characteristics and the vehicle-road coupling characteristics, solves the reasoning and prediction results of the vehicle state change corresponding to the driver offline knowledge base in the form of the knowledge data structure, and integrates into the knowledge reasoning vector set; the input signal of the step is the knowledge rule oriented function and the driver offline knowledge base output by step one corresponding to step four in the third step; the output signal is the knowledge reasoning vector set; Step three, hierarchical logical reasoning classification, based on the scene danger level and driving mode, a hierarchical logical reasoning classification method is established, and under this classification method, the knowledge reasoning vector set is divided into a typical knowledge reasoning vector subset, the input signal of this step is the knowledge reasoning vector set output by step four of step three in the third step; The output signal is the classified knowledge reasoning vector subset; Step four, reasoning knowledge architecture generation, this step arranges the knowledge reasoning vector subset according to the priority, and sets the search logic algorithm for the knowledge reasoning vector subset. The search logic algorithm standardizes the reasoning knowledge architecture, integrates the search logic algorithm with the knowledge vector subset, and forms the decision reasoning knowledge base K df The input signal of this step is the knowledge reasoning vector subset corresponding to step four of step three in the third step; the output signal is the decision reasoning knowledge base K df ; The specific steps of step five in the third step are as follows: Step one, data cleaning, which cleans the human situation assessment result H output by step one in the first step ms , the traffic situation assessment result T fs , the mixed situation fusion result M us , the online assessment result O output by step four in the first step er , and the historical data curve of the output τ * of step three in the fourth step, specifically including: redundant data cleaning and basic data cleaning. The second link, feature extraction, extracts the human situation assessment result H output by the first link corresponding to step five in the third step after cleaning ms , the traffic situation assessment result T fs , the mixed situation fusion result M us , the online assessment result O er , and the effective features of the historical data curve of tau * , and forms an effective feature set, specifically including "human-traffic" situation feature extraction, human-machine mixed decision consistency feature extraction, and feature fusion. Step three, similarity comparison and rating, the effective feature set output by step one corresponding to step five in the third step is compared with the effective features of the decision model knowledge corresponding to step two, the decision graph knowledge corresponding to step three, and the decision reasoning knowledge corresponding to step four, and the similarity level of the effective feature set output by step one corresponding to step five in the third step is calculated, which includes four parts: decision model knowledge similarity comparison, decision reasoning knowledge similarity comparison, decision graph knowledge similarity comparison, and similarity level determination; Step 4: New Knowledge Synthesis. This step first determines the human posture assessment result H corresponding to the valid feature set with a similarity level below a threshold. ms Traffic situation assessment results T fs Hybrid situation fusion result M us And the knowledge types corresponding to the effective features of the online evaluation results data, and then new knowledge is synthesized according to the knowledge data structure of the corresponding knowledge type, which specifically includes four parts: knowledge type classification, decision model knowledge synthesis, decision reasoning knowledge synthesis, and decision graph knowledge synthesis. To establish a new knowledge synthesis mechanism model, first, the input data of step five in the third step corresponding to link one is composed into a new knowledge synthesis dataset N ck ={H ms , T fs , M us , O er} , the configuration redundancy data and physical relationship redundancy data existing in N ck are cleaned up by using a similar data comparison algorithm, and considering that each element data has consistency on the time axis, the calculation formula of the data comparison number S num and the similar repetition rate D rt of the data in the comparison window in the similar data comparison algorithm is shown in the following formula (18): (18) In the formula, Tm is the timestamp of the comparison window, Da is the matrix formed by the comparison data corresponding to Tm, the subscript t0 is the first comparison time, and f Tm For N ck The sampling frequency, Δt is the window length, and d num The number of similar duplicate records within the window is used as a basis for filtering the data corresponding to each element in the cleaned dataset to obtain the cleaned dataset N. ’ ck ; To achieve N ’ ck Pattern determination requires N ’ ck Typical features are extracted from the data, and the data is classified according to similarity. N ’ ck Typical features F T Including "human-transportation" situational characteristics M ST And the consistency feature of human-machine hybrid decision-making C HT The "human-traffic" situational characteristics are used to characterize the statistical characteristics of the "human-traffic" situational data at the current moment, while the human-machine hybrid decision-making consistency characteristics are used to characterize the indicative characteristics of the human-machine hybrid decision-making effect at the current moment. T The calculation formula is shown in equation (19): (19) In the formula, f ST e ST and m ST M respectively ST The corresponding frequency domain characteristics, extreme value characteristics, and mean characteristics, C DK Human-machine decision consistency rate, according to F T N ’ ck The decision model knowledge K is respectively compared with the knowledge K. dd Decision Graph Knowledge Base K dm and the decision reasoning knowledge base K df The similarity of each knowledge element in the text is compared, and the similarity level D is evaluated. s The similarity comparison is divided into two parts: data transformation and semantic primitive parameterization, and similarity calculation. The data transformation part will convert N... ’ ck Convert them to K respectively dd K dm and K df The knowledge data structure in it will constitute K dd K dm and K df The smallest unit of knowledge in a knowledge base is defined as a semantic primitive. The parameterization process of a semantic primitive is the process of setting the semantic primitive value ψ for a specific knowledge base. For N ’ ck When there are data modalities in the data that cannot be converted into the knowledge data structure of the corresponding knowledge base, the corresponding D will be... s Set to 0; N ’ ck When the data in D can be transformed into a data modality of the knowledge data structure in the corresponding knowledge base, s The calculation formula is shown in equation (20): (20) where d is represents K dd , K dm and K df respectively. ’ ck The distance between the transformed knowledge and N ’ ck The transformed knowledge and K dd , K dm and K df are calculated for similarity, and similarity level is determined according to D s and C DK . The determination logic is that N s is considered to be able to synthesize a new knowledge only when the value of D s is within the threshold range of the corresponding knowledge base, and the value of C DK is lower than the threshold. ’ ck The new knowledge is synthesized, N ’ ck is represented according to the knowledge data structure of the corresponding knowledge base, and is merged into the corresponding knowledge base, thus completing the synthesis of the new knowledge.

5. The hybrid augmented intelligence based highly consistent human-machine hybrid decision making method of claim 1, wherein: The steps in the fourth step include the following steps: The specific steps of step two in the fourth step are as follows: Step one, online evaluation knowledge base framework integration, this step integrates the modeling process of online evaluation result data in the new knowledge synthesis mechanism model in step five of the third step, and combines it into an online evaluation knowledge base framework, and then outputs it to the decision online verification module in the highly automated driving system; Step two, online evaluation knowledge base update, this step integrates the online evaluation knowledge output by step four corresponding to step five in the third step according to the knowledge data structure of the synthesized new knowledge, and then outputs it to the decision online verification module in the highly automated driving system; The specific steps of step four in the fourth step are as follows: Step one, vehicle dynamics inverse solution, this step solves the ideal control quantity information corresponding to the current time by performing dynamics inverse solution on the vehicle state quantity at the current time, the input signal of this step is the vehicle state quantity corresponding to the current time; The output signal is the ideal control quantity information corresponding to the current time; Step two, control rule framework, this step establishes a control algorithm framework for calculating the expected control quantity, which is realized by using typical control algorithms in control theory, and is used to calculate the expected control quantity; Step three, expected control quantity calculation, this step combines the ideal control quantity information corresponding to the current time output by step one corresponding to step four in the fourth step and the control algorithm framework established by step two to calculate the final expected control quantity, the input signal of this step is the ideal control quantity information under the control rule framework; The output signal is the expected control quantity.

Citation Information

Patent Citations

  • A Highway Overtaking Behavior Decision-Making Method for Autonomous Vehicles

    CN106874597B

  • Generation method and device for decision network model of vehicle automatic driving

    CN107169567A

  • A method and system for transferring human control in autonomous vehicles

    CN107943046B