Methods, systems, and devices for a user - understandable and interpretable learning model
By capturing the semantic content of the original classifier by extending convolutional neural network (CNN), an interpretable learning system is built to generate user-understandable explanations, solving the problem that users cannot resolve the solution process in autonomous and semi-autonomous driving systems, and improving the clarity of user experience and driving manipulation.
Patent Information
- Application Number
- CN202110526326.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-17
- Filing Date
- 2021-05-14
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-05-14
AI Technical Summary
The existing autonomous and semi-autonomous driving systems are unable to understand the system's decision-making process due to the black box nature of the data-driven learning model, and lack the necessary tools and systems to provide clarity in driving manipulation, resulting in poor user experience.
Capturing the semantic content of the original classifier by extending convolutional neural networks (CNNs), an interpretable learning system is built to generate user-understandable explanations, including receiving input data, determining semantic functions, calculating semantic accuracy, extending interpretable classifiers and training connection sets to generate user-understandable output explanations.
Improve users' understanding of the decision-making process of autonomous and semi-autonomous driving systems, enhance user experience and technicians' insight into automated behavior, and provide clarity of driving manipulation.
Smart Images

Figure CN114118349B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the priority of U.S. Provisional Patent Application 63 / 071,135, filed on August 27, 2020, the content of which is incorporated herein by reference.
[0003] The technology generally relates to methods, systems, and devices for user - interpretable learning outputs for autonomous, semi - autonomous, and conventional vehicles, and more particularly to methods, systems, and devices of an interpretable learning system of an extended convolutional neural network (CNN) having a learning architecture and capturing the semantics learned by an original classifier to compute a predicted output with a user - understandable explanation. Background Art
[0004] The public is reluctant to use autonomous and semi - autonomous driving features. This is due to the novelty of the technology and the opaque nature of autonomous and semi - autonomous vehicles when performing various driving maneuvers. The public views these systems with the opaque nature and has no ability or limited ability to independently judge, evaluate, or predict the current or future autonomous driving maneuvers being performed. This is because the current data - driven learning models for autonomous and semi - autonomous vehicles are black boxes. These systems can very accurately perform and predict the classification of data samples regarding driving conditions and the surrounding environment, but for the user, there is no clarity regarding the current or upcoming automated process. Thus, from the perspective of a human user, the model is a black box that cannot correctly explain why the system obtains an output and an action based on its calculations. There is a lack of necessary tools and systems to provide sufficient feedback to the driver regarding driving maneuvers to improve the inherent discomfort level felt by the driver when using autonomous and semi - autonomous systems.
[0005] The use of connectionist learning models can be or is considered to successfully predict data sample classes with high accuracy. However, the level of interpretability of the black boxes used to implement connectionist learning models is very low.
[0006] Therefore, users can benefit from a deep understanding of why complex machine - learning models produce a certain output.
[0007] There is a desire to implement interpretable learning methods, systems, and devices that utilize a learning architecture to provide an extension to neural networks (NNs) such as convolutional neural networks (CNNs), process and capture the semantic content learned by an original classifier for further processing by the implemented CNN (or any other NN), and generate a user - understandable semantic explanation of the predicted output.
[0008] There is a desire for methods, systems, and devices to implement an extended architecture and, through the extension, process the semantic content of a given classifier in an interpretable classifier; construct and train a set of connections for an interpretable classifier for a new learning system that can verify the output to a user, and can use appropriate criteria to quantify the output and use semantic expressions to explain the success or rationale.
[0009] Further, other desirable features and characteristics of the present disclosure will become apparent from the following detailed description and the appended claims, in conjunction with the accompanying drawings and the foregoing technical field and background. SUMMARY OF THE INVENTION
[0010] Methods, systems, and devices are provided to improve autonomous or semi-autonomous vehicles through an interpretable learning system that extends a convolutional neural network (CNN) or other neural network (NN) with a learning architecture to capture the semantics learned by the original classifier to compute a user-understandable explanation of the predicted output.
[0011] In one exemplary embodiment, a method for constructing an interpretable user output is provided. The method includes: receiving input data of a feature set for processing at a neural network (NN) composed of multiple layers of an original classifier, where the original classifier has been frozen with a set of weights related to the features of the input data; determining a semantic function to classify data samples into semantic categories; determining the semantic accuracy level of each layer of the original classifier within the neural network, where the original classifier is a trained model; for training samples of the semantic categories, computing a representative vector having an average activation of nodes of each layer for evaluation by computing the distance between samples in a test set and available layers of each semantic category; computing multiple test samples for each layer and each semantic category that are closest to each other within each layer to specify a layer with the highest score, where the highest score represents the best semantics in each test sample; extending the specified layer to the neural network through a category branch to extract semantic data samples of the semantic content and for extending the interpretable classifier through the extension at the neural network based on multiple semantic categories; training a set of connections of the interpretable classifier of the neural network to compute a set of output explanations with an accuracy metric related to each output explanation based on at least one semantic category among multiple semantic categories; and comparing the accuracy metric of each output explanation by the trained interpretable classifier for each semantic category based on the extracted semantic data samples to generate an output explanation in a user-understandable format.
[0012] In various exemplary embodiments, the method includes: calculating multiple interpretations as outputs; and calculating counterfactual interpretations of the constructed interpretations, wherein the designated layer is the best semantic content layer. The method includes: configuring an interpretable classifier having multiple layers, wherein each layer is configured with a set of weights; labeling the extracted input data samples of the interpretable classifier with a defined set of semantic categories associated with the best semantic content layer; calculating activation centroids for each of the multiple layers of the interpretable classifier; and based on the activation centroids, calculating the layer of the interpretable classifier having the maximum semantic accuracy. The method includes: calculating the maximum semantic accuracy for each semantic category in the set of semantic categories based on a semantic function and optionally a redefined semantic function.
[0013] The method includes: based on the calculated amount of semantic accuracy, designating at least one layer of the original classifier by running the input data samples through each layer of the original classifier to observe the activation counts of the nodes included in each layer; for each semantic category, determining a set of distances between the activation of each node in the node layer for each sample and the average of the activations of the nodes in the corresponding node layer for all data samples in the training set; repeating the determination step for each successive input data sample until the average vector value of the nodes in each layer is calculated for each received input data sample and each semantic category in each layer; retaining the average vector value of each node in each layer of the original classifier; in each layer, scoring the average of the nodes in the vector, wherein the average is configured as a vector; summing the set of vector value scores for each layer received for all data samples in the dataset, wherein the vector values with fractions in each layer are used to determine the layer having the maximum value; and designating the layer having the maximum value as the best semantic content layer.
[0014] The neural network includes a convolutional neural network (CNN) and an extension of the CNN. The method further includes: by extending the original classifier, extending the interpretable classifier via an extension of the NN for driving comfort to extract semantic content via the designated layer to predict semantic categories corresponding to a set of features of vehicle dynamics and vehicle control. The method further includes: while freezing the set of weights of the original classifier, training the set of connections of the NN extension of the interpretable classifier for driving comfort; and adding an interpreter to the output of the NN extension to be related to the data samples of the extracted semantic content. The method further includes: by extending the original classifier, extending the interpretable classifier for trajectory prediction to extract semantic content via the designated layer to predict semantic categories related to a set of semantic features of sample data including convolutional social pooling. The method further includes training the set of connections of the interpretable classifier for trajectory planning, and adding an interpreter to the output to be related to the data samples of the set of semantic features of sample data including convolutional social pooling.
[0015] In another exemplary embodiment, a system for constructing an interpretable user output is provided. The system includes: a processor configured to receive input data for processing a feature set at a neural network composed of multiple layers of an original classifier, where the original classifier has been frozen with a set of weights related to the features of the input data; the processor is configured to determine a semantic function to classify data samples into semantic categories and determine the semantic accuracy level of each layer of the original classifier within the neural network, where the original classifier is a trained model; the processor is configured to calculate a representative vector with the average activation of nodes of each layer for training samples of the semantic categories to evaluate by calculating the distances of samples in the test set from the available layers of each semantic category; the processor is configured to calculate multiple test samples for each layer and each semantic category that are closest to each other within each layer to specify the layer with the highest score, where the highest score represents the best semantics in each test sample, and the specified layer is the best semantic content layer; the processor is configured to extend the best semantic content layer to the neural network through a category branch to extract a set of semantic data samples from the semantic content of the original classifier into the neural network and extend the NN with an interpretable classifier to define multiple semantic categories; the processor is configured to train a new set of connections of the interpretable classifier of the neural network to calculate a set of output explanations composed of explanations with an accuracy metric related to each output explanation based on at least one semantic category among the multiple semantic categories; and the processor is configured to compare the accuracy metric of each output explanation based on the extracted semantic data samples for each semantic category by the trained interpretable classifier and generate an output explanation in a user - understandable format.
[0016] In various exemplary embodiments, the system includes a processor configured to calculate multiple explanations as output and counterfactual explanations of the composed explanations. The system further includes a processor configured to: apply an interpretable classifier with multiple layers, where each layer is configured with a set of weights; label the extracted input data samples of the interpretable classifier with a defined set of semantic categories related to the best semantic content layer; calculate the activation centroid for each layer of the interpretable classifier; and based on the activation centroid, calculate the layer of the interpretable classifier with the maximum semantic accuracy.
[0017] The system further includes a processor configured to: calculate a maximum semantic accuracy rate for each semantic category in a defined set of semantic categories based on a semantic function and optionally a re - defined semantic function, where the re - defined semantic function provides more abstract semantic categories. The system further includes a processor configured to: by processing an input data sample through each layer of an original classifier to determine the activation number of nodes included in each layer, based on the calculated amount of semantic accuracy rate, specify at least one layer of the original classifier; in response to determining the node activation number, for each semantic category, determine a set of distances between the activation of each node in the node layer for each sample and the average of the activations of the nodes in the corresponding node layer for all data samples in the training set; repeat the determination for each successive input data sample until the average vector value of the nodes in each layer is calculated for each received input data sample and each semantic category; in each layer of the original classifier, retain the average vector value of each node; in each layer, score the average of the nodes in the vector, where the average is configured as a vector; sum the set of vector value scores for each layer received for all data samples in the dataset, where the vector values with scores in each layer are used to determine the layer with the maximum value; and specify the layer with the maximum value as the best semantic content layer. The system further includes a processor configured to: by extending the original classifier, extend the NN with an interpretable classifier for driving comfort through an extension of a convolutional neural network (CNN) to extract semantic content via a specified layer to predict semantic categories corresponding to a set of features of vehicle dynamics and vehicle control; only train the weight set of the new connections from the original CNN to the extended architecture, where the weight set of the original CNN trained for driving comfort is kept frozen; and add an interpreter to the output of the CNN to be related to the data samples related to the extracted semantic content. The system further includes a processor configured to: by extending the original classifier, extend the NN with an interpretable classifier for trajectory prediction to extract semantic content via a specified layer to predict semantic categories related to a set of features of sample data including convolutional social pooling; only train the weights of the new connections from the original CNN to the extended architecture, where the weight set of the original CNN trained for predicting driving trajectories is kept frozen; and add an interpreter to the output of the CNN to be related to the data samples related to the set of features of sample data including convolutional social pooling.
[0018] In yet another exemplary embodiment, a device for performing an interpretable classifier is provided. The device includes at least one processor deployed in a vehicle, the at least one processor being programmed to: receive input data of a feature set for processing at a convolutional neural network (CNN) constituted by a multi-classifier layer of an original classifier, wherein the original classifier has been frozen with a set of weights related to the features of the input data; determine a semantic function to classify data samples into semantic categories; determine a semantic accuracy level for each layer of the original classifier within the neural network, wherein the original classifier is a trained model; calculate a representative vector with an average activation of nodes of each layer for training samples of the semantic categories for evaluation by calculating distances of samples in a test set from available layers of each semantic category; calculate multiple test samples for each layer and each semantic category that are closest to each other within each layer to specify a layer with the highest score, the highest score representing the best semantics in each test sample; redefine the semantic function to change the semantic accuracy by determining the semantic accuracy of each layer of the original classifier through the trained model to generate more semantic content, wherein the redefinition of the semantic function is an optional step to provide more abstract semantic categories; extend the specified layer to the neural network through a category branch to extract semantic data samples derived from the semantic content and construct an interpretable classifier of the neural network to define multiple semantic categories; train a connection set of the interpretable classifier of the neural network to calculate an output interpretation set based on at least one semantic category among the multiple semantic categories with an accuracy metric related to each output interpretation, wherein the output interpretation is a well-defined syntactic sentence; and compare the accuracy metric of each output interpretation for each semantic category based on the extracted semantic data samples by the trained interpretable classifier to generate an output interpretation in a user-understandable format.
[0019] In various exemplary embodiments, the device includes at least one processor programmed to configure the interpretable classifier to have multiple layers, wherein each layer is configured with a set of weights; label the extracted input data samples of the interpretable classifier with a defined set of semantic categories; associate the defined set of semantic categories with a layer of best semantic content; calculate an activation centroid for each layer of the multiple layers of the interpretable classifier; and calculate a layer of the interpretable classifier with the maximum semantic accuracy based on the activation centroid.
[0020] The device further includes at least one processor programmed to: calculate the maximum semantic accuracy based on the redefined semantic function, which includes redefined semantic categories for the set of semantic categories; and create an output interpretation in a user-understandable format using the redefined semantic categories. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Hereinafter, exemplary embodiments will be described in conjunction with the following drawings, in which like reference numerals represent like elements, and wherein:
[0022] Figure 1 A block diagram depicting an example vehicle, which may include a processor of an explainable learning system, is shown according to an exemplary embodiment;
[0023] Figure 2 is a functional diagram illustrating an exemplary solution for using category branches or nodes of an explainable learning system according to an embodiment;
[0024] Figure 3A and 3B is an exemplary flow chart of an explainable learning system according to an embodiment;
[0025] Figure 4A , 4B , 4C, 4D, 4E, 4F, and 4G are exemplary flow charts according to an embodiment, showing the calculation of a specified layer in an original classifier, the original classifier containing most of the semantic data of the language with respect to the semantic category, configuring a sample data set with optimized calculation, and calculating the semantic accuracy of each layer of the explainable learning system;
[0026] Figure 5 is another flow chart of an explainable learning system according to an embodiment;
[0027] Figure 6 is an example diagram of an exemplary driving comfort prediction of an explainable learning system according to an embodiment;
[0028] Figure 7 is an example diagram for explaining vehicle trajectory prediction of an explainable learning system according to an embodiment;
[0029] Figure 8 is an exemplary flow chart of an explainable CNN for vehicle trajectory prediction in an explainable learning system according to an embodiment; and
[0030] Figure 9 is an example diagram of a CNN for driving comfort prediction of an explainable learning system according to an embodiment. DETAILED DESCRIPTION
[0031] The following detailed description is merely exemplary in nature and is not intended to limit the application and uses. Furthermore, there is no intention to be bound by any expressed or implied theory presented in the preceding technical field, background, brief summary or the following detailed description.
[0032] As used herein, the term "module" refers to any hardware, software, firmware, electronic control component, processing logic, and / or processor device, individually or in any combination, including but not limited to: application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), electronic circuits, processors (shared, dedicated, or grouped) executing one or more software or firmware programs and memories, combinational logic circuits, and / or other suitable components that provide the described functionality.
[0033] Embodiments of the present disclosure may be described herein in terms of functional and / or logical block components, as well as various processing steps. It should be understood that such block components may be implemented by any number of hardware, software, and / or firmware components configured to perform the specified functions. For example, embodiments of the present disclosure may employ various integrated circuit components, such as memory elements, digital signal processing elements, logic elements, look-up tables, etc., which may perform various functions under the control of one or more microprocessors or other control devices. Additionally, those skilled in the art will appreciate that embodiments of the present disclosure may be practiced in conjunction with any number of systems, and the systems described herein are merely exemplary embodiments of the present disclosure.
[0034] The performance of an automated planning system can be measured as a function of the value of the actions it generates and the cost of achieving them. The present disclosure describes methods, systems, and devices for the intersection of learning models and model-based planning work, which enable the interpretation of the internal workings of a black box, which can provide environmental information, and which helps guide a model-based planner in deciding what actual actions to take. Emphasis is placed on computing the semantics of such planning environments. By adding semantics to planning knowledge, the value of user decision-making and user trust can be increased. The present disclosure describes methods for assisting a user in understanding the high-level features computed by an AI system that performs functions based on the reasons of a complex decision-making process. Among other things, the present disclosure describes an interpreter component that is responsible for generating human-readable explanations, which include the format of the prediction and the associated semantic interpretation in well-defined syntactic sentences.
[0035] In an exemplary embodiment, the present disclosure describes an algorithmic method for learning semantics in two classifiers that are trained in two different domains related to autonomous driving to interpret predictive classifications of driving comfort and driving trajectories.
[0036] Autonomous and semi-autonomous vehicles are able to sense their environment and navigate based on the sensed environment. Such vehicles use multiple types of sensing devices to sense their environment, such as radar, lidar, image sensors, camera devices, etc.
[0037] The trajectory planning or generation of an autonomous vehicle can be regarded as a real-time plan for the vehicle to transition from one feasible state to another, satisfying the vehicle's limitations based on vehicle dynamics and being constrained by the navigation lane boundaries and traffic rules, while avoiding obstacles including other road users as well as ground unevenness and ditches.
[0038] The use of classifiers such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs) is considered a complex data-driven learning model, which is usually implemented as a black box from the user's perspective. In other words, there is no clear explanation of the actual workings of the internal algorithms that perform complex functions. In the field of autonomous or semi-autonomous driving, this situation becomes more prevalent when developers or customers are unclear about the basis for a particular prediction that is about to occur or is occurring, or indeed whether it is correct for a particular situation.
[0039] A performance metric that quantifies the success of a classifier learning system is the accuracy of its predictions. The accuracy metric quantifies how well the classifier predicts relative to its learned training set. This metric is usually not sufficient for users to understand why the classifier predicted a certain output. This understanding by users is important when the learning system is implemented in a real product and in actual execution. Users often need or want to understand the rationale of a certain learning system used in the decision-making, and in the case of customers, the presentation of this rationale can enhance the user experience and provide insights into the automated behavior of automotive systems for technicians.
[0040] In various exemplary embodiments, the present disclosure describes an architecture and method for interpreting each of the predictions of a classifier in an understandable user format.
[0041] In various exemplary embodiments, the present disclosure describes methods for successfully extending a given classifier, training and evaluating it to produce an explanation for its original predictions. The present disclosure describes the benefits of such semantic explanations in two domains: comfort driving and driving trajectories. In the first, the present disclosure explains why the classifier predicts that a person will feel comfortable or not given a particular driving style. For example, the semantic reasons provided by an interpretable classifier may include traffic congestion, pedestrians, cyclists, and jaywalkers. In the second domain, an interpretable classifier can explain why it is predicted that a vehicle will change lanes or maintain its current lane or brake. As an example, these explanations can include other vehicles cutting in, faster or slower traffic in adjacent lanes, and traffic buildup in the same lane causing deceleration. In this domain, it can also be demonstrated that various explanations also include contrastive reasons. This is because it is not always possible to apply one semantic reason to an instance while another explanation should not be used or is not applicable. For example, the reason for a vehicle to decelerate may be an increase in traffic volume rather than because a vehicle has cut into the lane ahead.
[0042] In various exemplary embodiments, the present disclosure describes methods, systems, and devices for computing a set of semantic categories for a given domain, including a desired level of abstraction of the categories; computing the best layer in the original classifier meaning, i.e., the layer of the original classifier that retains most of the semantics corresponding to the total input set; training an interpretable classifier network while keeping the original learning model weights unchanged; testing the interpretable classifier network and quantifying its interpretation accuracy, and determining and computing an interpretable user format of the interpretation for each prediction of the original classifier.
[0043] In various exemplary embodiments, the present disclosure describes systems, methods, and devices for computing a semantic category language; computing layers in an original classifier that retain most of the semantics related to the semantic category language and a sample data set; computing the semantic accuracy of each layer; improving the semantic language to a higher-level language through abstraction, synonyms, etc.; expanding the learning architecture; training new edges in the expansion that include an interpretable classifier; computing the interpretation accuracy of the interpretation output for each determined semantic category; and formatting an interpretable user interpretation.
[0044] Figure 1 A block diagram depicting an exemplary vehicle 10 is shown, which may include a processor 44 that implements an interpretable learning system 100. Generally, input feature data is received by the interpretable learning system (or simply "system") 100. The system 100 determines an interpretable output based in part on the received feature data.
[0045] As Figure 1 shown, vehicle 10 generally includes a chassis 12, a body 14, front wheels 16, and rear wheels 18. The body 14 is disposed on the chassis 12 and substantially encloses the components of the vehicle 10. The body 14 and the chassis 12 may together form a frame. The wheels 16 - 18 are rotatably coupled to the chassis 12 near respective corners of the body 14. In the illustrated embodiment, vehicle 10 is described as a passenger vehicle, but it should be understood that any other vehicle may also be used, including motorcycles, trucks, sport utility vehicles (SUVs), recreational vehicles (RVs), ships, airplanes, etc. Although the present disclosure is depicted in vehicle 10, it is contemplated that the provided methods are not limited to transportation systems or the transportation industry, but rather and are applicable to any service, device, or application that implements a CNN type learning model. In other words, it is believed that the methods, systems, and devices described for the interpretable learning system have wide applicability in a variety of different fields and applications.
[0046] As shown in the figure, vehicle 10 generally includes a propulsion system 20, a transmission system 22, a steering system 24, a braking system 26, a sensor system 28, an actuator system 30, at least one data storage device 32, at least one controller 34, and a communication system 36. In this example, the propulsion system 20 may include an electric motor such as a permanent magnet (PM) motor, and other electrical and non-electrical devices are also applicable. The transmission system 22 is configured to transfer power from the propulsion system 20 to the wheels 16 and 18 according to an optional speed ratio.
[0047] The braking system 26 is configured to provide braking torque to the wheels 16 and 18. In various exemplary embodiments, the braking system 26 may include friction braking, wire braking, a regenerative braking system such as an electric motor, and / or other suitable braking systems.
[0048] The steering system 24 affects the position of the wheels 16 and / or 18. Although depicted for illustrative purposes as including a steering wheel 25, in some exemplary embodiments contemplated within the scope of the present invention, the steering system 24 may not include a steering wheel.
[0049] The sensor system 28 includes one or more sensing devices 40a - 40n that sense observable conditions of the external environment and / or the internal environment of the vehicle 10 and generate sensor data related thereto.
[0050] The actuator system 30 includes one or more actuator devices 42a - 42n that control one or more vehicle features, such as but not limited to the propulsion system 20, the transmission system 22, the steering system 24, and the braking system 26. In various exemplary embodiments, the vehicle 10 may further include Figure 1 internal and / or external vehicle features not shown, such as various door, trunk, and cabin features, such as ventilation, music, lighting, touch screen display components, etc.
[0051] The data storage device 32 stores data that can be used to control the vehicle 10. In various exemplary embodiments, the data storage device 32 or a similar system may be located on the vehicle (in the vehicle 10) or may be remotely located in the cloud, on a server, or on a personal device (i.e., a smartphone, a tablet, etc.). The data storage device 32 may be part of the controller 34, separate from the controller 34, or part of the controller 34 and part of a separate system.
[0052] The controller 34 includes at least one processor 44 (integrated with or connected to the system 100) and a computer-readable storage device or medium 46. The processor 44 can be any custom or commercially available processor, central processing unit (CPU), graphics processing unit (GPU), application-specific integrated circuit (ASIC) (such as a custom ASIC implementing a neural network), field-programmable gate array (FPGA), a secondary processor among multiple processors associated with the controller 34, a semiconductor-based microprocessor (in the form of a microchip or chipset), any combination thereof, or any device commonly used to execute instructions. For example, the computer-readable storage device or medium 46 can include volatile and non-volatile storage in read-only memory (ROM), random access memory (RAM), and keep-alive memory (KAM). KAM is persistent or non-volatile memory that can be used to store various operating variables when the processor 44 is powered down. The computer-readable storage device or medium 46 can be implemented using any of many known storage devices, such as PROM (programmable read-only memory), EPROM (electrical PROM), EEPROM (electrically erasable PROM), flash memory, or any other electrical, magnetic, optical, or combination storage device capable of storing data, some of which represent executable instructions used by the controller 34 in controlling the vehicle 10.
[0053] The instructions can include one or more individual programs, each program including an ordered list of executable instructions for implementing a logical function. When executed by the processor 44, the instructions receive and process signals (such as sensor data) from the sensor system 28, execute logic, calculations, methods, and / or algorithms for automatically controlling components of the vehicle 10, and generate control signals transmitted to the actuator system 30 to automatically control components of the vehicle 10. Although only one controller 34 is shown in Figure 1 an embodiment of the vehicle 10 can include any number of controllers 34 that communicate via any suitable communication medium or combination of communication media and cooperate to process sensor signals, execute logic, calculations, methods, and / or algorithms, and generate control signals to automatically control features of the vehicle 10.
[0054] As an example, the system 100 can include any number of additional sub-modules embedded within the controller 34, which can be combined and / or further divided to similarly implement the systems and methods described herein. Additionally, inputs can be received from the sensor system 28, received from other control modules (not shown) associated with the vehicle 10, and / or determined / modeled by other sub-modules (not shown) within Figure 1 the controller 34 of the system 100. Further, the inputs may also be pre-processed, such as subsampling, noise reduction, normalization, feature extraction, missing data reduction, etc.
[0055] Figure 2 is a functional diagram showing an exemplary solution 200 using a class branch or node of an interpretable learning system according to an embodiment. In Figure 2 it, the functional diagram includes a multi-layer set of L1 (215), L2 (220),... L n (230) of an exemplary raw classifier 205 that receives input features. The input features to each layer are weighted according to the configuration provided by the raw classifier 205. An exemplary (or designated) layer L* (225) is selected from the multi-layer set by a learning model based on a computational solution using a semantic function to determine the layer containing the most available semantic information from each available layer L1 (215), L2 (220),... L n (230) of the layer set of the raw classifier 205. Since the determined or designated layer L* (225) contains the most semantic information, it is determined that this layer is the "best" candidate to act as an anchor for connection to an extended interpretable classifier (i.e., class branch 210). W0 is the initial (i.e., first) semantic language that is computed, and "W" is the word or class in the language that is learned to explain the prediction output Oc. Thus, a set of semantic classes is computed based on the data of the input features to the raw classifier 205. The class branch 210 extends the raw classifier 205 to extract semantic classes related to the learned prediction from the received semantic data samples and serves as the basis for an output explanation that is formatted in a user-understandable manner by an explanation maker 245. The interpretable classifier learns from the dataset why a certain prediction is output for a certain input feature. The explanation comes from the semantic classes. More simply put, for a given classifier (C), an explanation of why the classifier outputs a specific output (Oc) for a specific input (Ic) is constructed. Through the extension of the class branch 210 and further processing by another neural network, the extended CNN classifier (of the class branch 210) will constitute or construct an interpretable classifier (interpretable C): for each output Oc, interpretable C will output Oc with an explanation.
[0056] Figure 3A and 3B is an exemplary flowchart of a method 300 that can be performed by an Figure 2 interpretable learning system 200 according to an embodiment. In various exemplary embodiments, the exemplary method 300 includes various tasks or steps that enable a certain raw classifier 205 ( Figure 2)It can receive inputs in the form of [x1, x2, …, xn] in order to output classes related to each input. That is to say, the interpretation of the classes is predictive and depends on semantic functions f1, f2, …, fk, which calculate f1(x1, x2, …, xn), f2(x1, x2…, xn), …, fk(x1, x2, …, xn) for the parameters x1, …, xn in the dataset. The semantic functions f1(), …, fk() are the interpretations related to the predictions made for a given classifier or the original classifier. For each sample of semantic data in the dataset, the calculations of the functions are performed on the parameters x1, …, xn in the dataset.
[0057] By extending the architecture of a given classifier network, it is possible to train the extended connections of a new classifier network (i.e., another neural network), and the new classifier network can learn what function f() corresponds to the data input x1…xn and the class prediction output of the original classifier for the input x1…xn (this training may result in more than one function as a possible explanation and also includes why a certain function cannot explain the prediction). When evaluating the extended connections of the new classifier, the neural network needs to evaluate the output produced in the output interpretation in a human - understandable format.
[0058] In tasks 310, 320, and 330 of the exemplary method 300, the explainable learning system 200 defines a semantic language. The semantic class language is calculated based on the data in the input set of the existing classifier. This data is not used in the original training of the classifier. For example, in task 310, a set of functions for why an autonomous vehicle brakes, changes lanes, or continues without changing (f1, …, f6) are defined as the set of functions that are possible explanations: f1 = open road, f2 = left cut, f3 = right cut, f4 = deceleration, f5 = left lane is faster, f6 = right lane is faster.
[0059] In task 320, for a given dataset [x1, …, xn], the explainable learning system 200 calculates each of the 6 functions f1…f6. For example, f5(autonomous - lane, autonomous - speed, autonomous - in - front - lane, autonomous - in - front - speed, otherV1 - speed, otherV1 - lane, …, otherVk - speed, otherVk - lane).
[0060] In an exemplary case, the lane IDs are encoded such that the absolute difference in IDs for consecutive lanes is 1, and the lane to the left always has a smaller ID. The example calculation code that can be used in the semantic function is as follows:
[0061] In 325, calculate f5()
[0062] For all vehicles, otherVi:
[0063] If other Vi-lane = autonomous-lane-1, then add other Vi(speed, lane) to the list of other vehicles
[0064] For all vehicles in the list of other vehicles
[0065] And speed += other Vi in the list of other vehicles
[0066] Average left-lane speed = sum of speeds / size of (list of other vehicles)
[0067] If (average left-lane speed > autonomous-ahead speed) then
[0068] Return left-lane-faster = 1
[0069] Otherwise return left-lane-faster = 0
[0070] Therefore, at 327, calculate f1(), f2(), f3(), f4() and f6() according to the semantics of the f function.
[0071] In each exemplary embodiment, the elements of the language will be regarded as possible interpretations that can be used in the calculation to interpret the output of the classifier. In this language, the interpretable learning system 200 defines W0 as including elements in the context that can be recognized by the user to affect or change the classification output. For example, in a driving environment, W0 may be defined as including "accident", "traffic congestion", "pedestrian crossing", "hazard", etc. These are semantic structures that the user may use to describe the driving situation when trying to interpret a certain driving behavior.
[0072] In each exemplary embodiment of tasks 310 to 330, in order to construct semantic tags, the interpretable learning system 200 gives all or approximately all samples in the classifier dataset additional category tags obtained from the semantic language W0. In various embodiments, the tagging can be done manually or by calculation.
[0073] At task 330, the interpretable learning system 200 improves the semantic functions f1() to f6(). For example, in some domains, the interpretable learning system 200 realizes that a new function g() can be calculated from the f function after calculating the semantic accuracy. For example, if f1() = a pedestrian crossing the road and f2() = an animal crossing the road, then the interpretable learning system can define a new function g1(f1, f2) = VRU (vulnerable road user) is crossing the road = f1() or f2().
[0074] In task 340, the explainable learning system 200 calculates or designates (semantically) the best layer. To extend a given classifier or the original classifier to output an explanation for each of its prediction outputs, it is necessary to calculate the layers in the original classifier from which a layer will be designated for extending the learning architecture. The designated layer is denoted as layer L*. The purpose of the explainable learning system 200 is to determine the layer that contains the most semantic information or data. The original classifier does not consider which of its layers contains the most semantic information when performing classification. To calculate the designated layer L*, the centroid of each layer L in the original classifier must first be calculated. It should also be noted that the original classifier can be a CNN classifier. The centroid is calculated by the original classifier for each category w in W0 (the defined semantic language) and each layer L in the classifier on a given training set. Each centroid is referenced by the following function where the activation of layer L for sample x is a vector that includes the activations of all nodes in that layer when x is the sample to be evaluated:
[0075]
[0076] Next, in 340, given a test feature data set that has been labeled with their semantic categories, the explainable learning system 200 calculates the semantic accuracy for each layer L in the original classifier. The semantic accuracy metric (calculated for any layer L and semantic category Wi) is the average of the f L (W i ) semantic prediction successes for all samples in the test set. This means that in a given layer L, the distance of the activated samples is closest to the calculated centroid. Since the test feature data samples are also labeled with semantic labels, the explainable learning system 200 can also calculate how many samples are closest to their (semantic) corresponding centroid. The explainable learning system 200 only calculates the closest value as the value with the minimum distance. In various exemplary alternative embodiments, the explainable learning system 200 can be configured to implement different distance metrics, and the threshold affects the determination of the semantic accuracy or the quality of the metric. The semantic accuracy of the metric is as follows:
[0077]
[0078]
[0079] Then, the explainable learning system 200 finds the designated layer L* for the semantic language W0 by adopting or selecting the layer L that has the maximum value of semantic accuracy.
[0080] In Task 340, improve semantic functions f1() to f6(). For example, in some domains, a new function g can be calculated from the f functions. For example, if f1() = a pedestrian is crossing the road and f2() = an animal is crossing the road, then define a new function g1(f1,f2) = a VRU is crossing the road = f1() or f2().
[0081] In various exemplary embodiments, and in Task 330, the explainable learning system 200 is configured to improve the semantic language W0. A new language W1 can be defined as an improvement to the semantic language defined in Task 1. For example, for some complex domains, it may be desirable to output a simpler explanation to the user, which may include using abstractions, synonyms, or related word packs.
[0082] In Task 350, the explainable learning system 200 constructs an explainable classifier and trains a new set of connections within the explainable classifier to create a new neural network.
[0083] In various embodiments, the explainable learning system 200 is configured to train only the new connections or inputs received from the original classifier, which have been formed in the extended architecture through the extension of the class branches. The explainable learning system 200 in Task 350 extends, creates, or constructs an extension including the explainable classifier by creating step architecture changes when linking the original classifier.
[0084] In an exemplary embodiment, at 350, the explainable learning system 200 trains a new explainable classifier in two new nodes by creating a first node of a "semantic branch" and a second node of an "interpreter" (see Figure 2 ). The semantic branch node is connected to the classifier layer L* (which has been calculated in Task 4). The interpreter is connected to the semantic branch and outputs the actual explanation when the original classifier outputs a predicted class. The explainable learning system 200 trains the explainable classifier while the weights that have been learned in the original classifier are frozen or kept static. In other words, the explainable learning system 200 is only configured to require additional training and learning for the new set of weights of the semantic branch in order to capture the semantic category closest to the predicted class.
[0085] In Task 360, the explainable learning system 200 calculates the explanation accuracy as a measure of the accuracy of the explanation output for each determined semantic class. For example: for a given data sample x1…xn, the explainable learning system 200 calculates the explanation: the predicted autonomous vehicle makes a left lane change because the lane on its left is faster, rather than because another vehicle is cutting in. The validation in Task 5 only samples new samples of the dataset, rather than the generated samples used in the training set.
[0086] The interpretable learning system 200 implements simple interpreter logic in task 370: assuming a human - understandable syntax template, where the predicted class and the calculated explanation are output in the corresponding syntactic positions of the phrase: "Because {explanation output}, the output is {Oc}". It demonstrates the explanations successfully learned in two autonomous driving experiments: driving comfort and driving trajectory.
[0087] Figure 4A 、 4B 4C, 4D, 4E, 4F, and 4G are exemplary flowcharts that illustrate the calculations of a specified layer in the original classifier. The specified layer contains most of the semantic data of the language with respect to semantic classes, configures a sample data set with optimized calculation configurations, and, according to the embodiments, calculates the semantic accuracy rate of each layer of the interpretable learning system.
[0088] In various exemplary embodiments, Figures 4A to 4F In the process of describing the interpretable learning system 200, an exemplary set of assumptions is implemented. For example, the interpretable learning system 200 assumes that the user's vehicle is surrounded by 3 other vehicles identified as A, B, and C based on a data set including approximately 100 samples. In each of the 100 samples, the interpretable learning system 200 assumes that each sample is configured at time t = [my - vehicle - position, my - vehicle - speed, A - position, A - speed, B - position, B - speed, C - position, C - speed]. Among the 100 samples in this example, 40 samples include group 1 (for the training step), and another 40 samples include group 2 (for the training step). The remaining 20 samples are divided into group 3 (7 samples) and group 4 (13 samples). Both group 3 and 4 are used for the test step. Additionally, for each sample, the following sample assumptions 1 to 4 are also made as follows: (1) The interpretable learning system enables the assumptions of 40 samples (group 1) out of 100 samples to be marked with "Another vehicle cuts in front of me"; (2) The interpretable learning system assumes that another 40 samples (group 2) out of 100 samples can be marked with "My road is clear"; (3) Among the remaining 20 samples (group 3), the interpretable learning system assumes that 7 samples can be marked w / "Another vehicle cuts in front of me"; (4) Among the remaining 20 samples (group 4), the interpretable learning system assumes that 13 samples can be marked w / "My road is clear". Next, in Figure 4A it is assumed that the given extended classifier is configured with only 5 layers, namely L1(410), L2(420), L3(430), L4(440), and L5(450), to process each sample and generate a predicted class for each sample.
[0089] Next, in Figure 4B and 4CIn this case, the process of processing 100 samples by the interpretable learning system is illustrated by initially running (i.e., processing) the first 40 samples of Group 1 through the neural network. Next, the interpretable learning system monitors the activations in the nodes contained in each layer of the neural network to determine the set of nodes triggered in each execution run of each sample. For example, in Figure 4B , "Sample 1" is sent and processed at L1, where the observed set of node activations represented by the calculations "0.2, 0.3, 0.7, and 0.1" at L1 is shown, and so on, as shown in the subsequent successive layers L2 - L5 that generate the predicted class for Sample 1. Similarly, in Figure 4C , "Sample 2" is processed in L1 and includes another set of node activations represented by the calculations "0.2, 0.3, 0.7, and 0.1" in L1, and so on in the subsequent successive layers L2 - L5 that generate the predicted class for Sample 2. Additionally, similarly in Figure 4D , "Sample 3" is processed in L1 and includes another set of node activations represented by the calculations "0.2, 0.3, 0.7, and 0.1" in L1, and so on in the subsequent successive layers L2 - L5 that generate the predicted class for Sample 3.
[0090] Next, in Figure 4G , the interpretable learning system generates and retains an average set of numbers based on the calculations in each layer to create a trained classification model, where the learned semantic class represents "Another vehicle cuts in front of me" (i.e., for the class related to approximately all 40 samples of each group). That is, the 5 - vector averages are represented as follows:
[0091] X1 = [average of the first node in the first layer, average of the second node in the first layer, …, average of the fourth node in the first layer]
[0092] X2 = [average of the first node in the second layer, average of the second node in the second layer, …, average of the fourth node in the second layer]
[0093] X3 = [average of the first node in the third layer, average of the second node in the third layer, …, average of the fourth node in the third layer]
[0094] X4 = [average of the first node in the fourth layer, average of the second node in the fourth layer, …, average of the fourth node in the fourth layer]
[0095] X5 = [average of the first node in the fifth layer, average of the second node in the fifth layer, …, average of the fourth node in the fifth layer]
[0096] As Figure 4E and 4F shown, the processing steps in Figures 4A - 4D are repeated for all 40 samples.
[0097] Then, as Figure 4G shown, the average in each layer for "My path is open" (which is the category for all 40 samples in group 2) is represented by averaging the Y1 to Y5 sets with new vectors as follows:
[0098] Y1 = [average of the first node in the first layer, average of the second node in the first layer, …, average of the fourth node in the first layer]
[0099] Y2 = [average of the first node in the second layer, average of the second node in the second layer, …, average of the fourth node in the second layer]
[0100] Y3 = [average of the first node in the third layer, average of the second node in the third layer, …, average of the fourth node in the third layer]
[0101] Y4 = [average of the first node in the fourth layer, average of the second node in the fourth layer, …, average of the fourth node in the fourth layer]
[0102] Y5 = [average of the first node in the fifth layer, average of the second node in the fifth layer, …, average of the fourth node in the fifth layer]
[0103] Only groups 3 and 4 in the test set are scored. Based on processing the training set containing groups 1 and 2, the average for each layer and each semantic category is calculated. The distance is calculated on the test set containing groups 3 and 4. A layer is scored only when the test sample has the closest activation distance to the average of that layer for a specific semantic category.
[0104] For example, for layer L1: If the sample is one of the 7 cases in the test set (group 3) among the remaining 20 samples (given the category "Another vehicle cuts in front of me"), and the sample causes the activation of the node in layer 1 that is closest to X1 (instead of Y1), the interpretable learning system assigns a score of 1 to layer L1; otherwise, a score of 0 is given to layer L1. If the sample is one of the 13 samples (in group 4) among the 20 samples in the test set (given the category "Open road"), and the sample causes the activation in layer L1 that is closest to Y1 (instead of X1), the interpretable learning system assigns a score of 1 to layer L1; otherwise, a score of 0 is given to layer L1. Then, based on all the calculations performed for each of the 20 samples, the final score for layer L1 is determined by the sum of the 1 scores. The selected layer L* is the layer that has received the maximum total score (among the layers L1, L2, L3, L4, and L5).
[0105] Figure 5 is another flowchart of the interpretable learning system according to the embodiment. Figure 5 The flowchart of Figures 3A - 3B is a higher-level flowchart than Figure 5In it, the interpretable learning system 200 configures the training of samples of the semantic classification model of the interpretable learning system 200 in a single step of Task 2. In short, in Figure 5 In the flowchart 500 of, the interpretable learning system defines the W0 semantic language by performing the following steps to perform the tasks initially at Tasks 510 and 520: (1) At Task 510, the interpretable learning system defines the set of semantic categories W0 as possible interpretation outputs (e.g., elements in the environment that affect the classification output); (2) At Task 520, labels are calculated for each training sample with the categories in W0.
[0106] Next, at Task 530, the interpretable learning system 200 performs the calculation of L* only for the calculation of the semantic categories of each sample in the training set. The steps are: (1) For each layer L in C, for each category w in W0: For a given training set, calculate the centroid (average) of the activations in layer L through a function:
[0107] And
[0108] (2) For each layer L in C, calculate the semantic accuracy rate,
[0109] Given a test set labeled with semantic ground truth (W):
[0110]
[0111]
[0112] Find the layer in the classifier that best fits W0: L* = argmax(semantic accuracy rate(L)).
[0113] At Task 540, the interpretable learning system 200 determines W1 from WO for possible abstractions. For example, possible extensions can be implemented to define new categories W1 based on W0 in the selected layer L* (e.g., W1 can be derived using groupings of words such as synonyms or datasets (i.e., ontology databases) and generalizations that can replace technical terms with more user-understandable language). It should be noted that there is no relationship between the improvement of W0 to be W1 and the best semantic layer L*.
[0114] At task 550, the interpretable learning system 200 is configured to build an interpretable classifier by extending a new branch to layer L* (i.e., the semantic branch) connected to the original classifier and train the model with the new connections of the classifier, which will make the learned output semantic classes related to the output predictions of the original classifier become the corresponding classes in W0 (i.e., the class semantic language set). The trained interpretable classifier will be trained using the set of frozen weights of the original classifier to which it is connected. The trained classifier will learn only the weights in the semantic branch based on the new training samples that have been labeled in task 520. The training step can be configured to learn the weights of the semantic branch so that W0 is configured with the output of the new branch (e.g., using cross-entropy loss or other appropriate loss functions). At task 560, an interpretation validation sample set is calculated based on the accuracy metric for the output. For example, an "interpreter" is added to the labels of the new classifier neural network. The interpretable learning system can utilize the test samples to calculate the interpretation accuracy of each interpretation output for the determined semantic classes. At task 570, the interpretation = "because {the output of the semantic branch of interpretable C}, the output is {Oc}". The interpretable learning system trains the interpretable classifier while freezing the weights that have been learned in the given classifier C. Therefore, it only learns the new weights of the semantic branch to capture the semantic classes closest to the predicted classes. The simple interpreter logic assumes a human-readable syntax template where the predicted class and the calculated interpretation are output in the corresponding syntactic positions in the phrase: "because {interpretation output}, the output is {Oc}."
[0115] Figure 6 is an example diagram of the interpretable driving comfort prediction of the interpretable learning system 200 according to an embodiment. The CNN used by the interpretable learning system 200 can be configured to predict the comfort or discomfort of the participants during a stimulated automated (i.e., simulated) ride. In an exemplary embodiment, the CNN can be trained and evaluated on a study of over 100 participants, ultimately obtaining approximately 117K data points (for reference data and the original classifier, see Adaptive Driving Agent in the Proceedings of the 8th International Conference on Human Agent Interaction, Goldman et al 2020). By implementing Figure 5 the processing flow (i.e., the 5 tasks listed), an interpretable CNN can be constructed, and it can output human-readable interpretations for comfort prediction.
[0116] In Figure 6 the interpretable learning system 600 includes an input consisting of 24 features 610 related to the driving environment and vehicle dynamics to the original classifier (Oc) 605, and the original classifier is processed in weight layers L1 to L n as previously in Figures 2 - 5As described in. The expansion of CNN C makes it interpretable (Interpretable C 615), and for each output of Oc 605, Interpretable C 615 will interpret and output Oc 605.
[0117] Select or specify that layer L*635 is connected to the class branch 620, and its structure is sent to the trained text model w∈W0 = {pedestrian, traffic congestion, bicycle, open road} of the interpreter 625. The output of the original classifier 605 will indicate the predicted driver discomfort. The interpreter 625 will generate an explanation = the user feels uncomfortable due to traffic congestion. The designated L* layer discovery 640 for each layer indicates the calculated semantic accuracy.
[0118] In the following exemplary table are example results of an interpretable learning system according to an embodiment. In the following exemplary table, the example domain 1 of the interpretable CNN for driving comfort prediction outputs an interpretable output from the interpreter.
[0119]
[0120]
[0121] The interpretable output for test sample 1 is the phrase "The user is uncomfortable because of traffic". This phrase includes the explanatory predicate "because of traffic", which is the reason for the subject "The user is uncomfortable". In other words, it is a two-part structure consisting of the subject of the condition and the predicate explaining the condition. Figure 6 The constructed interpretable classifier extension has a raw classifier with a semantic branch connected to layer L*. Train the interpretable classifier while freezing the weights in the original classifier and adding an interpreter at the top of the network. Use the sample explanation = "Because of {the output of the semantic branch of Interpretable C}, the output is {Oc}".
[0122] Figure 7 is an example diagram for vehicle trajectory prediction of an interpretable learning system according to an embodiment. In Figure 7 it, the interpretable learning system 700 includes an interpretable classifier 705, which includes an encoder 710, a convolutional de-pooling module 715, and a decoder 720 including a selected layer L*. (An example of the type of raw classifier implemented by the convolutional de-pooling module 715 is described in "Convolutional Social Pooling for Vehicle Trajectory Prediction" by Deo.N and Trivedi, M.M in the Proceedings of the CVPR Workshop 2018. It can be from the centroid distance calculation layer L* (at 735). Explanation = vehicle deceleration because there may be other vehicles cutting in. For the class branch, in this second domain, the interpretable learning system does not explicitly calculate L* because the original architecture of the classifier indicates that the specified layer is the only reasonable choice. Before the specified layer L*, the neural network of the interpretable classifier 705 does not contain information about the autonomous vehicle after the specified layer, and the neural network continues as a long short-term memory (LSTM) neural network (but it is not configured as a CNN). In this case, the expertise is sufficient to select the specified layer L* without performing calculations to determine it. Similarly in Figure 7 , the correspondence between Oc and the output explanation is as follows: Oc is a prediction of whether the vehicle is going to change lanes and / or brake.
[0123] Figure 8 is an exemplary flowchart of an interpretable CNN for vehicle trajectory prediction of an interpretable learning system according to an embodiment. In Figure 8 , in task 810, the interpretable learning system defines the W0 semantic language as W0 = {open road, left cut-in, right cut-in, deceleration, left lane faster, right lane faster}. In task 820, training samples are received and calculated with labeled samples having W0. In an exemplary embodiment, the calculation may include functions such as the adjacent lane being faster than the autonomous lane. Adjacent lane speed = average speed (current time) of all vehicles in the lane adjacent to the autonomous one within a range of 0 to 90 feet (only at positive distances in the lane, not behind). Autonomous lane speed is the speed of the previous vehicle (current time). Adjacent lane speed > autonomous lane speed (left / right). Next, a given classifier is implemented, and in task 830, the layer (vector of nodes) that has received the highest score in terms of its semantic content (i.e., L* = the trajectory encoding layer via human experts) is obtained. L* is calculated for the best layer to extract semantics. For example, in a given classifier, a class may not be implemented (e.g., all classes related to pedestrians, animals crossing the road in W0... can all be named "VRU"), so this new class is used in W1, and the explanation will refer to VRU instead of a specific class. Also, a selection task 840 can be implemented to determine W1 from w0 for possible abstraction. In task 850, an interpretable model of the interpretable classifier is constructed and trained with new connections. In task 860, a set of validation samples is received to calculate the explanation with an accuracy metric. In task 870, new samples are calculated with the output explanation. The trained semantic model may output the following explanations: "lane change to the left because the left lane is faster than mine", "brake because of a cut-in from the right", "brake because of lane deceleration", and "drive normally because the road is open".
[0124] In the following exemplary table of W0, the occurrence times in the training set and the occurrence times in the test set of the interpretable learning system according to an embodiment are shown.
[0125]
[0126] In the following exemplary table of W0, the interpretation accuracy and interpretation ROC AUC of an interpretable learning system according to an embodiment are shown.
[0127]
[0128]
[0129] The following is an exemplary event list for the second domain and the original CNN that predicts user comfort while experiencing a simulated autonomous driving ride (see Adaptive Driving Agent in Proceedings of the 8th International Conference on Human Agent Interaction, Goldman et al 2020). The semantic language W0 used by the interpretable classifier is built on these labels and includes 8 categories: traffic, bicycle, pedestrian, open road, cut-in, danger, dense traffic, and jaywalker.
[0130] In this first example, the focus is on showing the feasibility of the method for each category, and it can be noted that these categories are a superset of the events already present in the input features of the comfort prediction CNN. The interpretable learning system can enable manual annotation of the language to label the training dataset in the interpretable CNN in the above table.
[0131] a. Approaching traffic congestion
[0132] b. Danger on the road
[0133] c. Driving in dense traffic
[0134] d. Jaywalker
[0135] e. Pedestrian crossing the road
[0136] f. Other vehicle merging into our lane
[0137] g. Cyclist
[0138] h. Driving next to a cyclist
[0139] i. Traffic jam
[0140] j. Open road
[0141] k. Other cars passing us
[0142] In an interpretable learning system, according to an embodiment, an activation centroid is calculated for layer L within W0 of the interpretable learning system. The interpretable learning system defines a set of semantic classes W0, labels each training sample with a class in W0, calculates the activation centroids of layer L and W in W0, and calculates layer L in a classifier, with a maximum semantic accuracy of In an exemplary table, results of a test layer calculated by a classifier with a maximum semantic accuracy of an interpretable learning system according to an embodiment are shown. The centroid semantic accuracy is determined for W0 = {traffic, bicycle, pedestrian, open road, cut-in, danger, dense traffic, and jaywalker}.
[0143]
[0144] Layer L is calculated in the classifier with maximum semantic accuracy as follows: Given a test set labeled with semantic ground truth (W),
[0145]
[0146] To find the layer in the classifier that performs best for W0: L* = argmax(semantic accuracy(L)).
[0147] Use the centroid calculation of the interpretable learning system to select Figure 6 L* = max_pooling1d_2 in the test layer of. The semantic accuracy = 1 if the correct semantic class is predicted (ignoring comfort elements), otherwise 0. *The prediction means that the test sample is close enough to the centroid, which is calculated based on all the same training samples.
[0148] Figure 9 is an example diagram of a CNN for driving comfort prediction of an interpretable learning system according to an embodiment. In Figure 9 it, a constructible classifier extends the original classifier with a semantic branch connected to layer L*, and this semantic branch will predict the corresponding class in W0.
[0149] An exemplary table of the interpretation accuracy of an interpretable learning system according to an embodiment is shown below. The following table shows the metrics calculated for all layers (not only for L*, the optimal layer). The following table shows all the final results. In this domain, as mentioned before, certain semantic information is part of the input function. Therefore, the input obtains the highest accuracy score. However, it can be seen that even though the original classifier was trained to predict comfort, layer L* has successfully performed in interpreting semantics, and the interpretable learning system defines the prediction calculated by the original CNN in W0. Example interpretation outputs calculated by the interpretable CNN include the following: "The user is uncomfortable because of traffic" (when the remaining input settings result in a very slow speed due to traffic congestion), "The user is uncomfortable because of jaywalkers" (when there are jaywalkers and the vehicle speed is relatively high at this time), "The user is uncomfortable because of cut-ins" (the vehicle speed is not too high, but when another vehicle passes by, the distance from the vehicle in front is very short).
[0150]
[0151] In the interpretation accuracy of the semantic output (SO) of sample s and the original semantic ground truth (SGT):
[0152]
[0153]
[0154] The number of training samples is 72432, and the number of test samples is 18059. It can be seen that selecting max_pooling1d_2 is a good approximation of the actual optimal value. For layer L* = max_pooling1d_1 and layer L* = max_pooling1d_2, the accuracy is very close to 1.0. This may be due to the fact that the input contains the semantic classes themselves.
[0155] In the second example (see the paper "Convolutional Social Pooling for Vehicle Trajectory Prediction" by Deo N. and Trivedi M. in the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshop 2018), convolutional social pooling includes a system and inputs the 3-second (X, Y) trajectories of all vehicles, with a maximum distance from the autonomous vehicle of 1 lane and 90 feet. The output is the predicted (X, Y) position of the autonomous vehicle in the next 5 seconds, as well as the longitudinal maneuver of the autonomous vehicle = {normal driving, braking} and the lateral maneuver = {keep lane, left lane change, right lane change}.
[0156] Figure 9 is an exemplary FIG. 910 of an interpretable CNN for trajectory prediction of an interpretable learning system according to an embodiment. As Figure 9As shown, the interpretable classifier 920 composed of layers L1(930), L2(940), L*(950) and L n (960) defines a semantic category set of W0 = {open road, left cut-in, right cut-in, deceleration, faster in the left lane, faster in the right lane} of W0 (through the category branch 970). The function of each category for training the model (output by the interpreter 980) is: (1) Open road: (from the HV to the previous vehicle), distance interval > 90 feet; (2) Cut-in (left / right) (within 90 feet within the 4-second interval limit of the paper), the autonomous lane should remain unchanged, the autonomous recognition ID changes and the road is not open, left / right is based on the previous lane ID of the new front; (3) Deceleration: 2 frames have the same autonomous lane ID, and the autonomous vehicle does not stop, the autonomous previous ID is not empty, the autonomous previous ID is the same in 2 frames, and it does not stop, the time interval between the autonomous and the autonomous previous decreases. The gap decreases > 0.15 * the old gap, and the road is not open in 2 frames; (4) The adjacent lane is faster than the autonomous lane: The adjacent lane speed = the average speed of all vehicles in the lane next to the autonomous within a distance of 0 to 90 feet (only at positive distances in the lane, not behind) (current time), the autonomous lane speed is the speed of the previous vehicle (current time), and the adjacent lane speed > the autonomous lane speed (left / right).
[0157] The CNN of the interpretable learning system is configured to predict whether the vehicle is going to make a left lane change, a right lane change, maintain its current lane, and whether to brake. The interpretable learning system assumes that a CNN classifier is given in advance and predicts the vehicle behavior in the next 5 seconds based on the data in the previous 3 seconds. It applies Figure 5 five steps of the process in to explain the prediction. In this domain, the interpretable learning system defines a semantic language W0, whose semantic categories are not included in the dataset of the original classifier. For this reason, 6 semantic categories are calculated: open road, left cut-in, right cut-in, deceleration, faster in the left lane, faster in the right lane. Each such category requires calculations on the settings in our available data, but initially none of the categories are used for predictive learning. The interpretable learning system labels the training set of the interpretable CNN with these categories. In this domain, it is clear through applying user expertise that expanding the CNN in the trajectory encoding layer is a reasonable decision.
[0158] In various exemplary embodiments, the interpretable learning system uses an interpretable classifier to calculate the explanation. For example, it can explain that the predicted left lane change is because the left lane is faster than the current lane of the autonomous vehicle: "Brake because of a cut-in from the right", "Brake because the lane decelerates", "Drive normally because the road is open".
[0159] In this domain, the interpreter component that interprets the CNN architecture consists of a binary multi-label vector, which is associated with 1 when the corresponding class is true and 0 otherwise. The vector [open road / closed road, left cut-in / non-left cut-in, right cut-in / non-right cut-in, deceleration / non-deceleration, left lane faster / slower or equal to ours, right lane faster / slower or equal to ours]. Therefore, this architecture may lead to more complex interpretations represented by combinations of multiple classes. For example, it can extend the interpreter to provide contrastive interpretations, such as "braking because of a right cut-in rather than lane deceleration", "left lane change because the left lane is faster rather than because another vehicle cut in front of me", "left lane change because the left lane is faster, my lane is decelerating, and a vehicle is cutting in". The interpretable learning system can make multiple interpretations in one explanation or make contrastive interpretations.
[0160] In various exemplary embodiments, the present disclosure presents methods that are believed to be capable of extending a given classifier, training a classifier, and evaluating a classifier to produce an explanation of the original prediction. An interpretable learning system is implemented to illustrate the benefits of such semantic interpretations in two domains: comfortable driving and driving trajectories. In the first one, the interpretable learning system explains why the classifier predicts that the user will feel comfortable or not given a specific driving style. The semantic reasons provided by the interpretable classifier include traffic congestion, pedestrians, cyclists, and jaywalkers. In the second domain, the interpretable classifier explains why it predicts that the vehicle will change lanes or maintain its current lane or brake. These explanations include other vehicles cutting in, whether the traffic in the adjacent lane is faster or not, and the traffic accumulated in the same lane causing deceleration. In this domain, it is also shown that the explanations include contrastive reasons because it can be demonstrated that one semantic reason applies while the other does not (e.g., deceleration may be caused by increased traffic flow rather than because a vehicle is cutting in).
[0161] The present disclosure in various exemplary embodiments provides value at the intersection of learning models and model-based planning work. Interpreting the semantics of the black box can be passed to the model-based planner to help select higher-value actions. The results can also be considered relevant when supporting human users in understanding the high-level features computed by the AI system.
[0162] The various tasks performed in conjunction with the supervised learning and training of the interpretable learning model can be executed by software, hardware, firmware, or any combination thereof. In fact, Figures 1 - 9 part of the process can be executed by different elements of the described system.
[0163] It should be recognized that Figures 1 - 9 the process can include any number of additional or alternative tasks, Figures 1 - 9 the tasks shown need not be executed in the order illustrated, andFigures 1 - 9 The processing can be incorporated into a more comprehensive program or procedure with additional features not described in detail herein. Additionally, Figures 1 - 9 one or more of the tasks shown can be omitted from Figures 1 - 9 the embodiments of the process shown, so long as the intended overall functionality remains intact.
[0164] The foregoing detailed description is merely illustrative in nature and is not intended to limit the embodiments of the subject matter or the application and uses of such embodiments. As used herein, the word “exemplary” means “serving as an example, instance, or illustration.” Any embodiment described herein as exemplary is not necessarily to be construed as preferred or advantageous over other embodiments. Additionally, there is no intention to be bound by any theory presented in the prior art field, background, or detailed description.
[0165] Although at least two exemplary embodiments have been presented in the foregoing detailed description, it should be understood that there are a large number of variations. It should also be understood that one or more exemplary embodiments are merely examples and are not intended to limit in any way the scope, applicability, or configuration of the present disclosure. Instead, the foregoing detailed description will provide those skilled in the art with a convenient roadmap for implementing one or more exemplary embodiments.
[0166] It should be understood that various changes can be made to the functionality and arrangement of the elements without departing from the scope of the present disclosure as set forth in the appended claims and their legal equivalents.
Claims
1. A method for constructing an interpretable user output, the method comprising: Receiving input data of a feature set for processing at a neural network (NN) composed of multiple layers of an original classifier, wherein the original classifier has been frozen with a set of weights; Determining, by a processor of a vehicle, a semantic function to classify data samples into semantic categories; Determining, by a processor of the vehicle, a semantic accuracy level for each layer among the multiple layers of the original classifier within the neural network, wherein the original classifier is a trained model; For training samples of the determined semantic categories, calculating, by a processor of the vehicle, a representative vector of the average activation of layer nodes with multiple layers for evaluation by calculating the distance between samples in a test set and available layers of each semantic category; Calculating, by a processor of the vehicle, multiple test samples for each layer and each semantic category among the multiple semantic categories, which are closest to each other in each layer, to specify a layer with the highest score, where the highest score represents the best semantics in each test sample in the processed test sample set; Expanding, by a processor of the vehicle, the specified layer among the multiple layers to the NN through a category branch to extract semantic data samples of semantic content, and expanding an interpretable classifier in the expansion of the NN to define multiple semantic categories determined by the NN; Training, by a processor of the vehicle, a connection set of the interpretable classifier of the NN to calculate an output interpretation set based on at least one semantic category among the multiple semantic categories with an accuracy metric related to each output interpretation; And Comparing, by the trained interpretable classifier for each semantic category, the accuracy metric of each output interpretation based on the extracted semantic data samples to generate an output interpretation in a user - understandable format; Automatically controlling the movement of the vehicle based on the input data and the neural network through the processor; And Using the output interpretation, providing an output for a user of the vehicle via the processor, the output including an explanation of the control of the movement of the vehicle.
2. The method according to claim 1, further comprising: Calculating multiple interpretations as output; And Calculating a counterfactual interpretation of the composed interpretation, wherein the specified layer is the best semantic content layer.
3. The method according to claim 2, further comprising: Configuring an interpretable classifier with multiple layers, wherein each layer is configured with a set of weights; Labeling the extracted input data samples of the interpretable classifier with a defined set of semantic categories related to the best semantic content layer; Calculating an activation centroid for each layer among the multiple layers of the interpretable classifier; And Based on the activation centroid, calculating the layer of the interpretable classifier with the maximum semantic accuracy.
4. The method according to claim 3, further comprising: Calculating the maximum semantic accuracy for each semantic category in the set of semantic categories based on the semantic function and optionally a re - defined semantic function.
5. The method according to claim 4, further comprising: Specifying at least one layer among the multiple layers of the original classifier based on the calculated semantic accuracy by running input data samples through each layer of the original classifier to observe the activation count of the nodes included in each layer. For each semantic category, determine a set of distances between the activation of each node in the node layer for each sample and the average of the activations of the nodes in the corresponding node layer for all data samples in the training set; Repeat the determination step for each successive input data sample until the average vector value of the nodes in each layer is calculated for each received input data sample and each semantic category in each layer; In each layer of multiple layers of the original classifier, retain the average vector value of each node; In each layer of multiple layers, score the average of the nodes in the vector, where the average of the nodes is configured as a vector; Sum the set of vector value scores for each layer in the multiple layers received for all data samples in the dataset, where the vector values with scores in each layer are used to determine the layer with the maximum value; And Designate the layer with the maximum value in the multiple layers as the best semantic content layer.
6. The method according to claim 1, wherein The neural network includes a convolutional neural network (CNN), and the neural network includes an extension of the CNN.
7. The method according to claim 6, further comprising: By extending the original classifier, extend the interpretable classifier for driving comfort through an extension of the NN to extract semantic content via a designated layer to predict a semantic category corresponding to a set of features of vehicle dynamics and vehicle control.
8. The method according to claim 7, further comprising: While freezing the set of weights of the original classifier, train the set of connections of the NN extension of the interpretable classifier for driving comfort; And Add an interpreter to the output of the NN extension to be related to the data samples of the extracted semantic content.
9. The method according to claim 6, further comprising: By extending the original classifier, extend the interpretable classifier for trajectory prediction to extract semantic content via a designated layer to predict a semantic category related to a set of semantic features of sample data including convolutional social pooling.
10. A system for constructing an interpretable user output, the system comprising: A processor configured to receive input data to process a set of features at a neural network composed of multiple layers of an original classifier, where the original classifier has been frozen with a set of weights related to the features of the input data; The processor is configured to determine a semantic function to classify data samples into semantic categories, and determine the semantic accuracy level of each layer in the multiple layers of the original classifier within the neural network, where the original classifier is a trained model; The processor is configured to calculate a representative vector having the average activation of the nodes in each layer for a set of samples for training semantic categories, to be evaluated by calculating the distance between each sample in a test set composed of the set of samples and each semantic category in the available layers in the multiple layers; The processor is configured to calculate multiple test samples for each layer in the multiple layers and each semantic category, which are closest to each other in each layer in the multiple layers, to designate the layer with the highest score in the multiple layers, where the highest score represents the best semantics in each test sample, and the designated layer is the best semantic content layer; The processor is configured to extend the best semantic content layer to the neural network through category branches to extract a set of semantic data samples from the semantic content of the original classifier into the neural network, and use the extension to extend the interpretable classifier of the neural network to define multiple semantic categories; The processor is configured to train a new set of connections of the interpretable classifier of the neural network to calculate a set of output interpretations composed of interpretations based on at least one semantic category among the multiple semantic categories with an accuracy metric related to each output interpretation; and The processor is configured to compare the accuracy metric of each output interpretation for each semantic category among the multiple semantic categories by the trained interpretable classifier based on the extracted semantic data samples, and generate an output interpretation in a user - understandable format; The processor is configured to automatically control the movement of the vehicle based on the input data and the neural network; and The processor is configured to use the output interpretation to provide an output to the user of the vehicle, the output including an explanation of the control of the movement of the vehicle.
Citation Information
Patent Citations
Method and apparatus for preview-based vehicle lateral control
CN1974297A
CONTROL SYSTEMS, CONTROL METHODS AND CONTROLS FOR AN AUTONOMOUS VEHICLE
DE102019112038A1