System and method for creating interpretable potential representation of artificial intelligence model
By applying auxiliary loss function to fixed canonical functions during the training of neural network models, the problems of model interpretability and explanatory ability are solved, and more efficient and accurate model inference is achieved.
Patent Information
- Application Number
- CN202410420597.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-29
- Filing Date
- 2024-04-09
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively improve the interpretability and explanatory nature of neural network models, resulting in reduced complexity and computational efficiency of the model, affecting the accuracy of the model and user trust.
By applying auxiliary loss functions during training of neural network models to fixed canonical functions, reduce redundancy, and generate potential representations based on interpretable artificial intelligence.
A more intuitive and interpretable potential representation of the model is realized, reducing the consumption of computing resources and improving the inference efficiency and accuracy of the model.
Smart Images

Figure CN120068924A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and more particularly to methods, non-transitory computer-readable storage media, and computer-implemented systems for visualizing latent representations of neural network models. Background Art
[0002] With the continuous development of computing technology and vehicle technology, features related to automation have become more powerful and widely available, and can control vehicles in a wider variety of environments. For example, for automobiles, the Society of Automotive Engineers (SAE) has established a standard (J3016) that identifies six levels of driving automation from "no automation" to "full automation". The SAE standard defines level 0 as "no automation", where the human driver performs all aspects of the dynamic driving task full-time, even when enhanced by warning or intervention systems. Level 1 is defined as "driver assistance", where the vehicle controls steering or acceleration / deceleration (but not both) in at least some driving modes, allowing the operator to perform all remaining aspects of the dynamic driving task. Level 2 is defined as "partial automation", where the vehicle controls steering and acceleration / deceleration in at least some driving modes, allowing the operator to perform all remaining aspects of the dynamic driving task. Level 3 is defined as "conditional automation", where, for at least some driving modes, the automated driving system performs all aspects of the dynamic driving task, expecting the human driver to respond appropriately to an intervention request. Level 4 is defined as "high automation", where, only for specific conditions, the automated driving system performs all aspects of the dynamic driving task even if the human driver does not respond appropriately to an intervention request. The specific conditions for level 4 can be, for example, a specific type of road (e.g., highway) and / or a specific geographic area (e.g., a properly mapped geographically isolated large urban area). Finally, level 5 is defined as "full automation", where the vehicle is able to operate without operator input under all conditions.
[0003] Artificial intelligence (AI) and machine learning technologies have advanced significantly. At the same time, research on the underlying computations in neural networks and / or AI models (hereinafter referred to as models) has become increasingly important. Models such as multi-layer perceptrons (MLPs), convolutional neural networks (ConvNets), recurrent neural networks (RNNs), and transformers have gained wide recognition for their ability to handle complex tasks and deliver excellent performance. One of the challenges that has emerged with the popularity of models is how to gain an understanding of the internal workings of these models.
[0004] That is, the complexity of the models stems from their hierarchical architectures, which can include multiple interconnected nodes / neurons and weighted connections. Due to these complex structures, it is difficult to understand the computations within these models during the inference phase. Despite the complexity, understanding these computations can facilitate the interpretation and understanding of the models' inference processes, leading to an increased level of trust in the models and promoting their increased adoption across various domains. As a result, interest has grown in developing new techniques and methods that reveal the inner workings of these models, thus bridging the gap between their architectural complexity and underlying inference.
[0005] Meanwhile, the lack of interpretability and explainability of the models has also received increasing attention, which hinders users' understanding of the decision-making process and their ability to trust and understand the model outputs. This lack of model transparency also poses challenges in identifying biases or errors in the models' predictions. Therefore, efforts have been invested in transforming "black-box" models into interpretable "white-box" models to address these issues and enable a comprehensive understanding of their internal mechanisms.
[0006] However, with the increased focus on the models, the importance of complex models has become apparent, characterized by a vast number of interconnected neurons and nodes. Thus, the very large number of neurons within these complex models can lead to interference, for example, inadvertently including irrelevant or less important nodes / neurons when attempting to visualize or provide a latent representation of the model for interpretability and explainability. The over-inclusion of key neurons and their role in decision-making undermines the goal of improved interpretability and explainability. Additionally, most of these neurons typically contribute little to the models' inference processes, resulting in reduced computational efficiency, inefficient resource allocation, and increased energy consumption.
[0007] Furthermore, the lack of accessible and intuitive interpretability and explainability of the models hinders the gradual improvement of model accuracy. Here, accessible and intuitive interpretability and explainability can mean that the visualization of the latent representation of the model is computationally manageable from a machine's perspective and easily understandable from a human's perspective. In some cases, the lack of interpretability and explainability in the models can allow unnoticed errors and biases to persist, potentially causing significant errors and undermining user confidence. This is especially true in AI-dependent systems (such as autonomous driving systems where the reliability and fairness of model predictions are prioritized). Additionally, poor interpretability and explainability prevent users or developers from making post-training improvements to the models, thus hindering the efficient allocation of limited computational resources, impeding the effectiveness of the models in completing tasks and learning from errors, and ultimately affecting their overall accuracy and performance.
[0008] Given the above problems and the increasing integration of AI models in autonomous driving systems, it is important to develop algorithms and systems that enhance or improve the interpretability and explainability of these models. This will help model developers and users understand the decision-making process of the models, thus ultimately improving system performance. Additionally, it is important for these algorithms and systems to reduce computational costs and associated energy consumption on limited resources while enhancing inference accuracy across various neural network / AI model types. Summary of the Invention
[0009] The present disclosure provides a method, a non-transitory computer-readable storage medium, and a computer-implemented system for training a neural network model and visualizing a latent representation of the neural network model.
[0010] In a first aspect of the present disclosure, a method for training a neural network model based on a latent representation is provided, the latent representation including a human-interpretable variable / data representation required to perform a specified task. The method includes: obtaining input data; training a neural network model based on the obtained input data; fixing a normalization function during the training of the neural network model by applying an auxiliary loss function on latent activations for the specified task to minimize redundancy; and generating an explainable artificial intelligence-based representation of the latent representation of the neural network model based on applying the auxiliary loss function on the latent representation.
[0011] In another aspect of the present disclosure, a method for visualizing a latent representation of a neural network model is provided. The method includes: obtaining input data; applying a neural network model, the neural network model being trained based on the latent representation including an explainable artificial intelligence-based data representation required to perform a specified task and based on the obtained input data, wherein the neural network model has a fixed normalization function; and generating an explainable artificial intelligence-based representation.
[0012] In yet another aspect of the present disclosure, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores instructions thereon that, when executed by one or more processors, cause the one or more processors to perform operations. The operations include: obtaining input data; training a neural network model based on the obtained input data; fixing a normalization function during the training of the neural network model by applying an auxiliary loss function on latent activations for the specified task to minimize redundancy; and generating an explainable artificial intelligence-based representation of the latent representation of the neural network model based on applying the auxiliary loss function on the latent representation.
[0013] It should be understood that all combinations of the above-described concepts and additional concepts described in more detail herein are considered to be part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter that appear at the end of this disclosure are considered to be part of the subject matter disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] To more clearly illustrate embodiments of the present disclosure or related technologies, the following drawings that will be described in conjunction with exemplary embodiments are briefly introduced. Obviously, the drawings only reflect some embodiments of the present disclosure, which means that those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts. The arrows in the drawings indicate relationships, through which the component starting from the arrow is used to train / apply the component pointed to by the arrow. Embodiments of the present disclosure will be more fully understood and appreciated in light of the following detailed description in conjunction with the drawings, in which:
[0015] Figure 1 is a block diagram showing an example of an artificial intelligence (AI) model during training according to some embodiments of the present disclosure;
[0016] Figure 2 is a block diagram illustrating another example of a trained artificial intelligence (AI) model during training according to some embodiments of the present disclosure;
[0017] Figure 3 is a block diagram illustrating another example of an AI model during training according to some embodiments of the present disclosure;
[0018] Figure 4 is a block diagram illustrating an example of a trained AI model suitable for performing a method of visualizing a model in inference according to some embodiments of the present disclosure;
[0019] Figure 5 is a flowchart illustrating an exemplary process for generating a human-interpretable representation of a latent representation of a neural network model according to some embodiments of the present disclosure;
[0020] Figure 6 is a flowchart illustrating an exemplary process for obtaining an auxiliary loss function according to some embodiments of the present disclosure;
[0021] Figure 7 is a flowchart illustrating another exemplary process for obtaining an auxiliary loss function according to some embodiments of the present disclosure;
[0022] Figure 8 is a flowchart illustrating an exemplary process for visualizing a latent representation of a neural network according to some embodiments of the present disclosure;
[0023] Figure 9Adepicts a perspective view of a road from a driver's perspective in an exemplary autonomous driving scenario for a lane centering task according to some embodiments of the present disclosure;
[0024] Figure 9B depicts a top view of a road in an exemplary autonomous driving scenario for a lane centering task according to some embodiments of the present disclosure;
[0025] Figure 10A depicts a perspective view of a road from a driver's perspective in an exemplary autonomous driving scenario for a distance keeping task according to some embodiments of the present disclosure;
[0026] Figure 10B depicts a top view of a road in an exemplary autonomous driving scenario for a distance keeping task according to some embodiments of the present disclosure; and
[0027] Figure 11 illustrates an exemplary hardware and software environment for an autonomous vehicle according to some embodiments of the present disclosure.
[0028] It should be understood that for simplicity and clarity of illustration, the elements shown in the figures are not necessarily drawn to scale. For example, the dimensions of some elements may be exaggerated relative to other elements for clarity. Additionally, where considered appropriate, reference numerals may be repeated in the figures to indicate corresponding or similar elements. Detailed Description
[0029] Referring to the accompanying drawings, embodiments of the present disclosure are described in detail with technical problems, structural features, achieved purposes, and effects as follows. Specifically, the terms in the embodiments of the present disclosure are only used for the purpose of describing specific embodiments and do not limit the present disclosure. In the following detailed description, many specific details are set forth to provide a thorough understanding of the present invention. However, those skilled in the art will understand that the present invention can be practiced without these specific details. In other instances, well-known methods, procedures, and components are not described in detail so as not to obscure the present invention. The subject matter of the present invention is particularly pointed out and clearly claimed at the end of the specification. However, with regard to the organization and method of operation of the present invention and its purposes, features, and advantages, the present invention can be best understood by reference to the following detailed description when read in conjunction with the accompanying drawings. Since the illustrated embodiments of the present invention can be largely implemented using electronic components and circuits known to those skilled in the art, in order to understand and appreciate the underlying concepts of the present invention and in order not to obscure or distract from the teachings of the present invention, the details will not be explained to a greater extent than considered necessary as shown above. For example, the specification and / or the drawings may refer to a processor or a processing circuit. The processor may be a processing circuit. The processing circuit may be implemented as a central processing unit (CPU) and / or one or more other integrated circuits such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a full-custom integrated circuit, etc., or a combination of these integrated circuits.
[0030] The following specification and / or drawings may refer to an image or an image frame. An image is an example of a media unit. Any reference to an image may be applied to the media unit as necessary. The media unit may be an example of a sensed information unit (SIU). Any reference to the media unit may be applied to any type of natural signal, such as but not limited to signals generated by nature, signals representing human behavior, signals representing operations related to vehicle signals, geodetic signals, geophysical signals, text signals, digital signals, time-series signals, etc. Any reference to the media unit may be applied to the SIU. The SIU may be of any type and may be sensed by any type of sensor, such as a visual light camera. An audio sensor. A sensor that can sense infrared, radar imaging, ultrasound, electro-optical, radiographic, light detection and ranging (LIDAR), thermal sensors, passive sensors, active sensors, etc. Sensing may include generating samples (e.g., pixels, audio signals, etc.) representing the signals transmitted or otherwise arriving at the sensor. The SIU may have one or more images, one or more video clips, text information about one or more images, text describing motion information, etc.
[0031] Any combination of any modules or units listed in any of the figures, any part of the specification, and / or any of the claims can be provided. Any of the units and / or modules illustrated in the present application can be implemented in hardware and / or code, instructions, and / or commands stored in a non-transitory computer-readable medium, and can be included in a vehicle, outside a vehicle, in a mobile device, in a server, etc. The vehicle can be any type of vehicle, such as a ground transportation vehicle, an aerial vehicle, or a watercraft. The vehicle is also referred to as a self-driving vehicle. It should be understood that autonomous driving includes at least partial automatic (semi-automatic) driving of the vehicle, which includes all types of level 2 or higher levels defined in the SAE standard.
[0032] As used herein, a neural network or artificial intelligence (AI) model that can be used interchangeably throughout the present disclosure can be general or dedicated to a specific application scenario, e.g., decision-making, classification, prediction, etc. In particular, the model can be customized for conventional tasks related to autonomous driving. These tasks can be classified, for example, as perception, localization and mapping, planning and decision-making, and control. Perception tasks involve the accurate detection and recognition of objects and entities in the surrounding environment. This includes identifying and classifying pedestrians, vehicles, traffic signs, traffic lights, and other relevant objects. Localization tasks focus on determining the precise position of the vehicle in its surrounding environment, which involves using sensors and data to estimate the vehicle's position relative to known reference points or maps. On the other hand, mapping tasks involve creating and updating a representation of the surrounding environment. Together, localization and mapping enable the autonomous driving system to understand the precise position of the vehicle and navigate it effectively. Planning tasks involve generating a sequence of actions or a trajectory based on the vehicle's current position and desired destination. Decision-making tasks require analyzing the current driving situation and determining appropriate actions, such as changing lanes, accelerating, braking, or yielding. Together, planning and decision-making enable the autonomous driving system to navigate the vehicle in a safer and more efficient manner. Control tasks generally include executing the planned actions and adjusting the vehicle's dynamics to follow the desired trajectory. This includes controlling the steering, acceleration, and braking systems to maintain proper control and stability of the vehicle. Control tasks ensure that the physical response of the vehicle is consistent with the planned actions.
[0033] The present disclosure proposes a method for obtaining a human-interpretable latent representation within an AI model. As explained in more detail below, the human-interpretable latent representation can be obtained by learning or fixing a specific regularization function that minimizes redundancy. Subsequently, learning or fixing the specific regularization function can be achieved by introducing an auxiliary loss function on the latent activations during training, thereby effectively mandating a human-interpretable representation.
[0034] As a preliminary matter, for the interpretability and explainability of neural networks and / or AI models to be accessible, the hundreds, thousands, or even millions of neurons within the model are scaled down in order to obtain a simplified and relatively compact representation of the entire set of neurons in the model. To some extent, this can facilitate obtaining an intuitive understanding of which part of the model input is being looked at and thus encoded by each of the limited number of neurons in the relatively compact representation in order to accomplish a given task. By simplifying the representation of the neurons, an understanding of how the neural network responds to the model input under a given task can be achieved.
[0035] However, in an attempt to obtain an enhanced or improved interpretable or explainable latent representation of the model for a user or other AI-supported system, it is desirable to further reduce the number of active neurons. This means that the mapping relationship used as the output policy of the network within the policy head needs to be transformed from non-linear to linear as much as possible, thereby aiming to achieve the irreducibility of the network. This irreducibility remains significantly meaningful both in the field of computational science and in practical applications.
[0036] Reducing the number of active neurons helps simplify the computational process and enhance the interpretability and explainability of the model's predictions. By promoting linearity, the mapping relationship in the policy head allows for a more direct understanding of how the model arrives at its decisions. This shift towards linearization enables a clearer mapping between the input features and the output inference (e.g., prediction), thus facilitating an easier explanation and understanding of the system's behavior.
[0037] It has been recognized that in the context of neural networks and AI models, gauge invariance is a fundamental concept related to the irreducibility and interpretability of the model. Similar to the concept of fixing the gauge in physics, fixing the gauge function in the model can help minimize redundancy and obtain a more compact latent representation of the model, and in particular a more compact set of active neurons that play an essential and / or irreplaceable role in the inference of the model, thereby enhancing its interpretability and explainability.
[0038] As used herein, the term gauge invariance refers to the symmetric and redundant degrees of freedom present within the model. Taking a simple and straightforward example in electrostatics, the electric potential Φ can be due to the electric field The independence with respect to the choice of the constant C is determined by a transformation in which C is arbitrarily imposed (i.e., Φ -> Φ + C). In technical terms, under this transformation of the electric potential, the electric field remains gauge invariant. At this point, the constant C is an inherent redundancy within the electric field system. For some specific applications of the work to be performed by the electric field system on the input electrons or those electrons that depend on the characteristics of the electric field system, the constant C is reducible and does not play a substantial role in the useful work or practical utilization of the electric field. On the contrary, from the perspective of the machine, maintaining the constant C or even multiple alternative values for C can potentially introduce unnecessary computational overhead in the electric field simulation. Additionally, from the human perspective, it may cause unnecessary costs in terms of interpretation and / or understanding, which may divert attention from more relevant aspects.
[0039] Therefore, considering the concept of gauge invariance, this redundancy factor is minimized as much as possible during the application of the system. For example, this applies in the case of neural networks or AI models, especially during inference or more generally during the prediction process, the goal of which is to maximize the degree of reduction of the mapping relationship between the input and the output, thereby enhancing the interpretability or explainability of the model. One possible approach involves transforming the policy head of the model into a more linear one by fixing the gauge function (i.e., implementing gauge invariance), rather than retaining a large number of reducible and non-linear redundancy factors, such as abundant less relevant neurons and associated connection weights. Compared with the case where the gauge changes due to the reducible mapping relationship of the non-linearity within the policy head that carries multiple redundant elements of the model, transforming the policy head of the model into a more linear one through the fixing of the gauge function facilitates a clearer interpretation and understanding of the model.
[0040] The present disclosure proposes a method for training a neural network model. During the inference phase, it is desired to enhance the interpretability or explainability of the trained model. In addition to providing the model inference output for tasks such as autonomous driving, this method aims to provide an enhanced potential representation of the inference process of the model. To this end, by applying an auxiliary loss function to fix the gauge function of the model, the policy head that requires the decision of the model exhibits more linear characteristics and reduces the number of active neurons / nodes in the original potential representation of the model. In doing so, the redundancy factors involved in the inference process of the model are largely removed from the potential representation. Through this enhancement, the input-output mapping relationship of the model (typically embodied by the policy head of the model) is transformed into a more linear one than when the auxiliary loss function is not applied.
[0041] Additionally, applying an auxiliary loss function to achieve a learned or fixed canonical function can be used as a reference point, marker, or basis for interpreting and understanding the model. Further, the fixed canonical can further reduce the number of active neurons in a relatively compact representation of the entire set of neurons within the model. As the number of active neurons / nodes in the original latent representation of the model decreases and the redundant factors involved in the inference process of the model (as shown in the latent representation) are removed, the resulting model requires less time and fewer computational resources during the inference phase. Thus, the inference of the model not only becomes task-oriented but also more interpretable, explainable, and computationally efficient. This results in a more concise and intuitive latent representation of the model, which enhances its interpretability and explainability to humans and other AI-supported systems.
[0042] Accordingly, a user, model developer, or trained model can modify and / or fine-tune the network structure of a previously trained AI model for a given task based on the generated enhanced latent representation, e.g., deactivate or even remove those nodes that are less relevant to the model's decision-making due to a reduced number of countable neurons. The deactivation or removal of less relevant nodes saves computational resources, simplifies the model's decision-making process, and improves computational efficiency. Additionally, by using the enhanced latent representation generated by the methods and systems disclosed in the present application, a model developer or an AI-supported system can check whether it is necessary to modify the type, quantity, format, etc. of the model input to better facilitate the generation of correct model decisions. In this way, the enhanced latent representation can improve the accuracy of model inference for model inputs used in a given task and enhance the safety and reliability of using such an AI model in applications such as autonomous driving systems.
[0043] Furthermore, as the human interpretability of the AI model increases, users and / or trainers can provide more accurate feedback to the model as reference data. This higher-quality reference data reduces the total amount of data required for the model to achieve a task-specific objective. This means that the model requires less time and fewer computational resources to train its parameters to achieve a working model.
[0044] It should be noted that the flexibility of the disclosed method is a significant feature because it does not rely on the specific complexity or implementation details of the model. Therefore, it can be effectively applied to a wide range of models, regardless of their architecture, size or complexity. Scalability ensures its compatibility with various types of models (including neural networks, deep learning models, reinforcement learning models or any other form of machine learning algorithms). In general, the disclosed method provides a resource-efficient and scalable method that ensures compatibility across various types of models and then makes it valuable in practical applications by fixing the canonical function of the model using auxiliary losses while slowing down the increased computing resources used for model reasoning and avoiding dependence on model-specific details, thereby obtaining enhanced and direct insights into the model.
[0045] Reference is now made to the drawings, wherein like numerals refer to like parts throughout. Figure 1 A block diagram 1000 showing an example of an artificial intelligence (AI) model during training is illustrated according to some embodiments of the present disclosure. Figure 1 As shown, the AI model 1200 may include: a model backbone 1202, a mixing block 1204, a potential layer 1206, a plurality of neurons 1208 within the potential layer 1206, and a strategy head 1210.
[0046] The model backbone 1202 may constitute a foundational part of the AI model 1200, which may be responsible for initial data processing, for example. In some examples, the model backbone 1202 may include various layers and modules designed to extract and transform information carried in the model input, which is provided by the training data 1100 during the training phase. The model backbone 1202 captures, extracts, and classifies the necessary features and representations from a large number of model inputs (e.g., frontal images or videos of roads, or annotations of lateral acceleration, etc.), which are required for subsequent analysis and decision making within the AI model 1200. In an embodiment, the model backbone 1202 may be a convolutional neural network (CNN) that learns different features such as lines and curves of roads.
[0047] The hybrid block 1204 can, for example, integrate and combine information from different parts (e.g., layers) of the model backbone 1202. It enhances the overall representation of the input data by facilitating the exchange of information and feature fusion between model inputs. The hybrid block 1204 ensures the effective sharing and utilization of relevant information, improving the overall performance and accuracy of the AI model 1200. In an embodiment, the hybrid block 1204 can be a multilayer perceptron (MLP), which may include: a channel hybrid MLP that allows communication between different channels, and / or a token hybrid MLP that allows communication between different spatial locations. These layers can be interwoven (i.e., combined) to enable interaction of two types of inputs.
[0048] The latent layer 1206 can contribute to a simplified and relatively compressed representation of the model input (which can be, for example, the training data 11000 during the training phase), and such a representation can explicitly or implicitly include a summary of the key features of the model input (such as features related to lane boundaries). In some embodiments, the latent layer 1206 can be obtained by discarding duplicate or irrelevant elements (e.g., neurons or nodes) using different data representation and approximation techniques. This allows for less data to be transmitted without substantial loss and for a simplified version of the model to be transmitted instead of the huge ontology model. Thus, computational efficiency can be improved because less data needs to be processed and transmitted from one area to another. Moreover, model accuracy can be maintained with little loss.
[0049] The latent layer 1206 can include a plurality of neurons 1208, each neuron dedicated to or focused on capturing and processing specific input features or patterns for a given task. In some examples, the latent layer 1206 can serve as a relatively compact version of the neurons representing the entire set of neurons within the AI model. That is, the number of neurons in the latent layer 1206 is limited compared to the original number of hundreds, thousands, or even millions of neurons within the complex structure of the model 1200. Thus, the collective behavior of these neurons 1208 can contribute to the overall processing of the input data within the AI model 1200 to complete the task.
[0050] The policy head 1210 represents a component that develops policies and generates a final output or makes a decision based on the analysis of the processed input data. The policy head 1210 also provides a higher-level understanding of the model input. That is, the policy head 1210 indicates the action to be taken based on the state of the model 1200 and the detected surrounding environment. In an embodiment, the policy head 1210 can be a trainable AI model.
[0051] During the training phase, the AI model 1200 receives, processes model inputs from the training data 1100, and generates model outputs, such as the (multiple) actions to be taken under a given task, which outputs are to be compared with a set of ground truths within the loss function module 1300. Examples of model inputs can be image signals depicting a front view image of a road. However, one of ordinary skill in the art can understand that there can also be other suitable forms of training data, such as audio signals, text annotations, or combinations of audio signals and image signals (e.g., video streams) along with text annotations. In some embodiments, the model input can be raw data from one or more sensors of the same vehicle or separate vehicles. For example, the model input can be an image including red, green, and blue (RGB) values of pixels captured by a camera sensor. The model input can be raw SIU, processed SIU, text information, information derived from SIU, etc. In different embodiments, the loading of the model input can be from a local disk, from a remote storage location via a suitable “cloud” network, etc. Obtaining the model input can include: receiving data, participating in the preprocessing of the data, preprocessing only a part of the data and / or receiving only another part of the data, and generating the preprocessed data, etc. Processing the model input can include at least one of the following: detection, noise reduction, improvement of signal-to-noise ratio, defining bounding boxes, etc. The model input can be received from one or more sources such as one or more sensors, one or more communication units, one or more memory units, one or more image processors, etc.
[0052] From the received model inputs (e.g., the training data 1100 during the training phase), the model backbone 1202 can extract features for performing the task, such as the curvature of the road, lane markings, etc. included in the image, and can pass the extracted features to the hybrid block 1204. Here, these features are combined, reduced from high-dimensional model input data to low-dimensional latent vectors, and fed into the latent layer 1206, which is represented as a relatively compact set of neurons as described above. In this way, when reaching the latent layer 1206, as the entire set of neurons of the AI model 1200 is compressed, the amount of data or the complexity of the data stream can be reduced. This compression further improves the computational efficiency because fewer data mappings need to be learned and processed.
[0053] The latent layer 1206 helps to learn the data characteristics and thus simplifies the data representation. Then, these data features can be stored in the individual neurons 1208. The policy head 1210 processes the information received from the latent layer 1206, which can include, for example, environmental information for performing a task, such as the curvature of the road, lane markings, and the vehicle's current position relative to the road, the vehicle's current speed and lateral acceleration, whether there are other vehicles nearby, etc. The information from the latent layer 1206 can also include information about the latent layer itself, such as a reduction in the number of neurons within the latent layer 1206 representing the entire neuron set of the complex model structure. The policy head 1210 outputs some model outputs based on the processed information. In an embodiment, the model outputs can include outputting driving operation decisions, such as instructions or actions to turn the steering wheel in order to increase the lateral acceleration and keep the vehicle centered within a curved lane.
[0054] In some embodiments, the model backbone 1202 and the hybrid block 1204 can be configured to map the model input into the latent layer 1206 according to some criteria that can be stored in a database of semantic relationships. In some embodiments, the model backbone 1202 learns to compress the input data dimensions to encode the latent representation of the features, while the policy head 1210 re - creates the encoded latent representation into a reconstructed output, such as a model output. For example, the model backbone 1202 can be configured to generate a compressed latent layer 1206 of the model input using a one - dimensional vector representing one or more elements of the model input. In one embodiment, the compressed latent layer 1206 can be represented as a vector V, where V = [E1, E2, E3, … EN], where E1 refers to element 1, E2 refers to element 2, E3 refers to element 3, and EN refers to element N. Each element can be a one - dimensional or multi - dimensional matrix. Each element can represent potentially useful features around the vehicle, such as lane boundary lines, lane centerlines, nearby vehicles, traffic signs, outlines of trees, etc.
[0055] The model backbone 1202 can be configured to encode meaningful information about various data attributes in its latent manifold, and the meaningful information can subsequently be utilized to perform related tasks. In such an embodiment, the latent layer 1206 helps to reduce the dimension of the input data and eliminate irrelevant information. Thus, the reduction in the dimension of the input data can reduce the computational consumption because fewer computer resources need to be allocated to process the reduced complexity and volume of the input data. Additionally, the model accuracy can be improved because the irrelevant information that may skew the modeling is eliminated.
[0056] In some embodiments, given a latent layer 1206, a policy head 1210 can be configured to determine the behavior that a vehicle needs to follow from a set of predefined tasks. The tasks determine the actions that an autonomous vehicle needs to take. Some examples of these tasks are lane centering, distance keeping, overtaking, changing lanes, intersection handling, and traffic light handling, etc.
[0057] The model output from the AI model 1200 (e.g., from the policy head 1210) can represent actions performed in an environment of a specific application scenario (e.g., autonomous driving), such as operating the throttle, brake pedal, or steering wheel, etc. Although Figure 1 the components depicted in Figure 1 are shown as constituting the AI model, it is readily understood that other AI models customized for specific application scenarios (e.g., decision-making, autonomous driving, etc.) can also be generalized to include similar components as shown in
[0058] Within the AI model 1200, a loss function (e.g., shown as the loss function block 1300 in Figure 1 ) plays an important role during the training phase of the AI model 1200. The loss function 1300 can be configured to evaluate the difference between the model prediction given by the policy head 1210 of the AI model 1200 and the actual value provided as the prior ground truth. By measuring this difference, the loss function 1300 can quantify the performance of the AI model, enabling it to optimize its learning process and enhance its prediction ability.
[0059] The main purpose of the loss function 1300 is to minimize the difference between the output of the model and the ground truth. It serves as a guide for the model to adjust its internal parameters and update its weights in order to minimize the overall loss. This process is typically accomplished through various optimization algorithms (including gradient descent, which iteratively updates the model based on the computed loss). Thus, the loss function 1300 can act as a key feedback mechanism during the training phase. By evaluating the model performance, it provides valuable information about the direction and magnitude of the necessary changes to improve the prediction. For example, by minimizing the loss based on iterations, the AI model 1200 becomes better suited to capture the underlying patterns and relationships within the training data 11000.
[0060] As depicted in Figure 1 and the subsequent figures, the thicker arrows represent the backpropagation of the error generated by various types of loss functions towards the network architecture of the AI model.
[0061] The choice of a specific loss function (e.g., mean squared error or cross-entropy) depends on the specific task and characteristics of the AI model 1200. These functions measure the dissimilarity between the predicted output of the model and the ground truth in different ways, allowing for customized optimization based on the specific problem domain. With the help of the loss function 1300, the AI model 1200 can learn from the provided feedback, adjust its internal parameters, and enhance its prediction ability according to the current task.
[0062] However, when considering the need to regularize the AI model 1200, the current loss function module 1300 alone is not sufficient because it may not linearize the policy head or improve the reducibility of the model. Therefore, an auxiliary loss function according to an embodiment of the present disclosure is introduced, for example, as shown in block 1500. The auxiliary loss function 1500 is designed to further reduce the number of neurons / nodes present in the original latent representation of the AI model 1200, making it truly interpretable and explainable for both humans and machines. This approach eliminates the need to allocate attention or computational resources to irrelevant neurons / nodes and their associated connections in the latent representation from either a human or machine perspective. The auxiliary loss function 1500 helps to remove redundancy and simplify the mapping relationships within the model, making it more interpretable and understandable for both humans and machines.
[0063] In some embodiments, the auxiliary loss function can be derived from (one or more) specific transformations applied to the original loss function 1300, such as Figure 1 shown by the transformation block 1400 as shown. The transformation represented by block 1400 receives information from the loss function 1300 (including its functional form) and performs a number of operations to transform the function before forwarding the transformed function to the auxiliary loss function 1500. It should be noted that the information received by block 1400 from the loss function 1300 may also include the predicted output of the model given by the policy head 1210 and provided to the loss function 1300, although block 1400 itself does not perform any transformation on the prediction of the model. The transformation performed by block 1400, such as a function transformation, can be, for example but not limited to, ordinary methods such as first-order or second-order differentials.
[0064] In some embodiments, an auxiliary loss function can be obtained as follows. First, a loss quantity can be determined. In some examples, the loss quantity can define the non-linearity of the policy head of an AI model for a specified task performed during the training phase of the AI model. In some examples, the loss quantity can manifest as the original loss function relied upon during the training phase of the AI model. Then, a first derivative of the determined loss quantity can be obtained to determine the non-linearity of the policy head. In some examples, instead of obtaining the first derivative, a second derivative of the determined loss quantity can be obtained to minimize the non-linearity of the policy head. The choice of the order of the derivative can depend at least in part on the task performed by the AI model, the characteristics of the (original) loss function, the degree to which irreducibility of the AI model is to be achieved, etc. Optionally, the absolute value of the obtained derivative (first derivative or second derivative, etc.) can be used as an absolute value function so as to add it back into the policy head during training to minimize the loss for a specific task and reduce the computational complexity and thus save computational resources. The resulting function, e.g., the derivative function and / or the absolute value function, can be used as an auxiliary loss function to be minimized together with the loss function. In some examples, the derivative function and / or the absolute value function can be added to the latent activation during training. In some cases, the latent activation can refer to the original loss function related to the latent layer of the model and its function. Thus, the latent variable ending with an enhanced latent representation can have a simple relationship (e.g., an approximate linear relationship) with the actual output / action suggested by the AI model.
[0065] In some embodiments, both the loss function 1300 and the auxiliary loss function 1500 can independently store the ground truth corresponding to the training data 100. Alternatively, the ground truth can be passed together with the information provided by the loss function 1300 to block 1400 for further transmission to the auxiliary loss function 1500. In this case, block 1400 does not process the ground truth. That is, block 1400 can be transparent to other information such as the predicted output of the model or the ground truth, in addition to the information related to the functional form of the loss function 1300.
[0066] The joint collaboration between the loss function 1300, the transformation 1400, and the auxiliary loss function 1500 (as indicated by reference numeral 1900) allows for a more flexible and general approach to fixing the specification of the AI model 1200 and thus optimizing the model. By transforming the original loss function 1300, the auxiliary loss function 1500 can capture additional patterns or relationships within the latent representation of the model, leading to a further enhancement of model interpretability and resource allocation. This promotes a more focused utilization of computational resources, thereby reducing any unnecessary complexity introduced by irrelevant elements within the model (e.g., neurons / nodes and associated weighted connections). By leveraging the transformation capabilities of the transformation block 1400, the auxiliary loss function 1500 contributes to an overall improvement in the interpretability of the model and its ability to convey meaningful but direct insights for both human understanding and machine-based decision-making processes.
[0067] In some embodiments, a method for training a neural network model based on a latent representation that includes a human-interpretable variable / data representation required to perform a specified task is provided. The method includes: obtaining input data; training a neural network model based on the obtained input data; fixing a specification function during training of the neural network model by applying an auxiliary loss function on latent activations for the specified task to minimize redundancy; generating an explainable artificial intelligence-based representation of the latent representation of the neural network model based on applying the auxiliary loss function in the latent application. By leveraging the concept of normalizing invariance by fixing the specification function, the desired irreducibility in the model can be achieved, resulting in a more interpretable and explainable representation characterized by a minimum number of active neurons. This approach enables a more intuitive understanding of the model's behavior and enhances its practical applicability in various domains. The reduction of redundancy and complexity in the neural network enhances its transparency, making it easier to explain the decision-making process. The enhanced interpretability thus achieved allows researchers to optimize the model inference performance from a human perspective by reallocating limited computational resources to a manageable number of active neurons.
[0068] In some embodiments, the exemplary training process 1000 may further include a visualization module 1600 for presenting the enhanced latent representation of the AI model resulting from the introduction of the auxiliary loss function 1500. Additionally, the graphical output from the visualization module 1600 may be displayed via the GUI module 1700. This enables a comparison between the latent representation without the auxiliary loss function and the enhanced latent representation with the auxiliary loss function, thereby contributing to the improvement of model performance and interpretability during the training phase.
[0069] In particular, the visualization module 1600 can be used to provide a visual representation of the latent features of a model by enabling researchers, including developers and users, to observe the changes and improvements in the interpretability and reducibility of the model brought about by the incorporation of the auxiliary loss function. By comparing the latent representations before and after including the auxiliary loss function 1500, improvements can be identified and insights into the learning and decision-making processes of the model can be obtained. The graphical output from the visualization module 1600 displayed via the GUI module 1700 allows users to understand the impact of the auxiliary loss function on the latent representation of the model and supports informed decision-making during the training process.
[0070] This visual feedback loop enables researchers to gain an intuitive understanding of how the black-box model works under a given task and potentially make informed decisions regarding the effectiveness of the auxiliary loss function, thus contributing to the continuous improvement of model training.
[0071] Figure 2 FIG. 2000 is a block diagram illustrating another example of a trained artificial intelligence (AI) model during training according to some embodiments of the present disclosure. In Figure 2 this figure, similar to the previous figures, the same elements are denoted by corresponding reference numerals and thus are not described further here. Figure 1 and Figure 2 The main difference between
[0072] is that in the exemplary block diagram 2000, informed learning techniques in the field of AI are employed to construct the auxiliary loss function instead of relying on (multiple) transformations of the original loss function that is designed to compute the loss between the prediction and the ground truth.
[0073] As Figure 2As shown, the perception quantity library 2100 can obtain a data set related to a given task in the training phase from the training data 1100 and store the data set in the perception quantity library 2100. In an example, each part of the training data 1100 in the obtained data set can include a physical quantity related to the task. In another example, each part of the training data 1100 in the obtained data set can be processed to obtain or can assist in obtaining a physical quantity related to the task. Then, based on some indications provided by the prior knowledge 2200, the perception quantity library 2100 passes a subset of the data set it stores to the informed learning module 2300. In an example, each element in this subset can be a relevant physical quantity selected from the perception quantity library 2100 under the guidance of the prior knowledge 2200 for comparison with the variables appearing in the latent representation of the AI model 1200.
[0074] The informed learning module 2300 can be configured to process the input from the perception quantity library 2100 (i.e., the above subset) based on the knowledge stored thereon and associated with a general task or a specific task (or a series of tasks), such as filtering the input to select the part of the input that is consistent with its informed knowledge, processing the input to obtain the part of the input that is consistent with its informed knowledge, classifying the input to obtain multiple clusters consistent with the informed knowledge, etc.
[0075] Upon receiving the subset provided by the perception quantity library 2100, the informed learning module 2300 can pass the information required to construct an appropriate auxiliary loss function to the auxiliary loss function 1500 based on the knowledge stored thereon. Additionally, upon receiving the subset, the informed learning module 2300 can combine the additional knowledge provided by the prior knowledge 2200 with the knowledge stored thereon to provide the information required to construct the auxiliary loss function as needed to the auxiliary loss function 1500.
[0076] In some embodiments, the auxiliary loss function can be obtained via informed learning as follows. First, a family of relevant / predetermined variables can be received. In some examples, the policy head 1210 may require a family of relevant / predetermined variables for latent activation for a specific task together with the input data. Then, the distance between the combination of the relevant variable family and the latent representation associated with the relevant variable family can be obtained. The obtained distance can be added to the loss function to train the policy head 1210 and the latent representation and encourage informed learning. In some examples, the above distance can be any distance metric commonly used to quantify the similarity or dissimilarity between data points or representations in a neural network or an AI model, such as Euclidean distance, Manhattan distance, cosine similarity, Hamming distance, Jaccard distance, etc.
[0077] Once an appropriate auxiliary loss function is constructed, it can be aggregated with the main / original loss function to form a composite loss function, where the aggregation can be performed at a loss aggregation module as shown in block 2400. This composite loss function not only minimizes the difference between the model predictions and the ground truth that backpropagates through the layers of the network architecture of the model during the training phase, but also aims to minimize the latent representation of the model as much as possible to obtain intuitive and human-interpretable representations.
[0078] By combining the auxiliary loss function with the main loss function, the neural network can be optimized for not only more accurate predictions that are less likely to produce errors that make it unsafe but also for improved transparency. This allows researchers to gain additional insights into the decision-making process whenever the auxiliary loss function is applied to see which features of the model are considered important in the model's decisions and to interpret the behavior of the model in a more intuitive way. Finally, this aggregation of losses enables the model to learn representations that are not only effective but also more understandable and interpretable by humans. This leads to a simplified and more efficient allocation of limited computational resources, which also reduces the energy consumption of the computational resources performing unnecessary tasks.
[0079] Figure 3 FIG. 3000 is a block diagram illustrating another example of an AI model during training according to some embodiments of the present disclosure. In Figure 3 which, similar to the previous figures, the same elements are denoted by corresponding reference numerals and thus will not be described further here. Figure 2 and Figure 3 The main difference between is that the exemplary block diagram 3000 additionally includes: a policy head training module 3100 and a policy head pool 3200. The policy head training module 3100 obtains data for training the policy head from the training data 1100. Once trained, the policy head is fed into the policy head pool 3200. Based on the task type during the AI model training phase, an appropriate policy head is selected from the policy head pool 3200 and provided to the policy head 1210 in the AI model 1200.
[0080] Thus, the model training process is enhanced by incorporating a dedicated module for training the policy head and a pool of policy heads with different training strategies. The policy head training module 3100 uses the training data to adjust the policy head parameters, while the policy head pool 3200 provides flexibility in selecting the most suitable policy head for the AI model based on the task type and / or input data characteristics.
[0081] Figure 4 FIG. is a block diagram illustrating an example of a trained AI model suitable for performing a method for visualizing a model during inference according to some embodiments of the present disclosure. As Figure 4As shown, the AI model 4200 may include: a model backbone 4202, a hybrid block 4204, a policy head 4210, a latent layer 4206, and a plurality of neurons 4208 within the latent layer 4206. Note that Figure 4 and Figures 1 - 3 One of the main differences between any of Figure 4 The process depicted in
[0082] represents the inference phase of the AI model. In this phase, the AI model 4200 has completed its training and has determined and attached the connections and weights of its internal components. It should also be noted that the lower right corner of the policy head module 4210 is marked with italic L and a, indicating that the policy head 4210 undergoes backpropagation of the error generated by both the loss function L and the auxiliary loss function a. As a result, the optimization of the policy head is affected by these two loss functions. Additionally, the auxiliary loss function a can be any one of the aforementioned auxiliary loss functions or a variant thereof.
[0083] The AI model 4200 receives the test data 4100 as input and provides its (multiple) outputs to the visualization module 4600 and the execution module 4800. The (multiple) outputs of the execution module 4800 (not shown) can be used to manipulate the mechanical, electronic, or mechatronic modules, devices, components, and equipment of the vehicle to achieve the autonomous driving function for a given task. These two modules 4600 and 4800 can be communicatively connected to better achieve the objectives of the inference phase of the AI model. The output of the visualization module 4600 can be provided to the GUI module 4700 to visually present to the user the enhanced latent representation as described in the present disclosure.
[0084] In addition, an illustrative representation of the enhanced latent representation presented by the GUI module 4700 is the thumbnail 4702. The thumbnail 4702 incorporates several irreducible neurons 4704 (depicted as four neurons for illustrative purposes only and not constructed as restrictive), which are obtained by fixing the specification of the AI model by utilizing an auxiliary loss function in combination with the original loss function of the AI model and reducing the neural ensemble contained in the original latent representation. As a result, the input / output mapping of the policy head brought about by the enhanced latent representation tends to linearize, thereby minimizing the involvement of redundant factors in the decision-making process of the model. This leads to improved interpretability of the AI model, for example, by examining the enhanced latent representation shown such as the exemplary thumbnail 4702.
[0085] As shown in the thumbnail 4702, one neuron is represented as a black dot, while the other three neurons are represented as white dots. This can, for example, indicate that the current task of the AI model is significantly related to the neuron represented by the black dot. In other words, this neuron can play a crucial role in effectively encoding the information payload related to the control output of the model within the input image frame.
[0086] The enhanced latent representation can adopt various GUI arrangements, and the present disclosure does not impose any restrictions thereon. The exemplary thumbnail representation used is for illustrative purposes only, and the interpretation of the relationship between neurons and tasks should be based on a specific context and model architecture.
[0087] Now turning to Figures 5 - 8 , these figures illustrate methods 5000, 6000, 7000, and 8000 corresponding to the methods and models described above, which can be used to train an AI model to obtain an auxiliary loss function and visualize the latent representation of the AI model. Here, the order of the steps described in methods 5000, 6000, 7000, and 8000 is exemplary and does not indicate the order in which methods 5000, 6000, 7000, and 8000 will execute the steps.
[0088] Referring to Figure 5 , in some embodiments, a method 5000 for generating a human-interpretable representation of the latent representation of a neural network model required to perform a specified task is provided. Method 5000 begins at 5002 where input data is obtained. Then, at 5004, a neural network model is trained based on the obtained input data. Thereafter, at 5006, during the training of the neural network model, the specification function is fixed by applying an auxiliary loss function on the latent activation for the specified task to minimize redundancy. Finally, at 5008, an explainable artificial intelligence-based representation of the latent representation of the neural network model is generated based on applying the auxiliary loss function on the latent application.
[0089] Reference Figure 6 In some embodiments, a method 6000 of obtaining an auxiliary loss function is provided. Method 6000 begins at 6002 where a loss amount that defines a non-linearity of a policy head for a specified task is determined. Then, at 6004, the second derivative of the determined loss amount is obtained to minimize the non-linearity of the policy head. Thereafter, at 6006, the absolute value of the obtained second derivative is taken, and at 6008, this absolute value is added to a latent activation during training to minimize the loss and ensure a simple mapping between predicted actions and actual action outputs.
[0090] Reference Figure 7 In some embodiments, a method 7000 of obtaining an auxiliary loss function via informed learning is provided. Method 7000 begins at 7002 where a family of relevant predetermined variables required for a policy head to perform a latent activation for a specified task is received along with input data. Then, at 7004, a distance between a combination of the family of relevant variables and a latent representation associated with the family of relevant variables is obtained simultaneously. Thereafter, at 7006, the obtained distance is added to the loss function to train the policy head and the latent representation and encourage informed learning.
[0091] Reference Figure 8 In some embodiments, a method 8000 of visualizing latent representations of a neural network model is provided. Method 8000 begins at 8002 where input data is obtained. Then, at 8004, a neural network model is applied, which is trained based on a latent representation including variable / data representations based on explainable artificial intelligence required for performing a specified task and based on the obtained input data. The neural network may have a fixed canonical function. Thereafter, at 8006, an explainable artificial intelligence-based representation is generated during inference.
[0092] Figures 9A - 10B Illustrations of different examples of operating an autonomous driving system for different tasks.
[0093] Figure 9A and Figure 9B respectively depict a perspective view 9000a and a top view 9000b of a road from a driver's perspective in an exemplary autonomous driving scenario for a lane centering task according to some embodiments of the present disclosure. As shown, from the driver's perspective, the field of view 9000a includes a ego vehicle 9010, other vehicles 9030 and 9040 driving in front of the ego vehicle 9010, and planned trajectories 9070a and 9070b of the ego vehicle 9010 (represented by thinner solid lines and thicker dashed lines, respectively). Additionally, the field of view 9000a also contains lanes, lane markings, lane boundaries (such as obstacles), green spaces, and urban contours, which are not denoted by reference marks for the sake of brevity.
[0094] The field of view 9000b is a top view captured from above the ego vehicle 9010, looking down at the portion of the road on which it is traveling. It can be observed that the ego vehicle 9010 is being driven along the first and second lanes on the far left, straddling two lanes. Another vehicle 9030 is traveling in the first lane on the far left, while another vehicle 9040 is traveling in the first lane on the far right. As an example, Figure 9A and Figure 9B can relate to a lane centering task in an autonomous driving scenario. In this task, the control objective of an AI model carried by or associated with the ego vehicle 9010 (e.g., deployed on the ego vehicle's autonomous driving system or deployed in the cloud that communicates with the ego vehicle via V2X) is to keep the ego vehicle within the center of the lane. The planned trajectory 9070a of the ego vehicle 9010 represents the planned trajectory inferred by the AI model trained in the case of an auxiliary loss function during the inference phase based on the current model input. The planned trajectory 9070b represents the planned trajectory inferred by the AI model trained without any auxiliary loss function during the inference phase based on the current model input.
[0095] Assuming that these two trajectories represent a complete solution for the same task, the trajectory 9070a is smoother and shorter compared to the trajectory 9070b. This indicates that the AI model performs better in trajectory planning when an auxiliary loss function is introduced, as the planned trajectory ensures a shorter and smoother transition from straddling the lane markings to a safe centered position. On the other hand, the trajectory 9070b exhibits more oscillations (which may imply involvement of the brake pedal) and a longer length, where the ego vehicle 9010 dangerously approaches another vehicle 9030 and moves at a certain distance away from the center of the lane of the ego vehicle 9010, thus indicating poor basic control performance (e.g., trajectory planning ability) and an error in the prediction of the trajectory 9070b.
[0096] Figure 10A and Figure 10B respectively depict a perspective view and a top view of the road from the driver's perspective in an exemplary autonomous driving scenario for a distance keeping task according to some embodiments of the present disclosure. As shown, from the driver's perspective, the field of view 10000a includes the ego vehicle 10010, other vehicles 10020 driving in front of the ego vehicle 10010, and the planned trajectories 10050a and 10050b of the ego vehicle 10010 (represented by solid and dashed lines respectively). Additionally, the field of view 10000a contains unmarked lanes, lane markings, lane boundaries (e.g., obstacles), green spaces, and city outlines.
[0097] Field of view 10000b is a top-down view captured from above ego vehicle 10010, looking down at the portion of the road it is driving on. It can be observed that ego vehicle 10010 is driving right behind another vehicle 10020, and both are in the second lane to the far left.
[0098] As an example, Figure 10A and Figure 10B It may involve a distance keeping task in an autonomous driving scenario. In this task, the control objective of an AI model carried by or associated with the ego vehicle 10010 (e.g., deployed on the ego vehicle's autonomous driving system or deployed in a cloud that communicates with the ego vehicle via V2X) is to maintain a safe distance from the vehicle ahead traveling in the same lane. The planned trajectory 10050a of the ego vehicle 10010 represents the planned trajectory inferred by the AI model trained under the auxiliary loss function based on the current model input during the inference phase. The planned trajectory 10050b represents the planned trajectory inferred by the AI model trained without any auxiliary loss function based on the current model input during the inference phase.
[0099] Assuming that the two trajectories represent complete solutions for the same task, trajectory 10050a is smoother and shorter than trajectory 10050b. In addition, according to trajectory 10050a, in the next time step T+1, the ego vehicle 10010a has decisively changed to a state where it is in a lane different from the other vehicle 10020a. On the contrary, in trajectory 10050b, in the next time step T+1, the ego vehicle 10010b is more hesitant and has not yet changed to a state where it is in a lane different from the other vehicle 10020a, and the distance between the ego vehicle 10010b and the other vehicle 10020a is much shorter than the relative position between them in the previous time step T. This is a more dangerous state and is not conducive to safe driving. In this way, without the auxiliary loss function, the vehicle is more likely to collide.
[0100] Therefore, it can be seen that when the auxiliary loss function is introduced, the AI model performs better in safety control because the planned trajectory ensures that the vehicle can quickly avoid unfavorable situations, such as the shortened safety distance caused by the braking of the vehicle in front (e.g., by smoothly changing lanes). On the other hand, although trajectory 10050b attempts to guide the ego vehicle to the adjacent lane, trajectory 10050b still leads to an unsafe situation and further reduces the following distance. Therefore, the disclosed auxiliary loss function not only provides users with an intuitive potential representation of the AI model, but also improves the completion rate of tasks during the actual model reasoning phase, enhances the reasoning performance of the model, and improves the user experience.
[0101] It should be understood that the examples described with reference to Figures 8 - 1 0 are for illustrative purposes only and should not be construed as limiting the scope of the present disclosure.
[0102] In some embodiments, the above functions / features may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored as one or more instructions or codes on a non-transitory computer-readable storage medium or a non-transitory processor-readable storage medium. Blocks of the methods or algorithms disclosed herein may be implemented in a processor-executable software module that may reside on a non-transitory computer-readable or processor-readable storage medium. The non-transitory computer-readable or processor-readable storage medium may be any storage medium accessible by a computer or a processor. By way of example and not limitation, such non-transitory computer-readable or processor-readable storage medium may include RAM, ROM, EEPROM, flash memory, CD-ROM or other optical disc storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks typically reproduce data magnetically, while discs reproduce data optically with a laser. The above combinations are also included within the scope of non-transitory computer-readable and processor-readable media. Additionally, operations of a method or algorithm may reside as one or any combination or collection of codes and / or instructions on a non-transitory processor-readable storage medium and / or a computer-readable storage medium, which may be incorporated into a computer program product.
[0103] Figure 11 An exemplary hardware and software environment of an autonomous vehicle 11000 in which the various technologies disclosed herein may be implemented. For example, vehicle 11000 is shown traveling on road 11010, and vehicle 11000 may include: a powertrain 11020 including a prime mover 11060 that is powered by an energy source 11040 and capable of providing power to a driveline 11080; and a vehicle operating system 11100 that includes a steering control 11120, a powertrain control 111400, and a braking control 11160. Vehicle 11000 may be implemented as any number of different types of vehicles, including vehicles capable of transporting people and / or goods and capable of traveling over land, over sea, through the air, underground, underwater, and / or in space, and it should be understood that the above components 11020 - 11160 may vary widely based on the type of vehicle in which they are used.
[0104] For simplicity, the embodiments discussed below will focus on wheeled land vehicles such as cars, vans, trucks, buses, motorcycles, all-terrain vehicles (ATVs), etc. In such embodiments, the energy source 11040 may include, for example, a fuel system (e.g., providing gasoline, diesel, hydrogen, etc.), a battery system, solar panels or other renewable energy sources and / or a fuel cell system. The prime mover 11060 may include one or more electric motors and / or internal combustion engines (etc.). The powertrain 11080 may include wheels and / or tires along with a transmission and / or any other mechanical driving components adapted to convert the output of the prime mover 11060 into vehicle motion and one or more brakes configured to controllably stop or slow down the vehicle 11000 and a direction or steering component adapted to control the trajectory of the vehicle 11000 (e.g., a rack and pinion steering linkage that enables one or more wheels of the vehicle 11000 to pivot about a generally vertical axis to change the angle of the rotation plane of the wheel relative to the longitudinal axis of the vehicle). In some embodiments, a combination of powertrains and energy sources may be used (e.g., in the case of an electric / gas hybrid vehicle), and in other embodiments, multiple electric motors (e.g., dedicated to individual wheels or axles) may be used as the prime mover 11060. In the case of a hydrogen fuel cell implementation, the prime mover 11060 may include one or more electric motors, and the energy source 11040 may include a fuel cell system powered by hydrogen fuel.
[0105] The direction control 11120 may include one or more actuators or sensors for controlling and receiving feedback from the direction or steering component to enable the vehicle 11000 to follow a desired trajectory. The powertrain control 111400 may be configured to control the output of the powertrain 11020 (e.g., control the output power of the prime mover 11060, control the gears of the transmission in the powertrain 11080, etc.), thereby controlling the speed and / or direction of the vehicle 11000. The brake control 11160 may be configured to control one or more brakes that slow down or stop the vehicle 11000, such as disc or drum brakes coupled to the wheels of the vehicle.
[0106] Other vehicle types (including but not limited to all-terrain vehicles or tracked vehicles, as well as construction equipment) may utilize different powertrains, drivelines, energy sources, direction controls, powertrain controls, and brake controls. Additionally, in some embodiments, some components may be combined, e.g., where the direction control of the vehicle is primarily handled by changing the output of one or more prime movers. Accordingly, the embodiments disclosed herein are not limited to the specific applications of the techniques described herein in autonomous vehicles, wheeled vehicles, land vehicles.
[0107] In the illustrated embodiment, full or semi-automatic control of the vehicle 11000 is implemented in the primary vehicle control system 11180, which may include: one or more processors 11220 and one or more memories 11240, each processor 11220 being configured to execute program code instructions 11260 stored in the memory 11240. The processor 11220 may include, for example, a (plurality of) graphics processing units (GPUs) and / or a (plurality of) central processing units (CPUs). The processor 11220 may also include an application specific integrated circuit (ASIC), other chip sets, logic circuits, and / or data processing devices. The memory 11240 may be used to load and store, for example, data and / or instructions for the control system 11180. The memory 11240 may include any combination of the following: suitable volatile memories (e.g., read only memory (ROM), dynamic random access memory (DRAM), random access memory (RAM)), non-volatile memories (e.g., flash memory, memory cards, storage media), and / or other storage devices. When the embodiments are implemented in software, the techniques described herein may be implemented with modules, procedures, functions, entities, etc. that perform the functions described herein. The modules may be stored in the memory and executed by the processor. The memory may be implemented within or external to the processor, where the memory may be communicatively coupled to the processor via various means known in the art.
[0108] The sensor 11300 may include various sensors adapted to collect information from the vehicle's surrounding environment for controlling the operation of the vehicle 11000. For example, the sensor 11300 may include: one or more detection and ranging sensors (e.g., RADAR sensor 11340, LIDAR sensor 11360, or both), a satellite navigation (SATNAV) sensor 11320, e.g., compatible with any one of various satellite navigation systems (e.g., GPS (Global Positioning System), GLONASS (Global Navigation Satellite System), BeiDou Navigation Satellite System (BDS), Galileo, Compass), etc. The radio detection and ranging (RADAR) sensor 11340, the light detection and ranging (LIDAR) sensor 11360, and the digital camera 11380 (which may include various types of image capture devices capable of capturing still images and / or video images) may be used to sense stationary and moving objects in the immediate vicinity of the vehicle. The camera 11380 may be a monochrome camera or a stereo camera and may record still images and / or video images. The SATNAV sensor 11320 may be used to determine the vehicle's position on the earth using satellite signals. The sensor 11300 may optionally include an inertial measurement unit (IMU) 11400. The IMU 11400 may include multiple gyroscopes and accelerometers capable of detecting the linear and rotational motion of the vehicle 11000 in three directions. One or more other types of sensors (e.g., wheel rotation sensors / encoders 11420) may be used to monitor the rotation of one or more wheels of the vehicle 11000.
[0109] In various embodiments, the removable hardware pod is vehicle agnostic and thus can be installed on various non-autonomous vehicles, including: cars, buses, vans, trucks, scooters, tractor-trailers, recreational vehicles, etc. Although autonomous vehicles typically include a full sensor suite, in many embodiments, the removable hardware pod can include a dedicated sensor suite that typically has fewer sensors than a full autonomous vehicle sensor suite and can include: an IMU, a 3D positioning sensor, one or more cameras, a LIDAR unit, etc. Additionally or alternatively, the hardware pod can collect data from the non-autonomous vehicle itself, such as by integrating with the vehicle's CAN bus to collect various vehicle data including: vehicle speed data, braking data, steering control data, etc. In some embodiments, the removable hardware pod can include a computing device that can aggregate the data collected by the removable pod sensor suite and the vehicle data collected from the CAN bus, and upload the collected data to a computing system for further processing (e.g., uploading the data to the cloud). In many embodiments, the computing device in the removable pod can apply a timestamp to each instance of the data before uploading the data for further processing. Additionally or alternatively, one or more sensors within the removable hardware pod can apply a timestamp to the data as it is collected (e.g., the LIDAR unit can provide its own timestamp). Similarly, the computing device within an autonomous vehicle can apply a timestamp to the data collected by the autonomous vehicle's sensor suite and can upload the timestamped autonomous vehicle data to a computer system for additional processing.
[0110] The output of sensor 11300 can be provided to a set of master control subsystems 11200, including, for example, a localization subsystem, a perception subsystem, a planning subsystem, and a control subsystem. The localization subsystem is primarily responsible for precisely determining the position and orientation (sometimes also referred to as "pose" or "pose estimation") of vehicle 11000 within its surrounding environment, typically within a certain reference frame. In some embodiments, the pose is stored as localization data in memory 11240. In some embodiments, a surface model is generated from a high-definition map and stored as surface model data in memory 11240. In some embodiments, detection and ranging sensors store their sensor data in memory 11240 (e.g., a radar data point cloud is stored as radar data). In some embodiments, calibration data is stored in memory 11240. The perception subsystem is primarily responsible for detecting, tracking, and / or identifying objects within the environment around vehicle 11000. Machine learning models, such as the machine learning models discussed above according to some embodiments, can be used to plan vehicle trajectories. The control subsystem 11200 is primarily responsible for generating appropriate control signals for controlling various controls in the vehicle control system 11180 in order to achieve the planned trajectory of vehicle 11000. Similarly, machine learning models can be used to generate one or more signals to control the autonomous vehicle 11000 to achieve the planned trajectory.
[0111] It should be understood that Figure 11 the set of components shown for the vehicle control system 11180 is merely an example. Separate sensors may be omitted in some embodiments. Additionally or alternatively, in some embodiments, Figure 11Multiple sensors of the same type as shown can be used for redundancy and / or to cover different areas around the vehicle. Additionally, in addition to the types described above, there can be other types of additional sensors to provide actual sensor data related to the operation and environment of the wheeled land vehicle. Similarly, different types of control subsystems and / or combinations of control subsystems can be used in other embodiments. Further, although the main control subsystem 11200 is illustrated as being separate from the processor 11220 and the memory 11240, it will be appreciated that in some embodiments, some or all of the functions of the main control subsystem 11200 can be implemented with program code instructions 11260 residing in one or more memories 11240 and executed by one or more processors 11220, and in some cases, the main control subsystem 11200 can be implemented using the same processor(s) and / or memory. The subsystems can be implemented at least in part using various dedicated circuit logics, various processors, various field programmable gate arrays (FPGAs), various application specific integrated circuits (ASICs), various real-time controllers, etc., and as described above, multiple subsystems can utilize circuits, processors, sensors, and / or other components. Additionally, the various components in the vehicle control system 11180 can be networked in various ways.
[0112] For example, the vehicle 11000 can include one or more network interfaces, such as network interface 1154, which is adapted to communicate with one or more networks 11500 (such as LAN, WAN, wireless networks, and / or the Internet, etc.) to allow the transfer of information with other vehicles, computers, and / or electronic devices, including, for example, central services such as cloud services from which the vehicle 11000 receives environmental data and other data for its automatic control.
[0113] Furthermore, for additional storage, the vehicle 11000 can also include one or more mass storage devices, such as floppy disks or other removable disk drives, hard disk drives, direct access storage devices (DASDs), optical drives (such as CD drives, DVD drives, etc.), solid state storage drives (SSDs), network attached storage, storage area networks, and / or tape drives, etc. Additionally, the vehicle 11000 can include a user interface 11520 to enable the vehicle 11000 to receive multiple inputs from a user or operator and generate outputs for the user or operator, the user interface 11520 being, for example, one or more displays, touchscreens, voice and / or gesture interfaces, buttons, and other tactile controls, etc. Otherwise, user input can be received via another computer or electronic device, such as via an application on a mobile device or via a web interface, for example, from a remote operator.
[0114] This disclosure relates to systems and methods for object detection and detection confidence. The disclosed methods may be adapted for autonomous driving, but may also be used in other applications, such as robotics, video analysis, weather forecasting, medical imaging, etc. This disclosure may be described with respect to an exemplary autonomous vehicle 11000. Although this disclosure primarily provides examples using autonomous vehicles, other types of devices may be used to implement the various methods described herein, such as robots, camera systems, weather forecasting devices, medical imaging devices, etc. Additionally, these methods may be used to control an autonomous vehicle, or for other purposes, such as but not limited to video surveillance, video or image editing, video or image search or retrieval, object tracking, weather forecasting (e.g., using radar data) and / or medical imaging (e.g., using ultrasound or magnetic resonance imaging (MRI) data).
[0115] Those of ordinary skill in the art understand that each of the units, algorithms, and steps described and disclosed in the embodiments of this disclosure is implemented using either electronic hardware or a combination of software for a computer and electronic hardware. Whether the functionality runs in hardware or software depends on the conditions of the application and the design requirements of the technical solution. Those of ordinary skill in the art may use different ways to implement the functionality for each specific application, and such implementation should not exceed the scope of this disclosure. Those of ordinary skill in the art can understand that since the working processes of the above systems, devices, and units are basically the same, he / she can refer to the working processes of the systems, devices, and units in the above embodiments. For ease of description and simplification, these working processes will not be described in detail.
[0116] If the software functional unit is implemented, used, and sold as a product, it can be stored in a readable storage medium in a computer. Based on this understanding, the technical solution proposed in this disclosure can be substantially or partially implemented in the form of a software product. Alternatively, a part of the technical solution beneficial to the prior art can be implemented in the form of a software product. The software product in the computer is stored in a storage medium, which includes a plurality of commands for a computing device (such as a personal computer, a server, or a network device) to run all or some of the steps disclosed in the embodiments of this disclosure. The storage medium includes a USB flash drive, a removable hard disk, a ROM, a RAM, a floppy disk, or other types of media capable of storing program code. Although this disclosure has been described in connection with what is considered to be the most practical and preferred embodiments, it should be understood that this disclosure is not limited to the disclosed embodiments, but is intended to cover various arrangements made without departing from the scope of the broadest interpretation of the appended claims.
[0117] However, other modifications, variations and substitutions are also possible. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive. In the claims, any reference signs placed in parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of other elements or steps than those listed in a claim. Further, as used herein, the term "a" or "an" is defined as one or more than one. Additionally, the use of introductory phrases such as "at least one" and "one or more" in the claims should not be construed to imply that the introduction of another claim element by the indefinite article "a" or "an" limits any particular claim containing such introduced claim element to inventions containing only one such element, even when the same claim includes the introductory phrases "one or more" or "at least one" and the indefinite article, such as "a" or "an". This also applies to the use of the definite article. Terms such as "first" and "second", unless otherwise stated, are used to arbitrarily distinguish between the elements so described. Thus, these terms are not necessarily intended to indicate a temporal or other prioritization of such elements. The fact that certain measures are recited in mutually different claims does not preclude the advantageous use of a combination of these measures. Although certain features of the invention have been illustrated and described herein, many modifications, substitutions, changes and equivalents are now contemplated by those of ordinary skill in the art. Accordingly, it is to be understood that the appended claims are intended to cover all such modifications and variations that fall within the true spirit of the invention.
[0118] It should be understood that, for clarity, the various features of the embodiments of the present disclosure described in the context of separate embodiments may also be provided in combination in a single embodiment. Conversely, the various features of the embodiments of the present disclosure described in the context of a single embodiment for brevity may also be provided separately or in any suitable sub-combination. Those skilled in the art will appreciate that the embodiments of the invention are not limited by what has been specifically shown and described above. Rather, the scope of the embodiments of the present disclosure is defined by the appended claims and their equivalents.
[0119] A previous description of the disclosed embodiments is provided so that others may make or use the disclosed subject matter. Various modifications to these embodiments will be apparent, and the general principles defined herein may be applied to other embodiments without departing from the spirit or scope of the previous description. Thus, the previous description is not intended to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein. Accordingly, the claims are not intended to be limited to the aspects shown herein, but are to be accorded the full scope consistent with the language of the claims, where the reference to an element in the singular is not intended to mean "one and only one" unless explicitly so stated, but rather "one or more". Unless specifically stated otherwise, the term "some" means one or more. All structural and functional equivalents of the elements of the various aspects described in the previous description (which are known or later will be known) are expressly incorporated herein by reference and are intended to be covered by the claims. In addition, nothing disclosed herein is dedicated to the public, whether or not such disclosure is explicitly recited in the claims. A claim element should not be construed as a means-plus-function unless the element is expressly recited using the phrase "means for". It should be understood that the specific order or hierarchy of blocks in the disclosed processes is an example of an illustrative approach. Based on design preferences, it is understood that the specific order or hierarchy of blocks in the process may be rearranged while remaining within the scope of the previous description. The appended method claims present the elements of the various blocks in a sample order and are not meant to be limited to the specific order or hierarchy presented.
[0120] The various examples shown and described are provided only as examples to illustrate the various features of the claims. However, the features shown and described with respect to any given example need not be limited to the associated example and may be used or combined with other examples shown and described. Additionally, the claims are not intended to be limited by any one example. The above method descriptions and process flow diagrams are provided only as exemplary examples and are not intended to require or imply that the blocks of the various examples must be executed in the order presented. As will be appreciated, the order of the blocks in the above examples may be executed in any order. Words such as "thereafter," "then," "next," etc. are not intended to limit the order of the blocks; these words are simply used to guide the reader through the description of the method. Additionally, any reference to a claim element in the singular, for example, using the articles "a," "an," or "the," should not be construed as limiting the element to the singular. The various illustrative logical blocks, modules, circuits, and algorithmic blocks described in connection with the examples disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and blocks have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present invention. The hardware used to implement the various illustrative logic, logic blocks, modules, and circuits described in connection with the examples disclosed herein may be implemented or performed with a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but in the alternative, the processor may be any processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Alternatively, some blocks or methods may be performed by circuitry specific to a given function.
[0121] Further embodiments are listed below.
[0122] Embodiment 1. A method for training a neural network model based on a latent representation, the latent representation including a human-interpretable data representation required to perform a specified task, the method comprising: obtaining input data; training a neural network model based on the obtained input data; fixing a normalization function during training of the neural network model by applying an auxiliary loss function on a latent activation for the specified task to minimize redundancy; and generating a human-interpretable representation of the latent representation of the neural network model based on applying the auxiliary loss function on the latent representation.
[0123] Example 2. The method according to Example 1, wherein the auxiliary loss function is obtained by the following operations: determining a loss amount that defines the non-linearity of the policy head for the specified task; obtaining the second derivative of the determined loss amount to minimize the non-linearity of the policy head; taking the absolute value of the obtained second derivative; and adding the absolute value to the latent activation during training.
[0124] Example 3. The method according to any one of Examples 1-2, wherein the auxiliary loss function is obtained via informed learning by the following operations: receiving, together with the input data, a family of relevant predetermined variables required for the policy head to perform latent activation for the specified task; simultaneously obtaining the distance between the combination of the relevant predetermined variable family and the latent representation associated with the relevant predetermined variable family; and adding the obtained distance to the loss function to train the policy head and the latent representation and encourage the informed learning.
[0125] Example 4. The method according to any one of Examples 1-3, wherein training of the policy head and the latent representation is performed simultaneously during training of the neural network model.
[0126] Example 5. The method according to any one of Examples 1-4, wherein the specified task includes positioning the vehicle at the center of the lane in which the vehicle is traveling.
[0127] Example 6. The method according to any one of Examples 1-5, further comprising: extracting relevant quantities from the input data, the relevant quantities including: the boundary lines of the lane in which the vehicle is traveling; the distance from the vehicle to the boundary lines; the curvature of the boundary lines; or information that allows the vehicle to stay at the center of the lane.
[0128] Example 7. The method according to any one of Examples 1-6, wherein the auxiliary loss function is determined by the following operations: also receiving a predetermined policy head together with the received family of relevant predetermined variables.
[0129] Example 8. The method according to any one of Examples 1-7, wherein the predetermined policy head and the received variables are shared among neurons within the latent representation.
[0130] Example 9. A method for visualizing the latent representation of a neural network model, comprising: obtaining input data; applying a neural network model that is trained based on the latent representation including a human-interpretable data representation required for performing a specified task and based on the obtained input data, wherein the neural network model has a fixed canonical function; and generating a representation based on explainable artificial intelligence.
[0131] Example 10. The method according to Example 9, wherein the neural network model is trained by visualizing the latent representation including an interpretable artificial intelligence-based representation required for performing a specified task.
[0132] Example 11. The method according to any one of Examples 9-10, wherein the latent representation is compared with a second latent representation, and a measurement of how close the vehicle is to the center of the lane is determined.
[0133] Example 12. A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations including: obtaining input data; training a neural network model based on the obtained input data; fixing a normalization function by applying an auxiliary loss function to latent activations for a specified task during training of the neural network model to minimize redundancy; and generating an interpretable artificial intelligence-based representation of the latent representation of the neural network model based on applying the auxiliary loss function to the latent representation.
[0134] Example 13. The non-transitory computer-readable storage medium according to Example 12, wherein the auxiliary loss function is obtained by: determining a loss amount that defines the non-linearity of a policy head for the specified task; obtaining a second derivative of the determined loss amount to minimize the non-linearity of the policy head; taking the absolute value of the obtained second derivative; and adding the absolute value to the latent activation.
[0135] Example 14. The non-transitory computer-readable storage medium according to any one of Examples 12-13, wherein the auxiliary loss function is obtained via informed learning by: receiving, together with the input data, a family of relevant predetermined variables required for the policy head to perform the latent activation for the specified task; simultaneously obtaining a distance between a combination of the family of relevant predetermined variables and a latent representation associated with the family of relevant predetermined variables; and adding the obtained distance to the loss function to train the policy head and the latent representation and encourage the informed learning.
[0136] Example 15. The non-transitory computer-readable storage medium according to any one of Examples 12-14, wherein training of the policy head and the latent representation is performed simultaneously during training of the neural network model.
[0137] Example 16. The non-transitory computer-readable storage medium according to any one of Examples 12-15, wherein the specified task includes positioning the vehicle at the center of the lane in which the vehicle is traveling.
[0138] Example 17. The non - transitory computer - readable storage medium according to any one of Examples 12 - 16, wherein the operation further includes extracting relevant quantities from the input data, and the relevant quantities include: the boundary lines of the lane in which the vehicle travels; the distance from the vehicle to the boundary lines; the curvature of the boundary lines; or information that allows the vehicle to stay at the center of the lane.
[0139] Example 18. The non - transitory computer - readable storage medium according to any one of Examples 10 - 17, wherein the auxiliary loss function is determined by the following operation: also receiving a predetermined policy head together with the received family of relevant predetermined variables.
[0140] Example 19. The non - transitory computer - readable storage medium according to any one of Examples 10 - 17, wherein the predetermined policy head and the received variables are shared among neurons within the latent representation.
[0141] Example 20. A computer - implemented system, including one or more memory devices storing instructions that, when executed by the one or more processors, cause the one or more processors to perform the method according to any one of Examples 1 - 11.
Claims
1. A method for training a neural network model based on a latent representation comprising a human-interpretable representation of data required to perform a specified task, the method comprising: Get input data; Training the neural network model based on the acquired input data; During training of the neural network model, a canonical function is fixed by applying an auxiliary loss function on the latent activations for the specified task to minimize redundancy; as well as Based on applying the auxiliary loss function on the latent representation, an explainable artificial intelligence-based representation of the latent representation of the neural network model is generated.
2. The method according to claim 1, wherein: The interpretable artificial intelligence-based representation is a human-interpretable interpretable representation or a machine-interpretable interpretable representation.
3. The method according to claim 1, wherein: The auxiliary loss function is obtained by the following operation: determining a loss amount that defines a nonlinearity of a policy head for the specified task; Obtaining a second-order derivative of the determined loss amount to minimize the nonlinearity of the strategy head; Take the absolute value of the obtained second-order derivative; as well as The absolute value is added to the latent activation during the training.
4. The method according to claim 1, wherein: The auxiliary loss function is obtained through informed learning by the following operation: receiving, along with the input data, a strategy header with a family of relevant predetermined variables required for the potential activation for the specified task; Simultaneously obtaining distances between combinations of the related predetermined variable families and potential representations associated with the related predetermined variable families; as well as The obtained distance is added to the loss function to train the policy head and the latent representation and encourage the informed learning.
5. The method according to claim 1, wherein: The designated task includes positioning the vehicle in the center of a lane in which the vehicle is traveling.
6. The method according to claim 3, further comprising: Extract relevant quantities from the input data, the relevant quantities comprising: The boundary line of the lane where the vehicle is traveling; the distance from the vehicle to the boundary line; the curvature of the boundary line; or Information allowing the vehicle to stay in the center of the lane.
7. The method according to claim 4, wherein: The auxiliary loss function is determined by the following operation: receiving a predetermined strategy header along with the received relevant predetermined variable family, Wherein the predetermined strategy head and the received associated predetermined variable family are shared among neurons within the latent representation.
8. A method for visualizing a latent representation of a neural network model, comprising: Get input data; applying a neural network model, the neural network model being trained based on the latent representation including the interpretable artificial intelligence based data representation required to perform a specified task and based on the acquired input data, wherein the neural network model has a fixed canonical function; and Producing Explainable AI-based Representations.
9. A non-transitory computer-readable storage medium having stored thereon instructions, which when executed by one or more processors cause the one or more processors to perform the method according to any one of claims 1-7.
10. A computer-implemented system comprising: One or more memory devices storing instructions which, when executed by one or more processors, cause the one or more processors to perform a method according to any one of claims 1-7.