Visualizing neurons in artificial intelligence models

By applying layer-by-layer correlation propagation and visual backpropagation in autonomous driving systems, the generation of representations based on interpretable artificial intelligence is solved, and the problem of the lack of interpretability of neural network models in autonomous driving systems is improved, and computing efficiency and model accuracy are improved.

CN119963796APending Publication Date: 2025-05-09ULTRABERRY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410420608.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-08
Filing Date
2024-04-09
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The lack of interpretability in existing neural network models in autonomous driving systems makes it difficult to understand the decision-making process of the model, identify potential biases or errors, and improve computational efficiency.

Method used

By acquiring neurons from multiple neurons in the AI ​​model, determining the input region of interest (ROI) related to the task, and applying operations such as layer-by-layer correlation propagation (LRP) and visual backpropagation (VBP), a representation based on interpretable artificial intelligence is generated.

Benefits of technology

The visualization and interpretation of the neural network model decision process is realized, the computing efficiency and model accuracy are improved, and the user's trust in the model and the safety of the autonomous driving system are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963796A_ABST
    Figure CN119963796A_ABST
Patent Text Reader

Abstract

A method of visualizing neurons in an artificial intelligence (AI) model for automatic driving is provided. The method comprises the following steps: acquiring one or more neurons from a plurality of neurons of an AI model for a task; determining, for each of the one or more neurons, a respective region of interest (ROI) of the input related to the task, where the respective ROI is encoded by the one or more neurons for the task; and generating an interpretable artificial intelligence-based representation of the respective ROI of the determined input for at least a portion of the one or more neurons by applying a first operation including layer-by-layer correlation propagation (LRP).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and more particularly to a method, a non-transitory computer-readable storage medium, and a computer-implemented system for visualizing neurons in an artificial intelligence (AI) model for autonomous driving. Background Art

[0002] As computing and vehicle technologies continue to advance, features related to automation have become more powerful and widely available, and are able to control vehicles in a wider variety of environments. For example, for automobiles, the Society of Automotive Engineers (SAE) has established a standard (J3016) that identifies six levels of driving automation from "no automation" to "full automation." The SAE standard defines Level 0 as "no automation," where the human driver performs all aspects of the dynamic driving task full time, even when enhanced by warning or intervention systems. Level 1 is defined as "driver assistance," where the vehicle controls steering or acceleration / deceleration (but not both) in at least some driving modes, allowing the operator to perform all remaining aspects of the dynamic driving task. Level 2 is defined as "partial automation," where the vehicle controls steering and acceleration / deceleration in at least some driving modes, allowing the operator to perform all remaining aspects of the dynamic driving task. Level 3 is defined as "conditional automation," where, for at least some driving modes, the automated driving system performs all aspects of the dynamic driving task, with the expectation that the human driver will respond appropriately to intervention requests. Level 4 is defined as "high automation," where, only for specific conditions, the automated driving system performs all aspects of the dynamic driving task, even if the human driver does not respond appropriately to intervention requests. Specific conditions for Level 4 may be, for example, specific types of roads (e.g., highways) and / or specific geographic areas (e.g., geographically isolated metropolitan areas that have been appropriately mapped). Finally, Level 5 is defined as "full automation," where the vehicle is able to operate without operator input under all conditions.

[0003] Artificial intelligence and machine learning have made significant progress, especially in the field of neural network models. These models, including multi-layer perceptrons (MLPs), convolutional neural networks (ConvNets), recurrent neural networks (RNNs), and transformers, have gained widespread recognition for their outstanding ability to handle complex tasks and achieve impressive performance across different fields. However, due to the complex and layered architecture of these models, it is challenging to understand the underlying calculations of these models. Typically, they consist of multiple interconnected layers with highly nonlinear activation functions. In addition, these models involve a large number of parameters, usually on the order of millions, requiring a lot of training to determine the optimal values ​​of these parameters. Although these complex architectures and parameters enable the models to capture complex patterns and relationships of input data, they also contribute to the opaque "black box" nature of these models, in which it is difficult for users to understand how the models reach their predictions, decisions, or actions. The combination of a large number of parameters and complex architectures increases the difficulty of explaining and understanding the inner workings of these models. The underlying calculations and decision-making processes within these models often remain opaque, making it challenging to understand how the models reach their predictions or classifications. The lack of transparency has raised concerns in various fields including legal, medical, and commercial applications, where explainability and accountability are important considerations.

[0004] This lack of explainability not only hinders humans’ ability to trust and interpret the decisions made by the models, but also hampers attempts to identify potential biases or errors in their predictions. Furthermore, the computational efficiency of the models suffers because the allocation of limited computer resources (e.g., neurons and associated electrical energy) to power their decision making and implementation to complete assigned tasks and learn from errors is unconstrained. Furthermore, model accuracy degrades over time because errors and biases that lead to significant errors in their predictions go undetected and unfixed.

[0005] Addressing these challenges is important for increasing the trust and adoption of neural network models in real-world applications by increasing computational efficiency and model accuracy, as model providers and even end users have an increasing need to clearly understand the decision-making process of the model while achieving computational cost and energy reduction. In addition, transparent and interpretable neural network models can facilitate the identification and mitigation of bias and discrimination patterns, thereby ensuring fairness and accountability in automated decision-making systems, while improving model accuracy to prevent serious errors caused by bias and prediction errors. Therefore, there have been efforts to transform neural network models into more transparent models or "white box" models.

[0006] Various techniques have been proposed, including the use of model post-hoc interpretability methods and explainable AI frameworks. Post-hoc interpretability methods aim to provide explanations for model predictions after they are generated, while explainable AI frameworks focus on designing models with built-in interpretability from scratch. However, post-hoc interpretability methods and explainable AI frameworks each have their own shortcomings. Post-hoc methods typically provide an approximation of model behavior, which may not accurately capture the decision-making process of the underlying model. Some post-hoc methods rely on model-specific details, making them less applicable to a variety of AI models and architectures. For very large and complex models, explainable AI frameworks may not scale well, slowing down the explainability process in practical scenarios. Therefore, addressing deficiencies such as limited accuracy, model relevance, and / or scalability in post-hoc interpretability methods and explainable AI frameworks has also become an emerging challenge in the pursuit of improving the interpretability and explainability of neural network models.

[0007] As neural network models are increasingly embedded in autonomous driving systems and have become an indispensable part of autonomous driving systems, it is important to develop algorithms and systems that enhance the interpretability of these models to make autonomous driving systems more trustworthy and acceptable, helping model developers or end users understand the neural network decision-making process of the system. It is also important to develop algorithms and systems that reduce the computational cost and related energy usage of limited resources while improving the accuracy of model reasoning for various AI models. Summary of the invention

[0008] Embodiments of the present disclosure provide a method, a non-transitory computer-readable storage medium, and a computer-implemented system for visualizing neurons in an artificial intelligence (AI) model, which is general or dedicated to a specific application scenario, such as decision-making, and more specifically autonomous driving. In some embodiments, the method may include: obtaining one or more neurons from a plurality of neurons for a task of the AI ​​model; determining a corresponding region of interest (ROI) of an input related to the task for each of the one or more neurons, wherein the corresponding ROI is encoded by the one or more neurons for the task; and generating an interpretable artificial intelligence-based representation of the corresponding ROI of the determined input for at least a portion of the one or more neurons by applying a first operation including layer-by-layer relevance propagation (LRP).

[0009] In some embodiments, the explainable artificial intelligence-based representation is a human-interpretable explainable representation or a machine-interpretable explainable representation.

[0010] In some embodiments, input is collected for the task by sensors, recorded human driving databases, and / or cloud storage.

[0011] Furthermore, in some embodiments, the input is a processed image frame, and the respective region of interest (ROI) for each of the one or more neurons comprises a set of pixels in the processed image frame corresponding to the respective ROI.

[0012] Furthermore, in some embodiments, the input is a sequence of processed image frames, and a corresponding region of interest (ROI) for each of the one or more neurons comprises a union of respective sets of pixels in the sequence of processed image frames, each set of pixels corresponding to a respective sub-region of interest (sub-ROI) of each processed image frame in the sequence of processed image frames.

[0013] Furthermore, in some embodiments, generating the explainable AI-based representation includes applying a second operation including visual back propagation (VBP).

[0014] Furthermore, in some embodiments, visual back propagation (VBP) is applied sequentially after layer-wise relevance propagation (LRP).

[0015] Furthermore, in some embodiments, an artificial intelligence (AI) model includes: a hybrid block and a model backbone.

[0016] In addition, in some embodiments, applying the first operation and the second operation includes: applying layer-by-layer relevance propagation (LRP) through a mixing block to obtain a weight mask for a feature map of one or more neurons; weighting the feature map of one or more neurons using the weight mask to obtain a weighted feature map of one or more neurons; and applying visual back propagation (VBP) to back propagate the weighted feature map of one or more neurons through a model backbone.

[0017] Furthermore, in some embodiments, the input is a spectrogram of a speech segment.

[0018] In addition, an embodiment of the present invention provides a non-temporary computer-readable storage medium having instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to perform operations including: obtaining one or more neurons from a plurality of neurons for a task of an AI model; determining a corresponding region of interest (ROI) of an input related to the task for each of the one or more neurons, wherein the corresponding ROI is encoded by the one or more neurons for the task; and generating an explainable artificial intelligence-based representation of the corresponding ROI of the determined input for at least a portion of the one or more neurons by applying a first operation including layer-by-layer relevance propagation (LRP).

[0019] In some embodiments, the explainable artificial intelligence based representation is a human-interpretable explainable representation or a machine-interpretable explainable representation.

[0020] In some embodiments, input is collected for the task by sensors, recorded human driving databases, and / or cloud storage.

[0021] Furthermore, in some embodiments, the input is a processed image frame, and the respective region of interest (ROI) for each of the one or more neurons comprises a set of pixels in the processed image frame corresponding to the respective ROI.

[0022] Furthermore, in some embodiments, the input is a sequence of processed image frames, and a corresponding region of interest (ROI) for each of the one or more neurons comprises a union of respective sets of pixels in the sequence of processed image frames, each set of pixels corresponding to a respective sub-region of interest (sub-ROI) of each processed image frame in the sequence of processed image frames.

[0023] Furthermore, in some embodiments, generating the explainable AI-based representation includes applying a second operation including visual back propagation (VBP).

[0024] Furthermore, in some embodiments, visual back propagation (VBP) is applied sequentially after layer-wise relevance propagation (LRP).

[0025] Furthermore, in some embodiments, an artificial intelligence (AI) model includes: a hybrid block and a model backbone.

[0026] In addition, in some embodiments, applying the first operation and the second operation includes: applying layer-by-layer relevance propagation (LRP) through a mixing block to obtain a weight mask for a feature map of one or more neurons; weighting the feature map of one or more neurons using the weight mask to obtain a weighted feature map of one or more neurons; and applying visual back propagation (VBP) to back propagate the weighted feature map of one or more neurons through a model backbone.

[0027] Furthermore, in some embodiments, the input is a spectrogram of a speech segment.

[0028] In addition, an embodiment of the present disclosure provides a computer-implemented system comprising: one or more processors; and one or more memory devices storing instructions, which when executed by the one or more processors cause the one or more processors to perform operations including: obtaining one or more neurons from a plurality of neurons for a task of an AI model; determining a corresponding region of interest (ROI) of an input related to the task for each of the one or more neurons, wherein the corresponding ROI is encoded by the one or more neurons for the task; and generating an explainable artificial intelligence-based representation of the corresponding ROI of the determined input for at least a portion of the one or more neurons by applying a first operation including layer-by-layer relevance propagation (LRP).

[0029] In some embodiments, generating the explainable AI-based representation includes applying a second operation including visual back propagation (VBP).

[0030] It should be understood that all combinations of the above concepts and additional concepts described in more detail herein are considered to be part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure are considered to be part of the subject matter disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the embodiments of the present disclosure or related technologies, the following drawings to be described in the embodiments are briefly introduced. Obviously, the drawings are only some embodiments of the present disclosure, and ordinary technicians in this field can obtain other drawings based on these drawings without creative work. The arrows in the drawings indicate the relationship, through which the component from which the arrow starts is used to train / apply the component pointed by the arrow. The embodiments of the present disclosure will be more completely understood and appreciated according to the following detailed description in conjunction with the drawings, in which:

[0032] Figure 1A is a block diagram illustrating an example of an artificial intelligence (AI) model according to some embodiments of the present disclosure;

[0033] Figure 1B is a block diagram illustrating an example of a trained artificial intelligence (AI) model suitable for performing a method for visualizing a model according to some embodiments of the present disclosure;

[0034] Figure 2 is a block diagram illustrating an example of a proposed computer-implemented system according to some embodiments of the present disclosure;

[0035] Figure 3A is a diagram illustrating an example of a region of interest (ROI) within an input image according to some embodiments of the present disclosure;

[0036] Figure 3B is a diagram illustrating an example of a set of sub-regions of interest (sub-ROIs) in an input image sequence according to some embodiments of the present disclosure;

[0037] Figure 4 is a diagram illustrating examples of spectrograms represented in three-dimensional (3D) and two-dimensional (2D) forms, respectively, to facilitate visualization of neurons within an AI model according to some embodiments of the present disclosure;

[0038] Figure 5 is a flow chart illustrating an example of the operation of a neuron in a visualization AI model according to some embodiments of the present disclosure;

[0039] Figure 6 is a flow chart illustrating another example of the operation of a neuron in a visualization AI model according to some embodiments of the present disclosure;

[0040] Figure 7 is a flowchart illustrating an example of the operation of applying two saliency map-based visualization techniques according to some embodiments of the present disclosure;

[0041] Figure 8 is an illustration of an example of generating a human interpretable representation according to some embodiments of the present disclosure;

[0042] Fig. 9 is an illustration of another example of generating a human interpretable representation according to some embodiments of the present disclosure;

[0043] Fig.10 is an illustration of yet another example of generating a human interpretable representation according to some embodiments of the present disclosure; and

[0044] Fig.11 An exemplary hardware and software environment for an autonomous vehicle according to some embodiments of the present disclosure is illustrated.

[0045] It should be understood that for simplicity and clarity of illustration, the elements shown in the figures are not necessarily drawn to scale. For example, for clarity, the size of some elements may be exaggerated relative to other elements. In addition, where deemed appropriate, reference numerals may be repeated in the drawings to indicate corresponding or similar elements. DETAILED DESCRIPTION

[0046] With reference to the accompanying drawings, the embodiments of the present disclosure are described in detail with technical problems, structural features, purposes and effects of implementation, as follows. Specifically, the terms in the embodiments of the present disclosure are only used for the purpose of describing specific embodiments, rather than limiting the present disclosure. In the following detailed description, many specific details are set forth to provide a thorough understanding of the present invention. However, those skilled in the art will understand that the present invention can be implemented without these specific details. In other cases, well-known methods, processes and components are not described in detail to avoid confusing the present invention. The subject matter of the present invention is specifically pointed out and clearly claimed in the concluding part of the specification. However, with respect to the organization and operation method of the present invention and its purposes, features and advantages, the present invention can be best understood by referring to the following detailed description when read in conjunction with the accompanying drawings. Because the illustrated embodiments of the present invention can be implemented using electronic components and circuits known to those skilled in the art for the most part, in order to understand and appreciate the underlying concepts of the present invention, and in order not to confuse or disperse the teachings of the present invention, the details will not be explained to a greater extent than what is considered necessary as shown above. For example, the specification and / or the accompanying drawings may relate to a processor or a processing circuit. The processor may be a processing circuit. The processing circuit may be implemented as a central processing unit (CPU) and / or one or more other integrated circuits such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a full custom integrated circuit, etc., or a combination of these integrated circuits.

[0047] The following description and / or drawings may relate to images or image frames. An image is an example of a media unit. Any reference to an image may be necessarily applied to a media unit. A media unit may be an example of a sensing information unit (SIU). Any reference to a media unit may be necessarily applied to any type of natural signal, such as but not limited to a signal generated by nature, a signal representing human behavior, a signal representing an operation related to a vehicle signal, a geodetic signal, a geophysical signal, a text signal, a digital signal, a time series signal, etc. Any reference to a media unit may be necessarily applied to a SIU. The SIU may be of any type and may be sensed by any type of sensor, such as a visual light camera. An audio sensor. Sensors that may sense infrared, radar imaging, ultrasound, electro-optical, radiography, light detection and ranging (LIDAR), thermal sensors, passive sensors, active sensors, etc. Sensing may include generating a sample (e.g., pixel, audio signal, etc.) representing a signal sent or otherwise arriving at a sensor. The SIU may have one or more images, one or more video clips, text information about one or more images, text describing motion information, etc.

[0048] Any combination of any modules or units listed in any of the drawings, any part of the specification and / or any claims may be provided. Any of the units and / or modules illustrated in the present application may be implemented with hardware and / or code, instructions and / or commands stored in a non-transitory computer-readable medium, and may be included in a vehicle, outside a vehicle, in a mobile device, in a server, etc. The vehicle may be any type of vehicle, such as a ground transportation vehicle, an aviation vehicle, or a watercraft. The vehicle is also referred to as an ego vehicle. It should be understood that autonomous driving includes at least partially automatic (semi-automatic) driving of a vehicle, including all L2 level types or higher level types defined in the SAE standard.

[0049] As used herein, an artificial intelligence (AI) model may be general or specific to a specific application scenario, such as decision making, classification, prediction, etc. In particular, an AI model may be customized for conventional tasks related to autonomous driving. These tasks may be classified, for example, as perception, positioning and mapping, planning and decision making, and control. Perception tasks involve accurate detection and recognition of objects and entities in the surrounding environment. This includes identifying and classifying pedestrians, vehicles, traffic signs, traffic lights, and other relevant objects. Positioning tasks focus on determining the precise position of the vehicle in its surroundings, which involves using sensors and data to estimate the position of the vehicle relative to a known reference point or map, while mapping tasks, on the other hand, involve creating and updating representations of the surrounding environment. Positioning and mapping together enable the autonomous driving system to understand the precise position of the vehicle and navigate it effectively. Planning tasks involve generating action sequences or trajectories based on the current position and desired destination of the vehicle. Decision-making tasks require analyzing the current driving situation and determining appropriate actions, such as changing lanes, accelerating, braking, or giving way. Planning and decision-making together enable the autonomous driving system to navigate the vehicle in a safer and more efficient manner. Control tasks typically include executing planned actions and adjusting the dynamics of the vehicle to follow the desired trajectory. This includes controlling the steering, acceleration, and braking systems to maintain proper control and stability of the vehicle. The control task ensures that the vehicle's physical responses are consistent with the planned maneuvers.

[0050] The present disclosure proposes a method for visualizing neurons in an artificial intelligence (AI) model. To this end, millions of neurons in the AI ​​model must be simplified as a prerequisite to obtain an equivalent and compact representation of the entire set of neurons in the model. This facilitates obtaining an intuitive understanding of the portion of the model input that each of a limited number of neurons focuses on to complete a given task. By simplifying the representation of neurons, it is possible to intuitively grasp the response of the neural network to the model input under a given task. The simplified and compact representation of neurons allows for a more focused analysis of the behavior of the model and the contribution of neurons to the overall functionality of the AI ​​model.

[0051] As used herein, the term region of interest (ROI) refers to a unique aspect of the model input encoded by a neuron in a compact representation of a neuron relative to a specific task for which the AI ​​model is specifically trained. For example, in a lane change task related to autonomous driving, the identification of road boundaries is crucial. This is because a collision with a road boundary can lead to a serious accident, such as a vehicle overturning or damage during a lane change. Therefore, for the lane change task, at least some of the neurons in the compact representation of the neurons of the AI ​​model (also referred to as active neurons) focus their attention on the part of the model input that represents the lane boundary or contains information about the lane boundary. Then, each of the active neurons can encode the corresponding part of the model input based on the determined corresponding ROI, so that the overall functionality of the AI ​​model (i.e., executing the lane change of the vehicle) can be achieved, which requires preventing the vehicle from overlapping or colliding with any detected lane boundary in the model input. Finally, in order to generate a human-interpretable representation through which the end user and / or model developer can obtain an intuitive understanding of the underlying workings of the AI ​​model, a first operation is applied. This operation utilizes layer-by-layer relevance propagation (LRP), a technique for highlighting the contribution of individual neurons in encoding different aspects of the entire input in order to complete the task. By leveraging LRP, the method generates representations that are more easily understandable and interpretable by humans. In this way, the disclosed method provides a valuable method for visualizing and understanding the functionality of neurons within an AI model that is customized for a specific application scenario (e.g., autonomous driving). In addition, by focusing on the specific contributions of individual neurons, computing resources can be allocated to specific neurons or nodes with the highest contributions for completing a specific task, thereby improving computing efficiency.

[0052] That is, based on the generated representation, the user, model developer, or the trained model itself can intuitively understand which specific parts of the model input each active node in the AI ​​model is responsible for or focused on in a given task, which nodes are not involved in processing the model input (and potentially not involved in the model's decision-making, such as model reasoning), which nodes are most concerned with the parts of the model input associated with the model's decision-making (and which nodes are less important to these parts), etc. Therefore, the user, model developer, or trained model can modify and / or fine-tune the network structure of the previously trained AI model for a given task based on the generated human-readable representation, such as deactivating or even removing those nodes that are less relevant to the model's decision-making, so as to save computing resources and improve computing efficiency. In addition, using the human-readable representation generated by the method and system disclosed in the present application, the model developer or the trained model can check whether it is necessary to modify the type, quantity, format, etc. of the model input to better facilitate the generation of correct model decisions, thereby improving the accuracy of model reasoning for the model input in a given task, and enhancing the safety and reliability of using such AI models in applications such as autonomous driving systems.

[0053] In addition, as the human interpretability of AI models increases, users and / or trainers can provide more accurate feedback to the AI ​​models as reference data. This higher quality reference data reduces the total amount of data required by the AI ​​model to achieve a working model. This means that the AI ​​model requires less time and fewer computing resources to train its parameters to achieve a working model.

[0054] It should be noted that the disclosed method does not necessarily need to participate in the training process of the AI ​​model, thus avoiding the need to increase computing resources. The flexibility of the disclosed method is a significant feature because it does not rely on the specific complexity or implementation details of the AI ​​model. Therefore, it can be effectively applied to a wide range of AI models, regardless of their architecture, size or complexity. Scalability ensures its compatibility with various types of AI models, including neural networks, deep learning models, reinforcement learning models, or any other form of machine learning algorithms. In general, the disclosed method provides a resource-efficient and scalable method that obtains insights from AI models by alleviating the need for additional computing resources and avoiding dependence on model-specific details, ensuring compatibility across various types of AI models and then making them valuable in practical applications.

[0055] Reference is now made to the drawings, wherein like numerals refer to like parts throughout. Figure 1A Depicted is a block diagram illustrating an example of an AI model 100 according to some embodiments of the present disclosure. Figure 1AAs shown, the AI ​​model 100 may include a model backbone 102 , a mixing block 103 , a strategy head 104 , a potential layer 105 , and a plurality of neurons 106 .

[0056] The model backbone 102 constitutes the foundational part of the AI ​​model 100, which is responsible for initial data processing. In some examples, the model backbone 102 may include various layers and modules designed to extract and transform information carried in the model input. It captures, extracts, and classifies necessary features and representations from a large amount of model input data 101 (e.g., frontal images or videos of roads, or annotations of lateral accelerations), which are necessary for subsequent analysis and decision making within the AI ​​model. In an embodiment, the model backbone 102 may be a convolutional neural network (CNN) that learns different features (e.g., lines and curves of roads).

[0057] The hybrid block 103 can integrate and combine information from different parts (e.g., layers) of the model backbone 102. It enhances the overall representation of the input data by facilitating the exchange of information and feature fusion between model inputs. The hybrid block ensures the effective sharing and utilization of relevant information, improving the overall performance and accuracy of the AI ​​model. In an embodiment, the hybrid block 103 can be a multilayer perceptron (MLP), which may include: a channel hybrid MLP that allows communication between different channels; and a token hybrid MLP that allows communication between different spatial locations. These layers are interwoven (i.e., combined) to enable interaction of two types of inputs.

[0058] The policy header 104 represents a component that develops a policy and generates a final output or implements a decision based on the analysis of the processed input data. The policy header 104 also provides a higher level understanding of the input data. In other words, the policy header indicates the action to be taken based on the state of the deep learning model and the detected surrounding environment. In an embodiment, the policy header 104 can be a trainable AI model.

[0059] The latent layer 105 is a simplified or compressed representation of the model input data 101, which may include a summary of key features of the model input data 101 (e.g., features related to lane boundaries). In some embodiments, the latent layer 105 may be obtained by discarding duplicate or irrelevant data using different data representations and approximation techniques. This allows less data to be transmitted without loss, and compact models to be transmitted instead of raw data. In this way, computational efficiency may be improved because less data needs to be processed and transferred from one area to another. Moreover, model accuracy may be maintained without loss.

[0060] The latent layer 105 may include multiple neurons 106, each of which is dedicated or focused on capturing and processing specific input features or patterns for a given task. In some examples, the latent layer 105 can be used as a compact representation of the entire set of neurons within the AI ​​model. That is, the number of neurons in the latent layer 105 is limited and tolerable in terms of human interpretation. Therefore, the collective behavior of these neurons 106 can contribute to the overall processing of input data within the AI ​​model 100 in order to complete the task.

[0061] In operation, the AI ​​model 100 receives and processes a model input 101 and generates a model output 107. An example of a model input 101 may be an image signal depicting a frontal view image of a road, as shown in thumbnail 101a in the figure. However, it will be appreciated by those skilled in the art that there may also be other suitable forms of model input, such as an audio signal, a text annotation, or a combination of an audio signal and an image signal (e.g., a video stream) together with a text annotation. In some embodiments, the model input 101 is raw data from one or more sensors of the same vehicle or a separate vehicle. For example, the model input 101 may be an image including red, green, and blue (RGB) values ​​of pixels captured by a camera sensor. The model input 101 may be a raw SIU, a processed SIU, text information, information derived from the SIU, and the like. In different embodiments, the loading of the model input 101 may be from a local disk, from a remote storage location through an appropriate "cloud" network, and the like. Obtaining the model input 101 may include: receiving data, generating data, participating in the processing of data, processing only a portion of the data, and / or receiving only another portion of the data. Processing of the model input 101 may include at least one of: detection, noise reduction, improvement of signal-to-noise ratio, defining a bounding box, etc. The model input 101 may be received from one or more sources such as one or more sensors, one or more communication units, one or more memory units, one or more image processors, etc.

[0062] From the received model input 101, the model backbone 102 extracts features from the model input 101, such as the curvature of the road included in the image, lane markings, etc., and passes the extracted features to the mixing block 103. Here, these features are combined with other layers and reduced from the high-dimensional model input data to a low-dimensional latent vector as a compressed latent layer 105. In this way, the data volume / complexity of the original data 102 can be reduced by forming the compressed latent layer 105. Such compression further improves computational efficiency because less data needs to be learned and processed, as described in more detail below.

[0063] The latent layer 105 of the model input 101 helps to learn the data characteristics and simplify the data representation. Each data feature is stored as a separate neuron 106. The strategy head 104 receives the latent layer 105, processes the information given by the latent layer, which may include the above-mentioned curvature of the road, lane markings, and additionally the current position of the vehicle relative to the road, the current speed and lateral acceleration of the vehicle, whether there are other vehicles nearby, etc., and outputs the model output 107 based on the processed information. In an embodiment, the model output 107 may include an output driving operation decision for turning the steering wheel 107a to increase the lateral acceleration and keep the vehicle centered in the curved lane.

[0064] In some embodiments, the model backbone 102 and the hybrid block 103 may be configured to map the model input 101 to a latent layer 105, which may be stored in a database of semantic relations. In some embodiments, the model backbone 102 learns input data dimensionality compression to encode the latent representation of the feature, and the strategy head 104 recreates the encoded latent representation as a reconstructed output, such as a model output 107. For example, the model backbone 102 may be configured to generate a compressed latent layer 105 of the model input 101 using a one-dimensional vector representing one or more elements of the model input 101. In one embodiment, the compressed latent layer 105 may be represented as a vector V, where V = [E1, E2, E3, ... EN], E1 refers to element 1, E2 refers to element 2, E3 refers to element 3, and EN refers to element N. Each element may be a one-dimensional or multi-dimensional matrix. Each element may represent a potentially useful feature around the vehicle, such as a lane boundary line, a lane centerline, a nearby vehicle, a traffic sign, a tree outline, etc.

[0065] The model backbone 102 can be configured to encode meaningful information about various data attributes in its potential manifold, which can then be utilized to perform related tasks. In such an embodiment, the potential layer 105 helps to reduce the dimensionality of the input data and eliminate irrelevant information. Therefore, the reduction in the dimensionality of the input data can reduce computational consumption because fewer computer resources need to be allocated to process the reduced complexity and volume of the input data. In addition, model accuracy can be improved because irrelevant information that may skew the modeling is eliminated.

[0066] In some embodiments, given the latent layer 105, the policy head 104 can be configured to determine the behavior that the vehicle needs to follow from a set of predetermined tasks. The tasks determine the actions that the autonomous vehicle needs to take based on the latent layer 105. Some examples of these tasks are lane keeping, overtaking, changing lanes, intersection handling, and traffic light handling.

[0067] Model output 107 may represent an action performed in the context of a specific application scenario (e.g., autonomous driving), such as manipulating an accelerator, brake pedal, or steering wheel, which is represented by thumbnail 107a in the figure. Although the components depicted in FIG. 1 are shown as constituting an AI model, it is easy to understand that an AI model customized for a specific application scenario (e.g., decision making, autonomous driving, etc.) may also be summarized as including the following: Figure 1A Similar components as shown in .

[0068] Figure 1B Depicted is a block diagram illustrating an example of a trained artificial intelligence (AI) model suitable for performing a method for visualizing a model according to some embodiments of the present disclosure. Figure 1B and Figure 1A The main difference between Figure 1B The depicted AI model 100 has completed its training process, and therefore, all parameters of the AI ​​model have been determined. Figure 1B The connections between the components are intentionally omitted. Figure 1A is depicted in order to illustrate the data flow during the model training process. Figure 1B As shown, arrows may indicate operations that will be described in more detail below. In an embodiment, the operation performed is LRP.

[0069] LRP is a technique used in the field of artificial intelligence and deep learning to understand the contribution and relevance of input features to model output. LRP allows neural network models to be interpreted and analyzed by propagating relevance scores backward through the layers of the network. In particular, LRP operates by assigning relevance scores or weights to the output neurons (i.e., neural activations) of the model and then propagating these scores back through the layers or model components. The output neurons can optionally be neurons 106 in the potential layer 105 in the present disclosure. This back propagation process is intended to highlight the importance of different input features and their classification used by neurons such as neurons 106 when formulating the decision-making process of the model. By applying LRP, a human-interpretable representation can be generated that emphasizes the areas of the input data that are most relevant to the decision of the model from the perspective of the neurons (i.e., visualizing the input pixels that really contribute to the decision of the model). Therefore, computational efficiency can be confirmed and further improved because limited computer resources can be further allocated to the most important inputs that are considered to be the most important inputs for the decision-making of the model based on the more centralized processing of the LRP application. LRP has a wide range of applications, including image classification, natural language processing, and other fields that utilize AI models.

[0070] like Figure 1BAs depicted, human interpretable representation 109 may be a graphical user interface (GUI) representation obtained by applying LRP operations to model input 101 in FIG. 1 . An example of such a representation is shown in FIG. Figure 1B , where such a representation shows a mapping between a selected active neuron (or group of active neurons) among the active neurons of the AI ​​model 100 and its (their) associated portion in the model input for performing a given task. Another example is thumbnail 109b, for which the given task may be speech-related, such as voice control in autonomous driving. The salient portions within the model input as outlined in thumbnail 109b (e.g., a 2D representation of a spectrogram of speech) may be associated with corresponding active neurons, thereby reflecting what the latent layer encodes to complete a speech-related driving task. Details regarding an example of a human-interpretable representation 109 are described below.

[0071] Figure 2 Depicted is a block diagram illustrating an example of a proposed computer-implemented system according to some embodiments of the present disclosure. As shown, the proposed computer-implemented system may include a processor 200. The processor 200 may be a general-purpose processor or a special-purpose processor, such as an ASIC (application-specific integrated circuit), an FPGA (field programmable gate array), a SOC (system on chip), a CPLD (complex programmable logic device), etc.

[0072] As shown in the figure, the processor 200 may include: an acquisition module 216 , a determination module 217 , an LRP 218 , a visual back propagation (VBP) 219 , a representation generation module 220 and a visualization engine 221 .

[0073] The acquisition module 216 may be configured to receive model information 223 of the AI ​​model from the network 222 in which the AI ​​model is deployed and obtain knowledge of the entire neuron set of the AI ​​model from the received model information 223. Subsequently, the acquisition module 216 may be configured to acquire one or more neurons used as a compact representation of the entire neuron set and indicated as neuron information 224 from a plurality of neurons in the entire neuron set of the AI ​​model for a given task. In some implementations, depending on the different tasks to be performed, the acquisition module 216 may selectively acquire different sets of one or more neurons from the entire neuron set of the AI ​​model. These acquired neurons form a compact representation that focuses on the relevant aspects of the input required for a given task.

[0074] The processor 200 may also include a determination module 217. The determination module 217 may be configured to determine a corresponding ROI of an input related to a given task for each of the acquired one or more neurons as represented by the neuron information 224 based on the received neuron information 224. As shown, the determination module 217 may include an LRP unit 218 configured to apply an LRP operation to the model input. As a non-limiting example, the model input may be a processed signal 215. The processed signal 215 may be output from the signal processor 214, which may be used in conjunction with the signal processor 214. Figure 2 The processor 200 is shown as being separate. The signal processor 214 can be, for example, a digital signal processor (DSP). In some examples, the processed signal 215 is output from the signal processor 214 in response to receiving the unprocessed raw signal 213. The unprocessed raw signal 213 can be obtained from one or more sensors 210, a recorded human driving database 212, or from a network 222.

[0075] Optionally, the determination module 217 may also include VBP 219. VBP is a technique commonly used in the fields of computer vision and deep learning to gain insight into the image regions that most significantly contribute to the model's predictions. VBP works by transferring the gradients from the output layer (which may alternatively be Figure 1A and 1B The potential layer 105 in the VBP is propagated back to the input layer to operate, thereby attributing the relevance score to each pixel or region along the way. These relevance scores represent the importance of this specific pixel or region in contributing to the decision-making of the model. By mapping these relevance scores back to the input image, VBP promotes the generation of visually interpretable heat maps or saliency maps, making it a powerful tool for visualizing and explaining neural network models, and verifying whether the predicted results output by the model are consistent with the true values ​​obtained. In other words, by highlighting important image areas, VBP provides intuitive insights into the decision-making process, thereby contributing to the transparency and interpretability of the model used in vision-related tasks. In this way, the model accuracy of the neural network model can be confirmed and improved. In addition, computational efficiency can be enhanced because limited computer resources can be allocated to specific pixels or regions that contribute the most to the decision-making of the model, while eliminating resources that are ineffectively allocated on unnecessary input processing.

[0076] In summary, the determination module 217 within the processor 200 allows for identification and determination of task-relevant ROIs in the processed signal 215 for one or more relevant (ie, active) neurons in the latent layer based on the received neuron information 224 while enhancing model accuracy and improving computational efficiency.

[0077] The processor 200 may also include a representation generation module 220. In some examples, the representation generation module 220 may include a visualization engine 221 that is responsible for generating a visualization output 226. The visualization output 226 may be a human interpretable representation of a corresponding ROI (e.g., processed signal 215) determined for the model input of at least a portion of the acquired one or more neurons. As a non-limiting example, the visualization output 226 may be presented in a Figure 2 An example of the output in the GUI can be represented by a thumbnail 231, which will be referred to below. Figure 8-10 to describe.

[0078] Figure 3A Depicts a diagram illustrating an example of a ROI in an input image according to some embodiments of the present disclosure. As shown, Figure 2 The processed signal 215 in the image may be indicated by reference numeral 330, which may be an image signal and is subjected to image processing such as grayscale and cropping. The processed signal 330 may include an ROI 332, which contains a rectangular region of pixels represented by reference numeral 334, which horizontally spans from x0 to x1 and vertically spans from y0 to y1. In an embodiment, the identification of the ROI within the processed signal 330 may be binary, in that the processed signal may fall within the ROI, or the signal may fall outside the ROI.

[0079] Figure 3B A diagram depicting an example of a set of sub-regions of interest (sub-ROIs) in an input image sequence according to some embodiments of the present disclosure. Sometimes, the task undertaken by the AI ​​model requires multiple inputs, rather than just a single input. For example, in the case of an overtaking task, the AI ​​model requires a sequence of image frames to accurately assess the motion of other vehicles and / or moving objects around the ego vehicle. Therefore, Figure 3B This is intended to illustrate this application scenario. Figure 3B, which depicts an example in which an image frame sequence 331 is fed into an AI model to analyze the surrounding environment during an overtaking process. The image frame sequence may include, for example, three consecutive processed images 3301-3303 captured and processed in chronological order. Each of the processed images 3301-3303 may include a corresponding sub-ROI. As a non-limiting example, the sub-ROI in the processed image 3301 includes a rectangular area of ​​pixels represented by reference numeral 3321, whose horizontal range x0 spans to x1 and the vertical range y0 spans to y1. The sub-ROI in processed image 3302 includes a rectangular area of ​​pixels represented by reference number 3322, whose horizontal range spans from x2 to x3 (where x2 is greater than x0 and x3 is greater than x1), and whose vertical range spans from y0 to y1; the sub-ROI in processed image 3303 includes a rectangular area of ​​pixels represented by reference number 3323, whose horizontal range spans from x2 to x3 and whose vertical range spans from y2 to y3 (where y2 is greater than y0 and y3 is greater than y1).

[0080] Thus, the ROI of the processed image frame sequence 331 may be represented by reference numeral 3340, where its pixels span horizontally from x0 to x3 and vertically from y0 to y3. It is understood that the number of image frames included in the image frame sequence 331 may be any suitable number and the present disclosure is not limited in this regard. It is also understood that the ROI is depicted for illustration purposes only. Figure 3A and 3B ROI (Region of Interest) and sub-ROI (Sub-Region of Interest) as depicted in FIG. In most cases, the ROI and sub-ROI may have irregular shapes. Therefore, the present disclosure does not limit the shape of the ROI and / or sub-ROI.

[0081] Figure 4 Depict examples of spectrograms in three-dimensional (3D) and two-dimensional (2D) form, respectively, according to some embodiments of the present disclosure. A spectrogram is a graphical representation of the frequency content of a signal as it changes over time. A spectrogram can be used for signal processing, for example, for audio and speech analysis. Figure 4 As shown in the lower portion of , an exemplary 2D spectrogram 434 plots the frequency spectrum of an acoustic signal (e.g., an audio signal obtained by one or more microphones arranged in a vehicle cabin) on the y-axis and time on the x-axis. The intensity (or color) of each point in the 2D spectrogram 434 represents the intensity or amplitude of the frequency component of the acoustic signal at a specific time. It provides a visual representation of how the frequency content of the signal changes over time, thereby allowing analysis and identification of various audio features, such as harmonics, formants, or transient events.

[0082] Additionally, a 3D spectrogram extends the concept of a 2D spectrogram by adding an additional third dimension, i.e., the intensity or amplitude of the frequency components plotted in the third dimension versus the intensity or color in their 2D counterparts. The third dimension of a 3D spectrogram can be visualized as a surface plot or a contour plot, where the height or color of the surface / contours represents the amplitude of the frequency component at a specific time and frequency. Figure 4 In FIG. 4 , the exemplary 3D spectrogram 430 corresponding to the 2D spectrogram 434 is shown in FIG. Figure 4 The upper part is shown by reference numeral 430 .

[0083] Figure 4 A visual representation of a spectrogram of a specific segment of an exemplary acoustic signal that captures both 2D and 3D formats is provided. As shown, the different mountain-shaped areas circled and marked as ROI 422 in the 3D spectrogram 430 correspond to the areas marked as ROI 432 in the 2D spectrogram 434. The content within the corresponding ROI 432 in the 3D spectrogram 430 or 2D spectrogram 434 may be related to the sound speech produced by humans (e.g., the driver of the vehicle), while other parts in the 3D spectrogram 430 or 2D spectrogram 434 carry other components, such as machine noise during vehicle operation, road environmental noise, and noise caused by other passengers in the vehicle cabin. In practical applications, especially in scenarios involving voice control functions for autonomous driving, ROI 432 represents a significant part of the input acoustic data of the AI ​​model customized for such application scenarios. That is, the ROI 432 representing the significant part is a part of particular interest to the (multiple) active neurons in the potential layer of the model because it plays an important role in recognizing and executing voice commands related to autonomous driving tasks.

[0084] In an embodiment, the ROI may be made to include multiple features that are important for task significance, and a 2D spectrogram 434 or 3D spectrogram 430 may be constructed within the identified ROI 432 to additionally map the "level of interest" or input relevance of each pixel to a potential neuron on a continuous scale, where the color within the 2D spectrogram 434 or the height of the 3D spectrogram 430 indicates the level of interest or significance of the potential neuron.

[0085] Return to reference Figure 1B , wherein an exemplary implementation of a human-readable representation 109 (i.e., a thumbnail 109b) is depicted with Figure 4 The 2D spectrum diagram shown in 434 is consistent. Figure 4As shown, when a 2D spectrogram representation 434 is used as a human-readable representation, the outline of the ROI 432 can indicate the time-frequency components of the speech signal (e.g., collected by a microphone or microphone array within a vehicle cabin) that are concentrated on / interested by the active node and encoded by the active node for model decision making. By viewing a human-readable representation such as the 2D spectrogram representation 434, a user, a model developer, or the model itself can determine whether the vocal component of the collected speech signal is significant enough (e.g., whether it occupies enough area in the spectrogram) to adjust the settings of the speech signal sensing device (e.g., a microphone) and better help the model extract useful information payload, thereby improving the accuracy of model reasoning. Alternatively, when active nodes are concentrated on portions of the model input that are significantly deviated from the ROI, the user, model developer, or model can save unnecessary computing resources and improve the computational efficiency involved in model reasoning by deactivating or removing these nodes from the network structure of the AI ​​model.

[0086] Now go to Figures 5 and 6 , these figures illustrate methods 500 and 600 corresponding to methods and models that can be used to visualize neurons in an AI model as discussed above. Note that the order of methods 500 and 600 is exemplary and does not represent the order in which the steps of methods 500 and 600 are to be performed.

[0087] Reference Figure 5 , method 500 begins, wherein, at 502, one or more neurons are obtained from a plurality of neurons for a task of an artificial intelligence (AI) model. Then, at 504, a corresponding region of interest (ROI) of an input related to the task is determined for each of the one or more neurons, wherein the corresponding ROI is encoded for the task by the one or more neurons. Thereafter, at 506, a human interpretable representation of the determined corresponding ROI of the input is generated for at least a portion of the one or more neurons by applying a first operation including layer-wise relevance propagation (LRP).

[0088] In some implementations of method 500, the ROI may be determined by a saliency map-based visualization technique, such as layer-by-layer relevance propagation (LRP). Preferably, the ROI may be determined by a combination of saliency map-based visualization techniques, such as a combination of LRP and visual back propagation (VBP). In some implementations, in order to facilitate human interpretation and post-hoc interpretability of the AI ​​model, the relationship between the human-interpretable representation and the determined ROI may be visualized by, for example, presenting via a GUI a mapping between the relevant neurons of one or more acquired neurons and the determined ROI of the relevant neurons. As a non-limiting example, for a given task (e.g., a lane change task), for example, there may be two active neurons related to the task in the latent layer of the AI ​​model (i.e., a compact form of the entire set of neurons), one for encoding the leftmost lane boundary and the other for encoding the rightmost lane boundary, which can be illustrated, for example, with reference to the accompanying drawings. Figure 8 , where neurons 8201 and 8202 respectively encode two lane boundaries 808. Therefore, the human interpretable representation can be shown via the GUI as (i) a mapping between two active neurons and the two lane boundaries that the two active neurons are looking at, or (ii) a mapping between a neuron and the corresponding lane boundary that the neuron is looking at, etc.

[0089] In addition, as described above, the number of neurons in the entire neuron set of the AI ​​model can be huge and therefore beyond human interpretation. Therefore, there is a need to obtain a compact form of the entire neuron set to simplify the black box AI model. The compact form is one or more neurons that have been acquired, and they can be related to conventional driving tasks, for example, each of the one or more neurons that have been acquired can encode a part of the model input related to a given task. Therefore, the underlying logic of generating a human-interpretable representation of multiple neurons includes two aspects. One aspect is to present a compact form of the entire neuron set of the AI ​​model, rather than hundreds or thousands or millions of neurons of the model, and the other aspect is to show the mapping between the selected number of neurons in the one or more neurons that have been acquired and the selected (multiple) neurons in the input that are focused / interested, so that the end user and / or model developer can understand the role of each neuron (in a compact form, for example, within the potential layer of the representation of the AI ​​model) when encoding the model input and then affecting the decision of the model. By presenting neurons in a compact form, computational efficiency can be enhanced because less processing is required. In addition, by focusing on the input that contributes most to the neuron allocation of processing inputs, model accuracy can be confirmed and improved.

[0090] In some embodiments, the method of 600 illustrates another possible implementation of the operation of neurons in a visualization artificial intelligence (AI) model. For example, method 600 begins, where, at 602, one or more neurons are obtained from a plurality of neurons for a task of an artificial intelligence (AI) model. Then, at 604, a corresponding region of interest (ROI) of an input related to the task is determined for each of the one or more neurons, wherein the corresponding ROI is encoded by the one or more neurons for the task. Thereafter, at 606, a human interpretable representation of the corresponding ROI of the determined input is generated for at least a portion of the one or more neurons by applying a first operation including layer-by-layer relevance propagation (LRP) and a second operation including visual back propagation (VBP).

[0091] In some embodiments, Figure 7 As shown, the method of 700 illustrates the operation of applying the first operation and the second operation, such as Figure 6 As shown in box 606 in . For example, method 700 can begin, whereby at 702, LRP is applied through a hybrid block of the AI ​​model to obtain a weight mask for a feature map of one or more neurons. Then, at 704, the feature map of one or more neurons is weighted using the acquired weight mask to obtain a weighted feature map of one or more neurons. Thereafter, VBP is applied to back-propagate the weighted feature map of one or more neurons through the model backbone. The combination of LRP to the hybrid block and VBP back-propagation together further enhances the insights obtained from the visualization of the neurons that contribute most to the predictions made and the regions of interest of the input image that contribute most to the predictions made. The mapping of the entire image data processing to the prediction path can be performed together, from which deviations and errors can be identified. Therefore, model accuracy can be verified and improved, errors in predictions can be more easily identified and fixed, and computational efficiency can be improved when limited computer resources are properly allocated to process the regions of interest and neurons that most affect the predictions from the neural network model.

[0092] Figures 8 to 10 are different illustrations of different examples of operating an autonomous driving system for different tasks. Figures 8 to 10 Exemplary representations of raw data 802 , 902 , and 1002 , exemplary representations of model inputs 804 , 904 , and 1004 , and exemplary human-interpretable representations 805 , 807 ; 905 ; and 1005 , 1007 , respectively, are shown.

[0093] like Figures 8 to 10As shown, raw data 802, 902, and 1002 may be captured by a camera deployed on a vehicle. The camera may be configured to capture real-time images from a front view perspective of the vehicle cabin. In some implementations, model inputs 804, 904, and 1004 may be compact representations of input images (i.e., raw data 802, 902, and 1002) that capture useful features through a dedicated processor (e.g., signal processor 214 shown in FIG. 1 ). Figures 8 to 10 As shown, the raw data 802, 902, and 1002 may be RGB images and may carry rich information. For example, the raw data includes not only images of roads, but also images of scenes around the roads (e.g., other vehicles, trees, traffic signs, the sky, etc.). In contrast, in some implementations, the model inputs 804, 904, and 1004 retain only useful features and are converted into grayscale images for storage and processing by the AI ​​model. For example, Figures 8 to 10 As shown, model inputs 804, 904, and 1004 can be converted from the colored raw data 802, 902, and 1002 to grayscale images including: tree outlines 806, 906, 1006; lane boundary lines 808, 908, 1008; lane centerlines 810, 910, 1010 (e.g., first lane centerlines 810a, 910a, 1010a; and second lane centerlines 810b, 910b, 1010b); traffic sign outlines 812, 912, 1012; traffic sign text 814, 914, 1014; and other vehicles 816, 916, 1016.

[0094] exist Figure 8, an exemplary driving task handled by an AI model embedded in or integrated into a vehicle may be a lane change task. Here, two exemplary human-interpretable representations 805 and 807 are provided for illustration purposes only. Human-interpretable representations 805 and 807 each contain a schematic representation of a potential layer of an AI model. As described above, for the purpose of simplifying the black box model, the potential layer may be a compact form of the entire neuron set of the AI ​​model and is equivalent to the entire neuron set of the AI ​​model. As shown in representation 805, two neurons 8201 and 8202 in the schematic potential layer are activated under the lane change task. A rectangular area 818 corresponding to the model input 806 or the original data 802 is arranged in a representation 805 in which the corresponding ROI (i.e., the left lane boundary and the right lane boundary 808) of the model input determined for the two active neurons is highlighted. In some implementations, a user or developer may select a portion of the neurons in the schematic potential layer within the GUI of the human-interpretable representation in order to obtain an intuitive impression or understanding of the portion that each active neuron in the model input is looking at / encoding for a given task. For example, in representation 807, only the right lane boundary 808 is highlighted within the rectangular region 819 that the selected neuron 8201 focuses on in relation to performing the lane change task.

[0095] Reference now Fig. 9 and Fig.10 ,exist Fig. 9 In FIG. 1 , an exemplary driving task may be a lane centering task, for which the lane center line 908 is heightened as an ROI in a rectangular region 918 indicating the portion of the model input for the lane centering task that the active neurons 903 are looking at / encoding. Fig.10 , an exemplary driving task may be a perception task in which the vehicle seeks to identify the surrounding environment to obtain useful contextual information. As shown in representation 1005, two neurons 1021 and 1022 in the schematic latent layer are activated under the perception task. A rectangular area 1018 corresponding to the model input 1006 or the original data 1002 is arranged within the representation 1005 in which the corresponding ROIs (i.e., the traffic sign outline 1012 and the traffic sign text 1014) determined for the model input for the two active neurons are highlighted. Alternatively, in representation 1007, only the traffic sign outline 1012 is highlighted within the rectangular area 1019 that the selected neuron 1021 focuses on and is related to performing the perception task.

[0096] like Figures 8 to 10As shown, by viewing such exemplary human-readable representations (e.g., 805, 807; 905; 1005, 1007), a user, a model developer, or the model itself can determine whether image portions associated with decisions for a given task are sufficiently significant relative to the model input (e.g., whether these image portions are processed by active nodes, whether they have enough corresponding pixels, whether they are fully encoded by active patterns, etc.), so as to adjust the number (e.g., a single image frame versus a series of image frames), type (e.g., the perspective of the image, e.g., the driver's perspective or a wide-angle perspective), or format (e.g., high-definition images, color images, grayscale images, heat maps, etc.) of the model input, and better help the model extract useful information payloads, thereby improving the accuracy of model reasoning. Alternatively, when active nodes are concentrated on model input portions that significantly deviate from the ROI, the user, model developer, or the model can save unnecessary computing resources and improve the computational efficiency involved in model reasoning by deactivating or removing these nodes from the model's network structure.

[0097] It should be understood that reference Figures 8 to 10 The described examples are for illustrative purposes only and should not be construed as limiting the scope of the present disclosure.

[0098] In some embodiments, the above functions / features can be implemented with hardware, software, firmware, or any combination thereof. If implemented in software, the functions can be stored as one or more instructions or codes on a non-transitory computer-readable storage medium or a non-transitory processor-readable storage medium. The blocks of the methods or algorithms disclosed herein can be implemented in a processor-executable software module, which can reside on a non-transitory computer-readable or processor-readable storage medium. A non-transitory computer-readable or processor-readable storage medium can be any storage medium that can be accessed by a computer or processor. As an example and not limitation, such a non-transitory computer-readable or processor-readable storage medium can include RAM, ROM, EEPROM, flash memory, CD-ROM or other optical disk storage, disk storage or other magnetic storage device, or any other medium that can be used to store the required program code in the form of instructions or data structures and can be accessed by a computer. As used herein, disks and optical disks include compact disks (CDs), laser disks, optical optical disks, digital versatile disks (DVDs), floppy disks, and Blu-ray disks, wherein disks generally reproduce data magnetically, and optical disks reproduce data optically with lasers. The above combinations are also included within the scope of non-transitory computer-readable and processor-readable media. Additionally, the operations of a method or algorithm may reside as one or any combination or set of codes and / or instructions on a non-transitory processor-readable storage medium and / or computer-readable storage medium, which may be incorporated into a computer program product.

[0099] Fig.11 An exemplary hardware and software environment of an autonomous vehicle 1100 in which the various techniques disclosed herein may be implemented is illustrated. For example, the vehicle 1100 is shown traveling on a road 1101, and the vehicle 1100 may include: a powertrain 1102 including a prime mover 1106, which is powered by an energy source 1104 and is capable of providing power to a drivetrain 1108; and a vehicle operating system 1110, which includes a direction control 1112, a powertrain control 1114, and a brake control 1116. The vehicle 1100 may be implemented as any number of different types of vehicles, including vehicles capable of transporting people and / or cargo and capable of traveling by land, by sea, by air, underground, under the sea, and / or in space, and it should be understood that the above-described components 1102-1116 may vary widely based on the type of vehicle in which they are used.

[0100] For simplicity, the embodiments discussed below will focus on wheeled land vehicles, such as cars, vans, trucks, buses, motorcycles, all-terrain vehicles (ATVs), etc. In such embodiments, the energy source 1104 may include, for example, a fuel system (e.g., providing gasoline, diesel, hydrogen, etc.), a battery system, a solar panel or other renewable energy source, and / or a fuel cell system. The prime mover 1106 may include one or more electric motors and / or internal combustion engines (etc.). The transmission 1108 may include wheels and / or tires together with a transmission and / or any other mechanical driving components suitable for converting the output of the prime mover 1106 into vehicle motion and one or more brakes configured to controllably stop or slow the vehicle 1100 and a direction or steering component suitable for controlling the trajectory of the vehicle 1100 (e.g., a rack and pinion steering linkage that enables one or more wheels of the vehicle 1100 to pivot about a substantially vertical axis to change the angle of the rotation plane of the wheel relative to the longitudinal axis of the vehicle). In some embodiments, a combination of powertrain and energy source may be used (e.g., in the case of an electric / gas hybrid vehicle), and in other embodiments, multiple electric motors (e.g., dedicated to individual wheels or axles) may be used as prime movers 1106. In the case of a hydrogen fuel cell implementation, prime mover 1106 may include one or more electric motors, and energy source 1104 may include a fuel cell system powered by hydrogen fuel.

[0101] Steering control 1112 may include one or more actuators or sensors for controlling and receiving feedback from a steering or steering assembly to enable vehicle 1100 to follow a desired trajectory. Powertrain control 1114 may be configured to control the output of powertrain 1102 (e.g., controlling the output power of prime mover 1106, controlling the gears of a transmission in drivetrain 1108, etc.), thereby controlling the speed and / or direction of vehicle 1100. Braking control 1116 may be configured to control one or more brakes that slow or stop vehicle 1100, such as disc or drum brakes coupled to the wheels of the vehicle.

[0102] Other vehicle types (including but not limited to all-terrain vehicles or tracked vehicles, and construction equipment) may utilize different powertrains, drivetrains, energy sources, directional controls, powertrain controls, and brake controls. In addition, in some embodiments, some components may be combined, for example, where the directional control of the vehicle is primarily handled by changing the output of one or more prime movers. Therefore, the embodiments disclosed herein are not limited to the specific application of the technology described herein in autonomous vehicles, wheeled vehicles, and land vehicles.

[0103] In the illustrated embodiment, full or semi-automatic control of the vehicle 1100 is implemented in a main vehicle control system 1118, which may include: one or more processors 1122 and one or more memories 1124, each processor 1122 being configured to execute program code instructions 1126 stored in the memory 1124. The processor 1122 may include, for example, a (multiple) graphics processing unit (GPU) and / or a (multiple) central processing unit (CPU). The processor 1122 may also include an application-specific integrated circuit (ASIC), other chipsets, logic circuits and / or data processing devices. The memory 1124 may be used to load and store data and / or instructions, such as for the control system 1118. The memory 1124 may include any combination of the following: suitable volatile memory (e.g., read-only memory (ROM), dynamic random access memory (DRAM), random access memory (RAM)), non-volatile memory (e.g., flash memory, memory card, storage medium) and / or other storage devices. When the embodiment is implemented in software, the technology described herein may be implemented with modules, processes, functions, entities, etc. that perform the functions described herein. The module may be stored in the memory and executed by the processor.The memory may be implemented within the processor or external to the processor, where the memory may be communicatively coupled to the processor via various means known in the art.

[0104] Sensors 1130 may include various sensors suitable for collecting information from the surrounding environment of the vehicle for controlling the operation of the vehicle 1100. For example, sensors 1130 may include: one or more detection and ranging sensors (e.g., RADAR sensor 1134, LIDAR sensor 1136, or both), satellite navigation (SATNAV) sensor 1132, for example, compatible with any of various satellite navigation systems (e.g., GPS (Global Positioning System), GLONASS (Global Navigation Satellite System), BeiDou Navigation Satellite System (BDS), Galileo, compass), etc. Radio detection and ranging (RADAR) 1134 and light detection and ranging (LIDAR) sensors 1136 and digital cameras 1138 (which may include various types of image capture devices capable of capturing still images and / or video images) can be used to sense stationary objects and moving objects in the immediate vicinity of the vehicle. Camera 1138 can be a monochrome camera or a stereo camera, and can record still images and / or video images. SATNAV sensor 1132 can be used to determine the location of the vehicle on the earth using satellite signals. The sensors 1130 may optionally include an inertial measurement unit (IMU) 1140. The IMU 1140 may include multiple gyroscopes and accelerometers capable of detecting linear and rotational motion of the vehicle 1100 in three directions. One or more other types of sensors (e.g., wheel rotation sensors / encoders 1142) may be used to monitor the rotation of one or more wheels of the vehicle 1100.

[0105] In various embodiments, the removable hardware pod is vehicle agnostic and can therefore be installed on a variety of non-autonomous vehicles, including: cars, buses, vans, trucks, mopeds, tractor trailers, sport vehicles, etc. Although autonomous vehicles typically contain a full sensor suite, in many embodiments, the removable hardware pod can contain a dedicated sensor suite that typically has fewer sensors than a fully autonomous vehicle sensor suite and can include: an IMU, a 3D positioning sensor, one or more cameras, a LIDAR unit, etc. Additionally or alternatively, the hardware pod can collect data from the non-autonomous vehicle itself, such as by integrating with the vehicle's CAN bus to collect various vehicle data including: vehicle speed data, braking data, steering control data, etc. In some embodiments, the removable hardware pod can include a computing device that can aggregate the data collected by the removable pod sensor suite and the vehicle data collected from the CAN bus, and upload the collected data to a computing system for further processing (e.g., uploading the data to the cloud). In many embodiments, a computing device in a removable pod can apply a timestamp to each instance of data before uploading the data for further processing. Additionally or alternatively, one or more sensors within the removable hardware pod can apply a timestamp to the data as it is collected (e.g., a lidar unit can provide its own timestamp). Similarly, a computing device within an autonomous vehicle can apply a timestamp to data collected by the autonomous vehicle's sensor suite, and can upload the timestamped autonomous vehicle data to a computer system for additional processing.

[0106] The output of the sensor 1130 can be provided to a set of main control subsystems 1120, including, for example, a positioning subsystem, a perception subsystem, a planning subsystem, and a control subsystem. The positioning subsystem is primarily responsible for accurately determining the position and orientation (sometimes also referred to as "pose" or "pose estimation") of the vehicle 1100 in its surroundings, usually in a certain reference frame. In some embodiments, the pose is stored in the memory 1124 as positioning data. In some embodiments, the surface model is generated from a high-definition map and stored in the memory 1124 as surface model data. In some embodiments, the detection and ranging sensors store their sensor data in the memory 1124 (for example, the radar data point cloud is stored as radar data). In some embodiments, calibration data is stored in the memory 1124. The perception subsystem is primarily responsible for detecting, tracking and / or identifying objects in the environment around the vehicle 1100. Machine learning models, such as those discussed above according to some embodiments, can be used to plan vehicle trajectories. The control subsystem 1120 is primarily responsible for generating appropriate control signals for controlling various controls in the vehicle control system 1118 so as to achieve the planned trajectory of the vehicle 1100. Similarly, the machine learning model can be used to generate one or more signals to control the autonomous vehicle 1100 to achieve the planned trajectory.

[0107] You should understand that Fig.11 The collection of components for vehicle control system 1118 shown in is only an example. In some embodiments, individual sensors may be omitted. Additionally or alternatively, in some embodiments, Fig.11 The same type of multiple sensors shown can be used for redundancy and / or to cover different areas around the vehicle. In addition, in addition to the above types, there can also be other types of additional sensors to provide actual sensor data related to the operation and environment of the wheeled land vehicle. Similarly, different types of control subsystems and / or combinations of control subsystems can be used in other embodiments. In addition, although the main control subsystem 1120 is illustrated as being separated from the processor 1122 and the memory 1124, it will be understood that in some embodiments, some or all of the functions of the main control subsystem 1120 can be implemented with program code instructions 1126 that reside in one or more memories 1124 and are executed by one or more processors 1122, and in some cases, the main control subsystem 1120 can be implemented using (multiple) the same processors and / or memories. The subsystem can be implemented at least in part using various dedicated circuit logic, various processors, various field programmable gate arrays (FPGAs), various application-specific integrated circuits (ASICs), various real-time controllers, etc., as described above, multiple subsystems can utilize circuits, processors, sensors, and / or other components. In addition, the various components in the vehicle control system 1118 can be networked in various ways.

[0108] For example, the vehicle 1100 may include one or more network interfaces, such as network interface 1154, which is suitable for communicating with one or more networks 1150 (e.g., a LAN, WAN, wireless network and / or the Internet, etc.) to allow information to be transmitted with other vehicles, computers and / or electronic devices, including, for example, a central service such as a cloud service, from which the vehicle 1100 receives environmental data and other data for its automatic control.

[0109] In addition, for additional storage, the vehicle 1100 may also include one or more mass storage devices, such as floppy disks or other removable disk drives, hard disk drives, direct access storage devices (DASD), optical drives (e.g., CD drives, DVD drives, etc.), solid-state storage drives (SSD), network attached storage, storage area networks, and / or tape drives, etc. In addition, the vehicle 1100 may include a user interface 1152 to enable the vehicle 1100 to receive multiple inputs from a user or operator and generate outputs for the user or operator, such as one or more displays, touch screens, voice and / or gesture interfaces, buttons and other tactile controls, etc. Otherwise, the user input may be received via another computer or electronic device, such as via an application on a mobile device or via a web interface, such as from a remote operator.

[0110] Disclosed herein are systems and methods relating to object detection and detection confidence. The disclosed methods may be suitable for autonomous driving, but may also be used for other applications, such as robotics, video analysis, weather forecasting, medical imaging, and the like. The present disclosure may be described with respect to an example autonomous driving vehicle 1100. Although the present disclosure primarily provides examples using autonomous driving vehicles, the various methods described herein may be implemented using other types of devices, such as robots, camera systems, weather forecasting devices, medical imaging devices, and the like. In addition, these methods may be used to control autonomous driving vehicles, or for other purposes, such as, but not limited to, video surveillance, video or image editing, video or image search or retrieval, object tracking, weather forecasting (e.g., using radar data), and / or medical imaging (e.g., using ultrasound or magnetic resonance imaging (MRI) data).

[0111] It is understood by a person skilled in the art that each of the units, algorithms, and steps described and disclosed in the embodiments of the present disclosure is implemented using electronic hardware or a combination of software and electronic hardware for a computer. Whether the function runs in hardware or in software depends on the conditions of the application and the design requirements of the technical solution. A person skilled in the art can use different ways to implement the functions for each specific application, and such implementation should not exceed the scope of the present disclosure. It is understood by a person skilled in the art that since the working processes of the above-mentioned systems, devices, and units are basically the same, he / she can refer to the working processes of the systems, devices, and units in the above-mentioned embodiments. For ease of description and simplification, these working processes will not be described in detail.

[0112] If the software functional unit is implemented, used and sold as a product, it can be stored in a readable storage medium in a computer. Based on this understanding, the technical solution proposed in the present disclosure can be basically or partially implemented in the form of a software product. Alternatively, a part of the technical solution that is beneficial to traditional technology can be implemented in the form of a software product. The software product in the computer is stored in a storage medium, which includes multiple commands for a computing device (such as a personal computer, a server or a network device) to run all or some of the steps disclosed in the embodiments of the present disclosure. The storage medium includes a USB disk, a mobile hard disk, a ROM, a RAM, a floppy disk or other types of media capable of storing program code. Although the present disclosure has been described in conjunction with what are considered to be the most practical and preferred embodiments, it should be understood that the present disclosure is not limited to the disclosed embodiments, but is intended to cover various arrangements made without departing from the scope of the broadest interpretation of the attached claims.

[0113] However, other modifications, changes and substitutions are also possible. Therefore, the description and drawings are considered to be illustrative, not restrictive. In the claims, any reference numerals placed in brackets should not be interpreted as limiting the claims. The word "comprising" does not exclude the existence of other elements or steps other than those listed in the claims. In addition, the term "one" or "an" as used herein is defined as one or more than one. In addition, the use of introductory phrases such as "at least one" and "one or more" in the claims should not be interpreted as implying that any particular claim containing such introduced claim elements is limited to an invention containing only one such element by introducing another claim element by the indefinite article "one" or "one", even when the same claim includes the introductory phrases "one or more" or "at least one" and indefinite articles such as "one" or "one". The same applies to the use of definite articles. Unless otherwise specified, terms such as "first" and "second" are used to arbitrarily distinguish the elements described by these terms. Therefore, these terms are not necessarily intended to indicate the time or other priority of such elements. The fact that certain measures are stated in mutually different claims does not mean that the combination of these measures cannot be used advantageously. Although certain features of the present invention have been illustrated and described herein, many modifications, substitutions, changes and equivalents may now occur to those skilled in the art. It is therefore to be understood that the appended claims are intended to cover all such modifications and changes that fall within the true spirit of the invention.

[0114] It should be understood that various features of embodiments of the present disclosure described in the context of separate embodiments for clarity may also be provided in combination in a single embodiment. Conversely, various features of embodiments of the present disclosure described in the context of a single embodiment for brevity may also be provided individually or in any suitable sub-combination. Those skilled in the art will appreciate that embodiments of the present invention are not limited by what has been specifically shown and described above. Instead, the scope of the embodiments of the present disclosure is defined by the appended claims and their equivalents.

[0115] The previous description of the disclosed embodiments is provided to enable others to make or use the disclosed subject matter. Various modifications to these embodiments will be apparent, and the general principles defined herein can be applied to other embodiments without departing from the spirit or scope of the previous description. Therefore, the previous description is not intended to be limited to the embodiments shown herein, but to conform to the broadest scope consistent with the principles and novel features disclosed herein. Therefore, the claims are not intended to be limited to the aspects shown herein, but to conform to the full scope consistent with the language claims, wherein the reference to the elements in the singular form is not intended to mean "one and only one", unless explicitly stated so, but "one or more". Unless otherwise specifically stated, the term "some" refers to one or more. All structural and functional equivalents of the elements of the various aspects described in the previous description (which are known or will be known later) are expressly introduced herein as a reference and are intended to be covered by the claims. In addition, whether such disclosure is explicitly stated in the claims, the content disclosed herein is not intended to be dedicated to the public. Unless the phrase "device for..." is used to explicitly state the elements, the claim elements should not be interpreted as device plus function. It should be understood that the specific order or hierarchy of blocks in the disclosed process is an example of an illustrative method. Based on design preferences, it is understood that the specific order or hierarchy of blocks in the process can be rearranged while remaining within the scope of the previous description. The attached method claims present elements of the various blocks in a sample order and are not meant to be limited to the specific order or hierarchy presented.

[0116] The various examples shown and described are provided only as examples to illustrate the various features of the claims. However, the features shown and described with respect to any given example are not necessarily limited to the associated examples, and may be used or combined with other examples shown and described. In addition, the claims are not intended to be limited by any one example. The above method descriptions and process flow charts are provided only as illustrative examples, and are not intended to require or imply that the blocks of the various examples must be executed in the order presented. As will be understood, the order of the blocks in the above examples can be executed in any order. Words such as "thereafter", "then", "next", etc. are not intended to limit the order of the blocks; these words are simply used to guide the reader through the description of the method. In addition, any reference to a claim element in the singular form, for example, using the article "one", "an" or "the", should not be interpreted as limiting the element to the singular. The various illustrative logic blocks, modules, circuits, and algorithm blocks described in conjunction with the examples disclosed herein may be implemented as electronic hardware, computer software, or a combination of both. In order to clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and blocks have been generally described above according to their functions. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the entire system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such implementation decisions should not be interpreted as causing departure from the scope of the present invention. Hardware for implementing the various illustrative logics, logic blocks, modules, and circuits described in conjunction with the examples disclosed herein may be implemented or executed with a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but in an alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Alternatively, some blocks or methods may be performed by circuits specific to a given function.

[0117] Further examples are listed below.

[0118] Embodiment 1. A method for visualizing neurons in an artificial intelligence (AI) model for autonomous driving, comprising: obtaining one or more neurons from a plurality of task-specific neurons of the AI ​​model; determining a corresponding region of interest (ROI) of task-related input for each of the one or more neurons, wherein the corresponding ROI is encoded by the one or more neurons for the task; and generating a human-interpretable representation of the corresponding ROI of the determined input for at least a portion of the one or more neurons by applying a first operation including layer-by-layer relevance propagation (LRP).

[0119] Embodiment 2. A method according to embodiment 1, wherein the input is collected for the task by sensors, a recorded human driving database and / or cloud storage.

[0120] Embodiment 3. A method according to any one of embodiments 1-2, wherein the input is a processed image frame, and the corresponding region of interest (ROI) for each of the one or more neurons includes a set of pixels in the processed image frame corresponding to the corresponding ROI.

[0121] Embodiment 4. A method according to any one of embodiments 1-3, wherein the input is a sequence of processed image frames, and the corresponding region of interest (ROI) for each of the one or more neurons includes: a union of respective pixel sets in the sequence of processed image frames, and the respective pixel sets correspond to respective sub-regions of interest (sub-ROIs) of respective processed image frames in the sequence of processed image frames.

[0122] Embodiment 5. A method according to any one of embodiments 1-4, wherein generating the human interpretable representation includes: applying a second operation including visual back propagation (VBP) to identify, from the determined corresponding ROIs, a specific ROI that contributes most to the prediction made by the artificial intelligence (AI) model to complete the task, so as to improve computational efficiency and enhance model accuracy.

[0123] Embodiment 6. A method according to any one of embodiments 1-5, wherein the visual back propagation (VBP) is applied sequentially after the layer-by-layer relevance propagation (LRP), and the LRP is applied to identify the one or more neurons from the multiple neurons that contribute most to the predictions made by the artificial intelligence (AI) model to complete the task, so as to improve computational efficiency and enhance model accuracy.

[0124] Embodiment 7. A method according to any one of embodiments 1-6, wherein the artificial intelligence (AI) model comprises: a hybrid block and a model backbone.

[0125] Embodiment 8. A method according to any one of embodiments 1-7, wherein the applying the first operation and the second operation comprises: applying the layer-by-layer relevance propagation (LRP) through the mixing block to obtain a weight mask for the feature map of the one or more neurons; weighting the feature map of the one or more neurons using the weight mask to obtain a weighted feature map of the one or more neurons; and applying the visual back propagation (VBP) to back propagate the weighted feature map of the one or more neurons through the model backbone.

[0126] Embodiment 9. A method according to any one of embodiments 1-8, wherein the input is a spectrogram of a speech segment.

[0127] Embodiment 10. A non-temporary computer-readable storage medium having instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to perform operations including: obtaining one or more neurons from a plurality of neurons for a task of an AI model; determining a corresponding region of interest (ROI) of an input related to the task for each of the one or more neurons, wherein the corresponding ROI is encoded by the one or more neurons for the task; and generating a human-interpretable representation of the corresponding ROI of the determined input for at least a portion of the one or more neurons by applying a first operation including layer-by-layer relevance propagation (LRP).

[0128] Embodiment 11. The non-transitory computer-readable storage medium of embodiment 10, wherein the input is collected for the task by sensors, a recorded human driving database, and / or cloud storage.

[0129] Embodiment 12. A non-temporary computer-readable storage medium according to any one of embodiments 10-11, wherein the input is a processed image frame, and the corresponding region of interest (ROI) for each of the one or more neurons includes a set of pixels in the processed image frame corresponding to the corresponding ROI.

[0130] Embodiment 13. A non-temporary computer-readable storage medium according to any one of embodiments 10-12, wherein the input is a sequence of processed image frames, and the corresponding region of interest (ROI) for each of the one or more neurons includes: a union of respective sets of pixels in the sequence of processed image frames, wherein the respective sets of pixels correspond to respective sub-regions of interest (sub-ROIs) of respective processed image frames in the sequence of processed image frames.

[0131] Embodiment 14. A non-transitory computer-readable storage medium according to any one of embodiments 10-13, wherein generating the human-interpretable representation includes: applying a second operation including visual back propagation (VBP) to identify, from the determined corresponding ROIs, a specific ROI that contributes most to the prediction made by the artificial intelligence (AI) model to complete the task, so as to improve computational efficiency and enhance model accuracy.

[0132] Embodiment 15. A non-transitory computer-readable storage medium according to any one of embodiments 10-14, wherein the visual back propagation (VBP) is applied sequentially after the layer-by-layer relevance propagation (LRP), and the LRP is applied to identify the one or more neurons from the multiple neurons that contribute most to the predictions made by the artificial intelligence (AI) model to complete the task, so as to improve computational efficiency and enhance model accuracy.

[0133] Embodiment 16. A non-transitory computer-readable storage medium according to any one of embodiments 10-15, wherein the artificial intelligence (AI) model includes: a hybrid block and a model backbone.

[0134] Embodiment 17. A non-temporary computer-readable storage medium according to any one of embodiments 10-16, wherein the applying the first operation and the second operation comprises: applying the layer-by-layer relevance propagation (LRP) through the mixing block to obtain a weight mask for the feature map of the one or more neurons; weighting the feature map of the one or more neurons using the weight mask to obtain a weighted feature map of the one or more neurons; and applying the visual back propagation (VBP) to back propagate the weighted feature map of the one or more neurons through the model backbone.

[0135] Embodiment 18. A non-transitory computer-readable storage medium according to any one of Embodiments 10-17, wherein the input is a spectrogram of a speech segment.

[0136] Embodiment 19. A computer-implemented system comprising: one or more processors; and one or more memory devices storing instructions, wherein the instructions, when executed by the one or more processors, cause the one or more processors to perform operations comprising: obtaining one or more neurons from a plurality of neurons for a task of an AI model; determining a corresponding region of interest (ROI) of an input related to the task for each of the one or more neurons, wherein the corresponding ROI is encoded by the one or more neurons for the task; and generating a human-interpretable representation of the corresponding ROI of the determined input for at least a portion of the one or more neurons by applying a first operation comprising layer-by-layer relevance propagation (LRP).

[0137] Embodiment 20. A system according to embodiment 19, wherein generating the human interpretable representation includes: applying a second operation including visual back propagation (VBP) to identify a specific ROI from the determined corresponding ROIs that contributes most to the prediction made by the artificial intelligence (AI) model to complete the task, so as to improve computational efficiency and enhance model accuracy.

Claims

1. A method for visualizing neurons in an artificial intelligence (AI) model for autonomous driving, comprising: Obtain one or more neurons from a plurality of task-specific neurons of the AI ​​model; determining, for each of the one or more neurons, a corresponding region of interest ROI of an input related to the task, wherein the corresponding ROI is encoded by the one or more neurons for the task; as well as By applying a first operation including layer-by-layer relevance propagation (LRP), an explainable artificial intelligence-based representation of the determined corresponding ROI of the input is generated for at least a portion of the one or more neurons.

2. The method according to claim 1, wherein: The interpretable artificial intelligence-based representation is a human-interpretable interpretable representation or a machine-interpretable interpretable representation.

3. The method according to claim 1, wherein: The input is collected for the task by sensors, a recorded human driving database and / or cloud storage, and the input is a processed image frame, and the corresponding ROI for each of the one or more neurons includes a set of pixels in the processed image frame corresponding to the corresponding ROI.

4. The method according to claim 1, wherein: The input is a sequence of processed image frames, and the corresponding ROI for each of the one or more neurons comprises a union of respective sets of pixels in the sequence of processed image frames, wherein the respective sets of pixels correspond to sub-ROIs of respective processed image frames in the sequence of processed image frames.

5. The method according to claim 1, wherein: Generating the explainable artificial intelligence-based representation includes: applying a second operation including visual back propagation (VBP) to identify, from the determined corresponding ROIs, a specific ROI that contributes most to the prediction made by the AI ​​model to complete the task, so as to improve computational efficiency and enhance model accuracy.

6. The method according to claim 5, wherein: The VBP is applied sequentially after the LRP, and the LRP is applied to identify the one or more neurons from the plurality of neurons that contribute most to the predictions made by the AI ​​model to complete the task, so as to improve computational efficiency and enhance model accuracy.

7. The method according to claim 6, wherein: The AI ​​model includes a hybrid block and a model backbone, and wherein applying the first operation and the second operation includes: applying the LRP through the mixing block to obtain a weight mask for a feature map of the one or more neurons; weighting the feature maps of the one or more neurons using the weight mask to obtain a weighted feature map of the one or more neurons; and The VBP is applied to back-propagate the weighted feature map of the one or more neurons through the model backbone.

8. The method according to claim 1, wherein: The input is a spectrogram of a speech segment.

9. A non-transitory computer-readable storage medium having stored thereon instructions, which when executed by one or more processors cause the one or more processors to perform the method according to any one of claims 1-8.

10. A computer-implemented system comprising: one or more processors; as well as One or more memory devices storing instructions which, when executed by the one or more processors, cause the one or more processors to perform a method according to any one of claims 1-8.