Method, non-transitory computer-readable storage medium and system for visualizing neurons in ai model
By visualizing AI model neurons through ROI determination and LRP/VBP, the opacity of neural networks in autonomous driving is addressed, improving transparency, efficiency, and accuracy.
Patent Information
- Application Number
- JP2024098841
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-08
- Filing Date
- 2024-06-19
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2044-06-19
AI Technical Summary
The opacity and complexity of neural network models in autonomous driving systems hinder explainability, leading to trust issues, difficulty in identifying biases, and inefficient computational use, which affects accuracy and reliability.
A method is proposed to visualize neurons of AI models by determining their corresponding regions of interest (ROIs) and applying Layer-wise Relevance Propagation (LRP) and Variability-Based Propagation (VBP) to generate human-explainable representations, allowing for a simplified understanding of model inputs and improved computational efficiency.
Enhances model transparency, reduces computational costs, and improves accuracy by focusing resources on critical model inputs, facilitating better model tuning and error identification.
Smart Images

Figure 2025078567000001_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to the field of computer technology, and more particularly, to a method for visualizing neurons of an AI model for autonomous driving, a non-transitory computer-readable storage medium, and a computer-implemented system. [Background technology]
[0002] As computational and vehicle technologies develop, automation features become more powerful and widely available, enabling vehicles to control a wider variety of environments. For example, for automobiles, the Society of Automotive Engineers (SAE) established a standard (J3016) that identifies six levels of driving automation ranging from "no automation" to "full automation." The SAE standard defines Level 0 as "no automation," where the human driver performs all dynamic driving tasks full-time, even if augmented by warning or intervention systems. Level 1 is defined as "driver-assisted," where the operator performs all remaining dynamic driving tasks by controlling steering or acceleration / deceleration (but not both) for at least some driving modes. Level 2 is defined as "partial automation," where the operator performs all remaining dynamic driving tasks by controlling steering and acceleration / deceleration for at least some driving modes. Level 3 is defined as "conditional automation," where the automated driving system performs all dynamic driving tasks for at least some driving modes, and expects the human driver to respond appropriately to intervention requests. Level 4 is defined as "high automation," where the automated driving system performs all dynamic driving tasks only for certain conditions, even if the human driver does not respond appropriately to a request to intervene. The specific conditions for Level 4 may be, for example, a specific type of road (e.g., highways) and / or a specific geographic area (e.g., a properly mapped geographically isolated metropolitan area). Finally, Level 5 is defined as "full automation," where the vehicle can operate without operator input under all conditions.
[0003] Artificial intelligence and machine learning have already made remarkable progress, especially in the field of neural network models. These models (including Multi-Layer Perceptrons (MLPs), Convolutional Neural Networks (ConvNets), Recurrent Neural Networks (RNNs), and Transformers) have been widely recognized for their outstanding ability to handle complex tasks and achieve impressive performance across a variety of domains. However, due to the complex and layered architecture of these models, it is challenging to understand the computations underlying these models. Typically, they are composed of multiple interconnected layers with highly nonlinear activation functions. Furthermore, these models concern a large number of parameters, usually reaching the order of millions, and therefore a large amount of training is required to determine the optimal values of these parameters. Although these complex architectures and parameters allow the models to capture complex modes and relationships with the input data, they also contribute to the opaque "black box" nature of these models. However, users have difficulty understanding how the models achieve predictions, decisions or actions. The combination of a large number of parameters and complex architectures makes the inner workings of these models difficult to explain and understand. The underlying computational and decision-making processes within these models often remain opaque, making it challenging to understand how the models achieve their predictions or classifications. This lack of transparency has attracted attention from a variety of fields, including legal, medical, and commercial applications, where interpretability and explainability are important considerations.
[0004] This lack of explainability not only inhibits human trust and the ability to explain decisions made by models, but also inhibits attempts to recognize potential biases and errors in predictions. Also, the computational efficiency of models suffers as they are not constrained in the allocation of limited computer resources (e.g., neurons and associated electrical energy) used to power their decisions and implementations to complete assigned tasks and learn from errors. Moreover, the accuracy of models degrades over time as errors and biases that cause significant inaccuracies in predictions go undetected and uncorrected.
[0005] Solving these challenges is crucial for increasing the reliability and adoption rate of neural network models in real-world applications by improving computational efficiency and model accuracy. Because model providers and even end users are increasingly demanding to clearly understand the decision-making process of the model while simultaneously achieving computational cost and energy reduction. Transparent and explainable neural network models can also facilitate the identification and mitigation of bias and discrimination modes, thereby ensuring fairness and measurability in automated decision-making systems and improving model accuracy to prevent serious errors due to bias and prediction errors. Therefore, there have been efforts to convert neural network models into more transparent, or "white-box" models.
[0006] Various techniques have been proposed, including post-mortem interpretable methods for models and the use of explainable AI frameworks. Post-mortem interpretable methods aim to provide an interpretation of model predictions after they are generated, while explainable AI frameworks focus on designing models with interpretability built in from the beginning. However, post-mortem interpretable methods and explainable AI frameworks each have their own deficiencies. Post-mortem methods usually provide an approximation of model actions, which may not accurately capture the decision-making process of the underlying model. Some post-mortem methods rely on model-specific details, making them less suitable for various AI models and architectures. For very large and complex models, explainable AI frameworks may not scale well, thereby slowing down the interpretable process in real-world scenarios. Thus, resolving deficiencies such as limited accuracy, model relevance, and / or scalability in post-mortem interpretable methods and explainable AI frameworks also represents a new challenge in the pursuit of improving the interpretability and explainability of neural network models.
[0007] As neural network models are increasingly incorporated into autonomous driving systems and become an integral part of the system, it is important to develop algorithms and systems that increase the interpretability of these models to make autonomous driving systems more trustworthy and acceptable, and to help model developers or end users understand the decision-making process of the system's neural network. It is also important to develop algorithms and systems that reduce the computational costs and associated energy usage of limited resources while at the same time increasing the accuracy of model inference for various AI models. Summary of the Invention
[0008] Embodiments of the present disclosure provide a method, a non-transitory computer-readable storage medium, and a system for visualizing neurons of an AI model, the AI model being general purpose or specialized for a particular application scenario, such as decision making or more specifically for autonomous driving. In some embodiments, the method includes obtaining one or more neurons from a plurality of neurons for a task of the AI model, determining, for each neuron of the one or more neurons, a corresponding ROI of an input related to the task, where the corresponding ROI is encoded by the one or more neurons for the task, and applying a first operation including an LRP to generate an explainable AI-based representation of the determined corresponding ROI of the input for at least a portion of the one or more neurons.
[0009] In some embodiments, the explainable AI-based representation is a human-explainable representation or a machine-explainable representation.
[0010] In some embodiments, input is collected for the task via sensors, a recorded human driving database, and / or cloud storage.
[0011] Also, in some embodiments, the input is a processed image frame and a corresponding ROI for each of the one or more neurons includes a set of pixels of the processed image frame that corresponds to the corresponding ROI.
[0012] Also, in some embodiments, the input is a sequence of processed image frames, and the corresponding ROI for each of the one or more neurons comprises a union of each pixel set of the sequence of processed image frames, each pixel set corresponding respectively to a sub-ROI of a respective processed image frame of the sequence of processed image frames.
[0013] Additionally, in some embodiments, generating the explainable AI-based representation includes applying a second operation that includes VBP.
[0014] Additionally, in some embodiments, the VBP is applied sequentially after the LRP.
[0015] In some embodiments, the AI model also includes a mixing block and a model backbone.
[0016] Also, in some embodiments, applying the first operation and the second operation includes applying LRP through a mixing block to obtain a weight mask to be used for the feature map of the one or more neurons, weighting the feature map of the one or more neurons using the weight mask to obtain a weighted feature map of the one or more neurons, and applying VBP to back-transmit the weighted feature map of the one or more neurons through the model backbone.
[0017] Also, in some embodiments, the input is a spectrogram of an audio segment.
[0018] An embodiment of the present disclosure provides a non-transitory computer-readable storage medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to: obtain one or more neurons from a plurality of neurons for a task of an AI model, determine, for each neuron of the one or more neurons, a corresponding ROI of an input related to the task, where the corresponding ROI is encoded by the one or more neurons for the task, and apply a first operation including an LRP to generate, for at least a portion of the one or more neurons, a human-explainable representation of the determined corresponding ROI of the input.
[0019] In some embodiments, the explainable AI-based representation is a human-explainable representation or a machine-explainable representation.
[0020] In some embodiments, input is collected for the task via sensors, a recorded human driving database, and / or cloud storage.
[0021] Also, in some embodiments, the input is a processed image frame and a corresponding ROI for each of the one or more neurons includes a set of pixels of the processed image frame that corresponds to the corresponding ROI.
[0022] Also, in some embodiments, the input is a sequence of processed image frames, and the corresponding ROI for each of the one or more neurons comprises a union of each pixel set of the sequence of processed image frames, each pixel set corresponding respectively to a sub-ROI of a respective processed image frame of the sequence of processed image frames.
[0023] Additionally, in some embodiments, generating the explainable AI-based representation includes applying a second operation that includes VBP.
[0024] Additionally, in some embodiments, the VBP is applied sequentially after the LRP.
[0025] In some embodiments, the AI model also includes a mixing block and a model backbone.
[0026] Also, in some embodiments, applying the first operation and the second operation includes applying LRP through a mixing block to obtain a weight mask to be used for the feature map of the one or more neurons, weighting the feature map of the one or more neurons using the weight mask to obtain a weighted feature map of the one or more neurons, and applying VBP to back-transmit the weighted feature map of the one or more neurons through the model backbone.
[0027] Also, in some embodiments, the input is a spectrogram of an audio segment.
[0028] An embodiment of the present disclosure also provides a computer-implemented system including one or more processors and one or more memory devices that store commands that, when executed by the one or more processors, cause the one or more processors to: obtain one or more neurons from a plurality of neurons for a task of an AI model, determine, for each neuron of the one or more neurons, a corresponding ROI of an input related to the task, where the corresponding ROI is encoded by the one or more neurons for the task, and apply a first operation including an LRP to generate a human-explainable representation of the determined corresponding ROI of the input for at least a portion of the one or more neurons.
[0029] In some embodiments, generating the explainable AI-based representation includes applying a second operation that includes VBP.
[0030] All combinations of the above concepts and additional concepts described in more detail herein are to be considered as part of this disclosure, for example, all combinations of claimed subject matter appearing at the end of this disclosure are to be considered as part of the subject matter disclosed in the body. [Brief description of the drawings]
[0031] In order to more clearly illustrate the embodiments of the present disclosure and related art, the figures described in the following embodiments are briefly introduced. Obviously, these figures are only some embodiments of the present disclosure, and those skilled in the art can obtain other figures without creative effort by referring to these figures. Arrows in the figures indicate relationships, by which the components where the arrows start can be used to train / apply the components where the arrows point. These figures can be used in combination with the following detailed description to more fully understand the embodiments of the present disclosure. [Figure 1A] FIG. 2 is a block diagram illustrating an example of an AI model according to some embodiments of the present disclosure. [Figure 1B] FIG. 1 is a block diagram illustrating an example of a trained AI model suitable for performing a method for visualizing a model, according to some embodiments of the present disclosure. [Diagram 2] FIG. 1 is a block diagram illustrating an example of a proposed computer-implemented system according to some embodiments of the present disclosure. [Figure 3A] FIG. 2 illustrates an example of an ROI in an input image, according to some embodiments of the present disclosure. [Figure 3B] FIG. 2 illustrates an example set of sub-ROIs in an input image sequence, according to some embodiments of the present disclosure. [Figure 4] 1A and 1B are diagrams illustrating example spectrograms in three-dimensional (3D) and two-dimensional (2D) form for visualizing neurons in an AI model, according to some embodiments of the present disclosure. [Diagram 5] 1 is a flowchart of an example of visualizing the operation of neurons in an AI model, according to some embodiments of the present disclosure. [Figure 6] 11 is a flowchart of another example of visualizing the operation of neurons in an AI model, according to some embodiments of the present disclosure. [Figure 7] 1 is a flowchart of an example of applying a visualization technique based on two saliency maps, according to some embodiments of the present disclosure. [Figure 8]FIG. 1 is a diagram of an example of generating a human-explainable representation according to some embodiments of the present disclosure. [Figure 9] FIG. 11 is another example of generating a human-explainable representation according to some embodiments of the present disclosure. [Figure 10] FIG. 11 is a diagram of yet another example of generating a human-explainable representation according to some embodiments of the present disclosure. [Figure 11] FIG. 1 illustrates an example hardware and software environment for an autonomous vehicle, according to some embodiments of the present disclosure. For simplicity and clarity of the drawings, the components illustrated in the figures are not necessarily drawn to scale. For example, the size of some components may be larger than others for clarity. Also, where appropriate, additional graphical symbols may be repeated to indicate corresponding or similar components. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0032] With reference to the drawings, the embodiments of the present disclosure are described in detail with respect to technical problems, structural characteristics, objectives to be achieved, and effects. Specifically, the terms in the embodiments of the present disclosure are used only to describe specific embodiments, and do not limit the present disclosure. In the following detailed description, many specific details are described in order to fully understand the present invention. However, it should be understood that those skilled in the art can practice the present invention without these specific details. In other circumstances, known methods, processes, and components are not described in detail to avoid confusion. The subject matter of the present invention is specifically pointed out and explicitly protected in the last part of this specification. However, the organization, operation method, objectives, features, and advantages of the present invention can be best understood by referring to the following detailed description and reading it in combination with the drawings. Because the illustrated embodiments of the present invention can be realized mainly using electronic components and circuits that are already known to those skilled in the art. Therefore, no more details than are deemed necessary to understand the basic concept of the present invention and not to confuse or scatter the teachings of the present invention will be described. For example, the specification and / or drawings may involve a processor or a processing circuit. The processor may be a processing circuit. The processing circuitry may be implemented as a central processing unit (CPU) and / or one or more other integrated circuits, such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a fully custom integrated circuit, or a combination of these integrated circuits.
[0033] The following specification and / or drawings may involve images and / or image frames. An image is an example of a media unit. Any reference to an image may be applied to a media unit as appropriate. A media unit may be an example of a sensing information unit (SIU). Any reference to a media unit may be applied to any type of natural signal as appropriate, such as, but not limited to, naturally generated signals, signals representing human actions, signals representing actions related to vehicular signals, geodetic signals, geophysical signals, text signals, digital signals, time series signals, etc. Any reference to a media unit may be applied to an SIU as appropriate. An SIU may be of any type and may be sensed by any type of sensor, for example, a vision camera. Acoustic sensors may sense infrared, radar imaging, ultrasonic, electro-optical, radiography, light detection and ranging (LIDAR), thermal sensors, passive sensors, active sensors, etc. Sensing may include generating samples (e.g., pixels, audio signals, etc.) representative of a signal transmitted or otherwise reaching the sensor. An SIU may contain one or more images, one or more video clips, text information about one or more images, or text describing motion information, and so on.
[0034] Any combination of any modules or units listed in any of the accompanying drawings, any part of the specification, and / or any claim may be provided. Any of the units and / or modules illustrated in this application may be realized using codes, commands and / or instructions stored in hardware and / or non-transitory computer readable media and may be included inside a vehicle, outside a vehicle, on a mobile device, on a server, etc. The vehicle may be any type of vehicle, such as, for example, a ground transportation vehicle, an air vehicle, or a water transportation tool. The vehicle may also be referred to as a private vehicle. It should be understood that autonomous driving includes at least partially autonomous (semi-autonomous) driving of a vehicle, including all L2 level types or higher level types defined by the SAE standard.
[0035] As used herein, an AI model may be general purpose or dedicated to a particular application scenario, e.g., decision making, classification, prediction, etc. In particular, an AI model may be customized for common tasks related to autonomous driving. These tasks can be categorized, for example, into perception, localization and mapping, planning and decision making, and control. Perception tasks include accurate detection and identification of objects and entities in the surrounding environment. This includes identification and classification of pedestrians, vehicles, traffic signs, traffic lights, and other relevant objects. Localization tasks are centered on determining the exact location of the vehicle in the surrounding environment, which involves using sensors and data to estimate the vehicle's position relative to known reference points or a map. Mapping tasks, on the other hand, are related to creating and updating a representation of the surrounding environment. Both localization and mapping enable the autonomous driving system to know the exact location of the vehicle and navigate the vehicle efficiently. Planning tasks are related to generating a sequence or trajectory of actions based on the vehicle's current location and the desired destination. Decision-making tasks involve analyzing the current driving situation and determining the appropriate action, such as changing lanes, accelerating, braking, or yielding. Planning and decision-making together enable an autonomous driving system to navigate the vehicle in a safer and more efficient manner. The control task typically involves executing the planned actions and adjusting the vehicle dynamics to follow the desired trajectory. This includes controlling the steering, acceleration, and braking systems to maintain proper control and stability of the vehicle. The control task ensures that the vehicle's physical responses are consistent with the planned actions.
[0036] This disclosure proposes a method for visualizing neurons in an AI model. To achieve this, we need to simplify the millions of neurons in the AI model as a prerequisite to obtain an equivalent and compact representation of the entire set of neurons in the model. This makes it easier to intuitively understand the portion of the model input that each of the limited number of neurons focuses on to complete a specific task. The simplified representation of neurons provides an intuitive understanding of the response of the neural network to the model input under a given task. The simplified and compact representation of neurons allows for a more focused analysis of the model's behavior and the contribution of neurons to the overall functionality of the AI model.
[0037] As used herein, the term ROI refers to a unique aspect of the model input encoded by a neuron in the compact representation of the neuron for a particular task for which the AI model was specifically trained. For example, in lane change tasks related to autonomous driving, identification of road boundaries is crucial because collisions with road boundaries can cause serious accidents, such as rolling over or damage to the vehicle during lane change. Thus, for lane change tasks, at least some neurons (also referred to as active neurons) of the compact representation of the neurons of the AI model focus on the parts of the model input that represent lane boundaries or contain information about lane boundaries. Then, each of the active neurons can encode the corresponding parts of the model input based on the determined corresponding ROI to realize the overall functionality of the AI model (i.e., performing lane changes for the vehicle). This is necessary to prevent the vehicle from overlapping or colliding with any detected lane boundaries in the model input. Finally, a first operation is applied to generate a human-explainable representation from which an end user and / or model developer can obtain an intuitive understanding of the underlying behavior of the AI model. Such manipulation employs LRP, a technique for highlighting the contribution of each neuron in encoding different aspects of the overall input to complete a task. By leveraging LRP, the method generates representations that are easily understandable and explainable to humans. In this manner, the disclosed method provides a valuable way to visualize and understand the function of neurons in an AI model customized for a particular application scenario (e.g., autonomous driving). Also, focusing on the specific contribution of each neuron allows for the allocation of computational resources to the particular neurons or nodes that contributed most to completing a particular task, thereby improving computational efficiency.
[0038] That is, based on the generated representation, a user, model developer, or the trained model itself can intuitively understand which particular portions of the model inputs each active node in the AI model is responsible for or focuses on in a given task, which nodes are not involved in processing the model inputs (and potentially are not involved in the model's decision-making, such as model inference), which nodes are most interested in the portions of the model inputs that are relevant to the model's decision-making (and which nodes are less critical to these portions), etc. Thus, based on the generated human-readable representation, the user, model developer, or trained model can modify and / or fine-tune the network structure of the trained AI model for a given task, for example, by activating or removing nodes that are less relevant to the model's decision-making, thereby saving computational resources and improving computational efficiency. Furthermore, using the human-readable representation generated by the methods and systems disclosed in the present application, the model developer or trained model can inspect whether to modify the type, number, format, etc. of the model inputs to better facilitate the generation of accurate model decisions, thereby improving the accuracy of model inference of the model inputs in a given task and enhancing the safety and reliability of using such AI models in applications such as autonomous driving systems.
[0039] In addition, with the increase in human explainability of AI models, users and / or trainers can provide more accurate feedback to AI models as reference data. Such high-quality reference data reduces the total amount of data required for an AI model to achieve a working model. This means that the AI model needs to train parameters in less time and with fewer computational resources to achieve the working model.
[0040] It should be noted that the disclosed method does not necessarily need to be involved in the training process of the AI model, thus avoiding the need for increased computational resources. The flexibility of the disclosed method is a notable feature, since it does not depend on the specific complexity or realization details of the AI model. Thus, it can be effectively applied to a wide range of AI models, regardless of their architecture, size, or complexity. Scalability ensures compatibility with various types of AI models, including neural networks, deep learning models, reinforcement learning models, or any other form of machine learning algorithms. Overall, the disclosed method provides a resource-efficient and scalable way, which mitigates the need for additional computational resources and ensures compatibility across various types of AI models by avoiding dependency on specific details of the model, further adding value to practical applications and obtaining insights from the AI models.
[0041] Referring now to the drawings, in which like numerals in all accompanying drawings indicate like elements, FIG. 1A is a block diagram illustrating an example of an AI model 100 according to some embodiments of the present disclosure. As shown in FIG. 1A, the AI model 100 may include a model backbone 102, a mixing block 103, a policy head 104, a latent layer 105, and a number of neurons 106.
[0042] The model backbone 102 constitutes the foundational part of the AI model 100 and is responsible for initial data processing. In some examples, the model backbone 102 may include various layers and modules designed to extract and transform information contained in the model input. It captures, extracts, and classifies the necessary features and representations from the large amount of model input data 101 (e.g., road frontal images or videos, or lateral acceleration annotations). These necessary features and representations are necessary for subsequent analysis and decision making within the AI model. In an embodiment, the model backbone 102 may be a convolutional neural network (CNN) that learns different features (e.g., road tracks and curves).
[0043] The mixing block 103 can integrate and combine information from different parts (e.g., layers) of the model backbone 102. It enhances the overall representation of the input data by facilitating information exchange and feature fusion between model inputs. The mixing block ensures effective sharing and utilization of relevant information, improving the overall performance and accuracy of the AI model. In an embodiment, the mixing block 103 may be a multi-layer perceptron (MLP) that includes a channel mixing MLP that enables communication between different channels and a token mixing MLP that enables communication between different spatial locations. These layers are interleaved (i.e., combined) to realize the interaction of both types of inputs.
[0044] The policy head 104 is a component that represents the development policy and generates the final output or realizes the decision based on the analysis of the processed input data. The policy head 104 further provides a higher level of understanding of the input data. In other words, the policy head dictates the action to be taken based on the state of the deep learning model and the detected surrounding environment. In an embodiment, the policy head 104 may be a trainable AI model.
[0045] The latent layer 105 is a simplified or condensed representation of the model input data 101 that may include an outline of important features of the model input data 101 (e.g., features related to lane boundaries). In some embodiments, the latent layer 105 can be obtained by discarding duplicate or irrelevant data using different data representation and approximation techniques. This allows for the transfer of less data without loss and allows for the transfer of compact models rather than the original data. In this way, computational efficiency can be improved since less data needs to be processed and transferred from one domain to another, and model accuracy can be maintained without loss.
[0046] The latent layer 105 can include multiple neurons 106, with each neuron dedicated or focused to capture and process a particular input feature or pattern for a given task. In some examples, the latent layer 105 is used as a compact representation of the entire set of neurons in the AI model. That is, from a human explanation perspective, the number of neurons in the latent layer 105 is limited and acceptable. Thus, the collective operation of these neurons 106 can contribute to the overall processing of input data in the AI model 100 to complete a task.
[0047] During operation, the AI model 100 receives and processes model inputs 101 to generate model outputs 107. An example of the model input 101 may be an image signal depicting an image of a front view of a road, as shown in thumbnail 101a in the figure. However, those skilled in the art will appreciate that there may be other suitable forms of model inputs, such as audio signals, text annotations, or a combination of audio and image signals (e.g., video streams) and text annotations. In some embodiments, the model input 101 is raw data from one or more sensors of the same or different vehicles. For example, the model input 101 may be an image including red, green, and blue (RGB) values of pixels captured by a camera sensor. The model input 101 may be raw SIUs, processed SIUs, text information, information derived from SIUs, etc. In different embodiments, the model input 101 may be loaded from a local disk, from a remote storage location over a suitable "cloud" network, etc. The acquisition model input 101 may include receiving data, generating data, participating in processing of data, processing only a portion of the data, and / or receiving only another portion of the data. The processing of the model input 101 may include at least one of detection, noise reduction, improving the signal-to-noise ratio, defining a bounding box, etc. The model input 101 may be received from one or more sources, such as one or more sensors, one or more communication units, one or more memory units, one or more image processors, etc.
[0048] The model backbone 102 extracts features from the received model input 101, such as road curvature, lane markers, etc., contained in the image, and passes the extracted features to a mixing block 103, where these features are combined with other layers to reduce these features from high-dimensional model input data to low-dimensional latent vectors as a compressed latent layer 105. In this way, the data volume / complexity of the raw data 102 can be reduced by forming the compressed latent layer 105. Such compression further improves computational efficiency by requiring less data to be processed as well as training, as described in more detail below.
[0049] The latent layer 105 of the model input 101 helps to learn data characteristics and simplify data representation. Each data feature is stored as a separate neuron 106. The policy head 104 receives the latent layer 105 and processes the information provided by the latent layer. This information may include the curvature of the road, lane markers, and the current position of additional vehicles relative to the road, the current speed and lateral acceleration of the vehicle, whether there are other vehicles nearby, etc., and outputs the model output 107 based on the processed information. In an embodiment, the model output 107 may include an output maneuver decision to turn the steering wheel 107a to increase the lateral acceleration and keep the vehicle centered in the curved lane.
[0050] In some embodiments, the model backbone 102 and the mixing block 103 may be configured to map the model inputs 101 to a latent layer 105, which may be stored in a database of semantic relations. In some embodiments, the model backbone 102 learns to reduce the dimensionality of the input data to encode latent representations of features, while the policy head 104 recreates the encoded latent representations into a reconstructed output, such as the model output 107. For example, the model backbone 102 may be configured to generate a compressed latent layer 105 of the model inputs 101 using a one-dimensional vector representing one or more elements of the model inputs 101. In one embodiment, the compressed latent layer 105 may be represented as a vector V, where V=[E1, E2, E3, ...EN], where E1 refers to element 1, E2 refers to element 2, E3 refers to element 3, and EN refers to element N. Each element may be a one-dimensional or multi-dimensional matrix. Each element can represent a potentially useful feature around the vehicle, such as lane markings, lane centerlines, nearby vehicles, traffic signs, tree outlines, etc.
[0051] The model backbone 102 may be configured to encode meaningful information about various data attributes in its latent manifold, and these meaningful information can then be used to perform related tasks. In such an embodiment, the latent layer 105 contributes to reducing the dimensionality of the input data and removing irrelevant information. Thus, reducing the dimensionality of the input data can reduce computational costs, since fewer computer resources need to be allocated and processed for the reduced complexity and volume of the input data. Also, the accuracy of the model can be increased, since irrelevant information that may distort the modeling is removed.
[0052] In some embodiments, given the latent layer 105, the policy head 104 may be configured to determine an action that the vehicle should follow from a set of predefined tasks. The tasks determine the action that the autonomous vehicle should take based on the latent layer 105. Some examples of these tasks are lane keeping, overtaking, lane changing, intersection handling, signal handling, etc.
[0053] The model output 107 may represent actions to be performed in the context of a particular application scenario (e.g., autonomous driving), such as the operation of an accelerator pedal, a brake pedal, or a steering wheel, which is represented by a thumbnail 107a in the figure. Although the components shown in Figure 1 are shown to constitute an AI model, it can be readily understood that an AI model customized for a particular application scenario (e.g., decision-making, autonomous driving, etc.) may also be generalized to include similar components as shown in Figure 1A.
[0054] FIG. 1B is a block diagram showing an example of a trained AI model suitable for performing a method for visualizing a model according to some embodiments of the present disclosure. The main difference between FIG. 1B and FIG. 1A is that, as shown in FIG. 1B, the AI model 100 has already completed the training process, so all parameters of the AI model are determined. Therefore, in FIG. 1B, the connections between each component are intentionally omitted, while in FIG. 1A, the data stream during the model training process is illustrated. As shown in FIG. 1B, the arrows can indicate operations that will be described in more detail below. In an embodiment, the operation performed is LRP.
[0055] LRP is a technique used in the fields of artificial intelligence and deep learning to understand the contribution and relevance of input features to model output. LRP allows neural network models to be interpreted and analyzed by back-transmitting relevance scores through each layer of the network. In particular, LRP assigns relevance scores or weights to the model's output neurons (i.e., neural activations) and then manipulates these scores by transmitting them back through layers or model components, which in this disclosure may optionally be neurons 106 in latent layer 105. This back-transmitting process aims to highlight the importance of different input features and their classification that neurons, such as neuron 106, use in creating the model's decision-making process. By applying LRP, a human-explainable representation can be generated that highlights the regions from the neuron's perspective of the input data that are most relevant to the model's decision (i.e., visualizing the input pixels that truly contribute to the model's decision). Thus, limited computer resources can be further allocated to inputs that are considered most important to the model's decision-making for more intensive processing based on LRP application, thereby confirming and further improving computational efficiency. LRP has a wide range of applications including image classification, natural language processing, and other fields that leverage AI models.
[0056] As shown in FIG. 1B, the human-explainable representation 109 may be a graphical user interface (GUI) representation obtained by applying an LRP operation to the model input 101 of FIG. 1. An example of such a representation is illustrated by thumbnail 109a in FIG. 1B, which shows a mapping between one active neuron (or a group of active neurons) selected from among the active neurons of the AI model 100 and the relevant parts in the model input for performing a given task. Another example is thumbnail 109b, where the given task used for this thumbnail may be voice-related, such as voice control during autonomous driving. The salient parts outlined in thumbnail 109b in the model input (e.g., a 2D representation of a spectrogram of a voice) may be associated with the corresponding active neuron, thereby reflecting the contents encoded by the latent layer to complete the voice-related driving task. In the following, an example of the human-explainable representation 109 is described in detail.
[0057] 2 is a block diagram illustrating an example of a proposed computer-implemented system according to some embodiments of the present disclosure. As shown in the figure, the proposed computer-implemented system may include a processor 200. The processor 200 may be a general-purpose processor or a special-purpose processor, such as an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), an SOC (System on Chip), a CPLD (Complex Programmable Logic Device), etc.
[0058] As shown, the processor 200 may include an acquisition module 216 , a determination module 217 , an LRP 218 , a VBP 219 , a representation generation module 220 , and a visualization engine 221 .
[0059] The acquisition module 216 may be configured to receive model information 223 of the AI model from the network 222 that deploys the AI model, and acquire knowledge of the entire neuron set of the AI model from the received model information 223. The acquisition module 216 may then be configured to acquire a compact representation to be used as the entire neuron set from a plurality of neurons of the entire neuron set of the AI model for a given task, and acquire one or more neurons, denoted as neuron information 224. In some implementations, the acquisition module 216 may selectively acquire different sets of one or more neurons from the entire neuron set of the AI model depending on different upcoming tasks. These acquired neurons form a compact representation that focuses on relevant aspects of the input required for the given task.
[0060] The processor 200 may also include a determination module 217. The determination module 217 may be configured to determine, for each of the one or more neurons acquired as represented by the neuron information 224, a corresponding ROI of input related to a predetermined task based on the received neuron information 224. As shown in the figure, the determination module 217 may include an LRP unit 218 configured to apply an LRP operation to the model input. As a non-limiting example, the model input may be a processed signal 215. The processed signal 215 may be output from a signal processor 214, which may be separate from the processor 200 shown in FIG. 2. The signal processor 214 may be, for example, a digital signal processor (DSP). In some examples, the processed signal 215 is output from the signal processor 214 in response to the received raw unprocessed signal 213. The raw unprocessed signal 213 may be acquired from one or more sensors 210, a recorded human driving database 212, or a network 222.
[0061] Optionally, the decision module 217 may also include a VBP 219. VBP is a technique commonly used in the fields of computer vision and deep learning to gain insight into the image regions that contributed most to the model's prediction. VBP manipulates gradients by sending them from the output layer (which may alternatively be the latent layer 105 in Figs. 1A and 1B) back to the input layer, thereby attributing a relevance score to each pixel or region along the way. These relevance scores indicate the importance of this particular pixel or region in contributing to the model's decision-making. By mapping these relevance scores to the input image, VBP facilitates the generation of visually explainable hot or salient maps, making it a powerful tool for visualizing and interpreting neural network models, and verifying whether the prediction results output by the model match the true values obtained. That is, by highlighting important image regions, VBP provides intuitive insight into the decision-making process, thereby contributing to the transparency and explainability of models used in vision-related tasks. In this way, the model accuracy of neural network models can be confirmed and increased. In addition, limited computer resources can be allocated to the specific pixels or regions that contribute most to the model's decision-making, while eliminating resources that are ineffectively allocated to unnecessary input processing, thereby improving computational efficiency.
[0062] In summary, the determination module 217 in the processor 200 enables identifying and determining task-relevant ROIs in the processed signal 215 for one or more relevant (i.e., active) neurons of the latent layer based on the received neuronal information 224, while improving the accuracy of the model and computational efficiency.
[0063] The processor 200 may also include a representation generation module 220. In some examples, the representation generation module 220 may include a visualization engine 221 responsible for generating a visualization output 226. The visualization output 226 may be a human-interpretable representation of the corresponding ROI (e.g., the processed signal 215) determined for at least some of the model inputs of the acquired one or more neurons. As a non-limiting example, the visualization output 226 may be displayed in a GUI 230 as shown in FIG. 2. An example of the output in the GUI may be represented by a thumbnail 231, which is described below with reference to FIGS. 8-10.
[0064] 3A is a diagram illustrating an example of an ROI in an input image, according to some embodiments of the present disclosure. As shown, the processed signal 215 in FIG. 2 may be denoted by reference numeral 330. The processed signal 215 may be an image signal and may undergo image processing such as grayscaling or cropping. The processed signal 330 may include an ROI 332 including a rectangular region of pixels represented by reference numeral 334, whose horizontal range extends from x0 to x1 and whose vertical range extends from y0 to y1. In an embodiment, an identifier of the ROI in the processed signal 330 may be binary, either because the processed signal falls within the ROI or because the signal falls outside the ROI.
[0065] FIG. 3B is a diagram illustrating an example of a set of sub-ROIs in an input image sequence, according to some embodiments of the present disclosure. The task performed by the AI model requires multiple inputs, not just a single input. For example, for an overtaking task, the AI model requires a sequence of image frames to accurately assess the movement of other vehicles and / or moving objects surrounding the ego-vehicle. Thus, FIG. 3B is intended to illustrate such an application scenario. In FIG. 3B, an example is shown in which an image frame sequence 331 is fed to the AI model to analyze the surrounding environment during the overtaking process. The sequence of image frames may, for example, include three consecutively processed images 3301-3303 that are captured and processed in chronological order. Each of the processed images 3301-3303 may include a corresponding sub-ROI. As a non-limiting example, the sub-ROI in the processed image 3301 includes a rectangular region of pixels, represented by reference numeral 3321, whose horizontal extent extends from x0 to x1 and whose vertical extent extends from y0 to y1. The sub-ROI in processed image 3302 includes a rectangular region of pixels represented by reference numeral 3322, whose horizontal range extends from x2 to x3 (where x2 is greater than x0 and x3 is greater than x1) and whose vertical range extends from y0 to y1. The sub-ROI in processed image 3303 includes a rectangular region of pixels represented by reference numeral 3323, whose horizontal range extends from x2 to x3 and whose vertical range extends from y2 to y3 (where y2 is greater than y0 and y3 is greater than y1).
[0066] Thus, the ROI of the processed image frame sequence 331 may be represented by reference numeral 3340, with its pixels extending horizontally from x0 to x3 and vertically from y0 to y3. It can be understood that the number of image frames included in the image frame sequence 331 may be any suitable number, and the disclosure is not limited in this respect. It can also be understood that the ROIs (regions of interest) and sub-ROIs (sub-regions of interest) shown in Figures 3A and 3B are depicted for illustrative purposes only. In most cases, the ROIs and sub-ROIs may have irregular shapes. Thus, the disclosure does not limit the shape of the ROIs and / or sub-ROIs.
[0067] FIG. 4 illustrates examples of spectrograms in three-dimensional (3D) and two-dimensional (2D) form, according to some embodiments of the present disclosure. A spectrogram is a graphical representation of the frequency content of a signal as it changes over time. Spectrograms are used in signal processing, such as audio and speech analysis. As shown at the bottom of FIG. 4, an exemplary 2D spectrogram 434 plots the spectrum of an acoustic signal (e.g., an audio signal obtained from one or more microphones placed in a vehicle cabin) on the y-axis and time on the x-axis. The intensity (or color) of each point in the 2D spectrogram 434 represents the strength or amplitude of the frequency components of the acoustic signal at a particular time. This provides a visual representation of how the frequency content of a signal changes over time, thereby enabling analysis and identification of various audio features, such as harmonics, resonant peaks, and transient events.
[0068] Also, a 3D spectrogram extends the concept of a 2D spectrogram by adding an additional third dimension (i.e., the intensity or amplitude of the frequency components plotted in the third dimension, and the intensity or amplitude of the frequency components represented by the intensity or color of the corresponding portion in the 2D). The third dimension of a 3D spectrogram can be visualized as a surface plot or a contour plot, where the height or color of the surface / contour represents the amplitude of the frequency components at a particular time and frequency. In FIG. 4, an exemplary 3D spectrogram 430 corresponding to the 2D spectrogram 434 is shown at the top of FIG. 4 with reference numeral 430.
[0069] FIG. 4 provides a visual representation of a spectrogram capturing a particular fragment of an exemplary acoustic signal in both 2D and 3D formats. As shown, a distinct peak-shaped region circled and labeled ROI 422 in 3D spectrogram 430 corresponds to a region labeled ROI 432 in 2D spectrogram 434. The content within the corresponding ROI 432 in 3D spectrogram 430 or 2D spectrogram 434 may be related to voice utterances generated by a human (e.g., a vehicle driver), while other parts of 3D spectrogram 430 or 2D spectrogram 434 may include other components such as mechanical noise during vehicle operation, road environment noise, and noise caused by other passengers in the vehicle cabin. In practical applications, particularly in scenarios related to voice control functions for autonomous driving, ROI 432 represents a salient portion within the input acoustic data of an AI model customized for such application scenario. That is, the salient ROI 432 is an area of particular interest to the active neurons in the model's latent layer, since it plays an important role in identifying and executing voice commands relevant to the autonomous driving task.
[0070] In an embodiment, an ROI may be over-created to include multiple features important to task salience, and a 2D spectrogram 434 or 3D spectrogram 430 may be constructed within the identified ROI 432, further mapping the "interest level" or input relevance of each pixel to a potential neuron on a continuous scale, where the color in the 2D spectrogram 434 or the height in the 3D spectrogram 430 indicates the interest level or salience of the potential neuron.
[0071] Referring back to FIG. 1B, the illustrated example embodiment of the human-readable representation 109 (i.e., thumbnail 109b) corresponds to the 2D spectrogram representation 434 shown in FIG. 4. As shown in FIG. 4, when the 2D spectrogram representation 434 is used as the human-readable representation, the contour of the ROI 432 can indicate the time-frequency components of the audio signal (e.g., collected by a microphone or microphone array in the vehicle cabin) that are focused on / interested by the active node and are encoded by the active node for model determination. By viewing a human-readable representation such as the 2D spectrogram representation 434, a user, model developer, or the model itself can identify whether the vocal component of the collected audio signal is sufficiently prominent (e.g., whether it occupies a sufficient area in the spectrogram) to adjust the settings of the audio signal detection device (e.g., microphone) and enable the model to extract a useful information payload, thereby increasing the accuracy of the model inference. Alternatively, if active nodes are concentrated in portions of the model input that are significantly outside the ROI, a user, model developer, or model can save unnecessary computational resources and improve computational efficiency for model inference by deactivating or removing these nodes from the network structure of the AI model.
[0072] 5-6, which illustrate methods 500 and 600 corresponding to the methods and models used to visualize neurons in an AI model, as described above. It should be noted that the order of methods 500 and 600 is exemplary and does not indicate the order in which the steps of methods 500 and 600 should be performed.
[0073] 5, the method 500 begins by obtaining, at 502, one or more neurons from a plurality of neurons for a task of an AI model. Next, for each neuron of the one or more neurons, a corresponding ROI of an input related to the task is determined, at 504, where the corresponding ROI is encoded by the one or more neurons for the task. Thereafter, at 506, a first operation including an LRP is applied to generate a human-interpretable representation of the determined corresponding ROI of the input for at least a portion of the one or more neurons.
[0074] In some implementations of the method 500, the ROI can be determined by a visualization technique based on a saliency map (e.g., LRP). Preferably, the ROI can be determined by a combination of visualization techniques based on a saliency map, such as a combination of LRP and VBP. In some embodiments, to facilitate human explanation and post-hoc explainability of the AI model, the relationship between the human-explainable representation and the determined ROI can be visualized by a mapping between the determined ROI of the associated neuron and the associated correlated neuron of one or more neurons, e.g., obtained via a GUI. As a non-limiting example, for a given task (e.g., a lane change task), for example, in the latent layer (i.e., the compact form of the entire neuron set) of the AI model, there may be two active neurons associated with this task, one used to encode the leftmost lane boundary and the other used to encode the rightmost lane boundary, see, for example, FIG. 8 of the accompanying drawings. Here, neurons 8201 and 8202 encode two lane boundaries 808, respectively. Thus, a human-explainable representation can be displayed via a GUI as, for example, (i) a mapping between two active neurons and the two lane boundaries that the two active neurons focus on, or (ii) a mapping between one neuron and the corresponding lane boundary that this neuron focuses on.
[0075] Also, as mentioned above, the number of neurons in the entire neuronal set of an AI model may be huge, and thus beyond human explanation. Therefore, in order to simplify a black-box AI model, it is necessary to obtain a compact form of the entire neuronal set. The compact form is one or more neurons obtained, which can be related to a normal driving task, e.g., each of the obtained one or more neurons can encode a part of the model input related to a given task. Thus, the underlying logic of generating a human-explainable representation of multiple neurons includes two aspects. One is to present a compact form of the entire neuronal set of an AI model, rather than hundreds or thousands or millions of neurons of the model, and the other is to show a mapping between a selected number of neurons among the obtained one or more neurons and the content that the selected neuron(s) are paying attention to / interested in in the input, so that an end user and / or model developer can understand the role of each neuron (in a compact form, e.g., within the latent layer of the representation of the AI model) in encoding the model input and in influencing the model's decision-making. By presenting the neurons in a compact form, computational efficiency is improved since less processing is required. Also, the accuracy of the model can be confirmed and improved by focusing on the inputs that contributed most to the assignment of neurons to process the inputs.
[0076] In some examples, the method of 600 illustrates another possible embodiment for visualizing the operation of neurons in an AI model. For example, the method 600 begins by obtaining, at 602, one or more neurons from a plurality of neurons for a task of the AI model. Next, at 604, for each neuron of the one or more neurons, a corresponding ROI of an input related to the task is determined, where the corresponding ROI is encoded by the one or more neurons for the task. Thereafter, at 606, a first operation including LRP and a second operation including VBP are applied to generate a human-interpretable representation of the determined corresponding ROI of the input for at least a portion of the one or more neurons.
[0077] In some embodiments, as shown in FIG. 7, the method 700 shows the operation of applying the first operation and the second operation as shown in block 606 of FIG. 6. For example, the method 700 starts by applying LRP through a mixing block of an AI model at 702 to obtain a weight mask for a feature map of one or more neurons. Then, at 704, the obtained weight mask is used to weight the feature map of one or more neurons to obtain a weighted feature map of one or more neurons. Then, VBP is applied to transmit the weighted feature map of one or more neurons back through the model backbone. Combining LRP with the mixing block and VBP back transmission further enhances the insight gained from the visualization of the neurons that contributed most to the prediction made and the regions of interest of the input image that contributed most to the prediction made, respectively. The mapping from the entire image data processing to the predicted progression path can be performed together, from which biases and errors can be recognized. Thus, the accuracy of the model can be verified and enhanced, errors in prediction can be more easily identified and corrected, and computational efficiency can be improved when accurately allocating limited computer resources to process the regions of interest and neurons that have the greatest impact on the prediction from the neural network model.
[0078] 8-10 are diagrams illustrating different examples of operating an autonomous driving system for different tasks, showing example representations of raw data 802, 902, and 1002, example representations of model inputs 804, 904, and 1004, and example human-explainable representations 805, 807, 905, and 1005, 1007, respectively.
[0079] As shown in FIGS. 8-10, raw data 802, 902, and 1002 may be captured by a camera disposed in the vehicle. The camera may be configured to capture real-time images from a front-view perspective of the vehicle cabin. In some embodiments, model inputs 804, 904, and 1004 may be compact representations of input images (i.e., raw data 802, 902, and 1002) that capture useful features via a dedicated processor (e.g., signal processor 214 shown in FIG. 1). As shown in FIGS. 8-10, raw data 802, 902, and 1002 may be RGB images and can carry rich information. For example, raw data includes not only images of the road but also images of the scene around the road (e.g., other vehicles, trees, traffic signs, sky, etc.). In comparison, in some embodiments, model inputs 804, 904, and 1004 retain only useful features and are converted to grayscale images to facilitate storage and processing of the AI model. For example, as shown in Figures 8-10, model inputs 804, 904 and 1004 may be converted from colored raw data 802, 902 and 1002 into grayscale images including tree contours 806, 906, 1006, lane boundary lines 808, 908, 1008, lane centerlines 810, 910, 1010 (e.g., first lane centerlines 810a, 910a, 1010a and second lane centerlines 810b, 910b, 1010b), traffic sign contours 812, 912, 1012, traffic sign text 814, 914, 1014, and other vehicles 816, 916, 1016.
[0080] In FIG. 8, an exemplary driving task processed by an AI model embedded or integrated in a vehicle may be a lane change task. Here, two exemplary human-explainable representations 805 and 807 are provided for illustrative purposes only. Each of the human-explainable representations 805 and 807 includes a schematic representation of a latent layer of an AI model. As described above, for the purpose of simplifying the black-box model, the latent layer is a compact form of the entire neuron set of the AI model, which may be equivalent to the entire neuron set of the AI model. As shown in the representation 805, two neurons 8201 and 8202 in the exemplary latent layer are activated under the lane change task. A rectangular region 818 corresponding to the model input 806 or raw data 802 is placed in the representation 805 highlighting the corresponding ROIs of the model inputs determined for the two active neurons (i.e., the left lane boundary and the right lane boundary 808). In some embodiments, a user or developer can select a portion of the neurons in the exemplary latent layer in the GUI of the human-explainable representation to get an intuitive impression or understanding of what each active neuron in the model input is focusing on / encoding for a given task. For example, in representation 807, only the right lane boundary 808 is highlighted within the rectangular region 819 associated with the performance of the lane change task to which the selected neuron 8201 is focused.
[0081] 9 and 10, in FIG. 9, the exemplary driving task may be a lane center task, for which the lane center line 908 is enhanced into an ROI in a rectangular region 918, which indicates the portion of the model input for the lane center task that the active neurons 903 are focusing on / encoding. In FIG. 10, the exemplary driving task may be a perception task, under which the vehicle is required to identify the surrounding environment to obtain useful context information. As shown in the representation 1005, two neurons 1021 and 1022 in the exemplary latent layer are activated under the perception task. The rectangular region 1018 corresponding to the model input 1006 or raw data 1002 is placed in the representation 1005 to highlight the corresponding ROIs (i.e., traffic sign contour 1012 and traffic sign text 1014) of the model input determined for the two active neurons. Alternatively, in the representation 1007, only the traffic sign contours 1012 are highlighted within the rectangular region 1019 that is relevant to the performance of the perceptual task to which the selected neuron 1021 is focused.
[0082] As shown in FIGS. 8-10, by looking at such exemplary human-readable representations (e.g., 805, 807, 905, 1005, 1007), a user, model developer, or the model itself can determine whether image portions relevant to decision-making for a given task are sufficiently salient to the model inputs (e.g., whether these image portions are processed by active nodes, have enough corresponding pixels, are fully encoded by active nodes, etc.), and can then adjust the number (e.g., single image frames vs. a series of image frames), type (e.g., driver's field of view or wide-angle field of view), or format (e.g., high-definition images, color images, grayscale images, heat maps, etc.) of model inputs to better assist the model in extracting useful information payloads, thereby improving the accuracy of model inference. Alternatively, if active nodes are concentrated in portions of the model input that are significantly outside the ROI, a user, model developer, or model can save unnecessary computational resources and improve computational efficiency for model inference by deactivating or removing these nodes from the model's network structure.
[0083] It should be understood that the examples described with reference to FIGS. 8-10 are for illustrative purposes only and should not be construed as limiting the scope of the present disclosure.
[0084] In some embodiments, the above functions / features may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, the functions may be stored as one or more commands or codes in a non-transitory computer-readable storage medium or a non-transitory processor-readable storage medium. The blocks of the methods or algorithms disclosed herein may be implemented in a processor-executable software module, which may reside on a non-transitory computer-readable storage medium or a processor-readable storage medium. The non-transitory computer-readable storage medium or a processor-readable storage medium may be any storage medium accessible by a computer or a processor. By way of example and not limitation, such non-transitory computer-readable storage medium or a processor-readable storage medium may include RAM, ROM, EEPROM, flash memory, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage, or any other medium used to store desired program code in the form of commands or data structures and accessible by a computer. As used herein, disks and optical disks include compact disks (CDs), laser disks, optical disks, digital versatile disks (DVDs), floppy disks, and Blu-ray disks, where disks typically reproduce data magnetically, whereas disks reproduce data optically with laser light. Combinations of the above are also included within the scope of non-transitory computer-readable media and processor-readable media. Furthermore, the operations of a method or algorithm may reside as one or any combination or set of code and / or commands in a non-transitory processor-readable storage medium and / or computer-readable storage medium, which may be incorporated into a computer program product.
[0085] FIG. 11 illustrates an exemplary hardware and software environment for an autonomous vehicle 1100 in which various techniques disclosed herein can be implemented. For example, the vehicle 1100 is shown traveling on a road 1101, and includes a power system 1102 including a prime mover 1106 that can be powered by an energy source 1104 and provide power to a drivetrain 1108, and a vehicle operation system 1110 including directional control 1112, power system control 1114, and brake control 1116. The vehicle 1100 can be implemented as any number of different types of vehicles, including vehicles that can transport people and / or cargo, travel over land, over sea, travel in the air, underground, under the sea, and / or in space. And, it should be understood that the above components 1102-1116 can vary widely based on the type of vehicle in which these components are used.
[0086] For simplicity, the embodiments described below focus on wheeled land vehicles such as cars, vans, trucks, buses, motorcycles, all-terrain vehicles (ATVs), etc. In such embodiments, the energy source 1104 may include, for example, a fuel system (e.g., providing gasoline, diesel, hydrogen, etc.), a battery system, solar panels, or other renewable energy, and / or a fuel cell system. The prime mover 1106 may include one or more motors and / or internal combustion engines, etc. The drive train 1108 may include wheels and / or tires, a driveline and / or any other mechanical driving components suitable for converting the power output of the prime mover 1106 into vehicle motion, as well as one or more brakes configured to controllably stop or slow the vehicle 1100, and a steering or steering component suitable for controlling the trajectory of the vehicle 1100 (e.g., a rack gear and pinion steering linkage, whereby one or more wheels of the vehicle 1100 can pivot about a substantially vertical axis to change the angle of the plane of rotation of the wheels relative to the longitudinal axis of the vehicle). In some embodiments, a combination of power system and energy source may be used (e.g., in the case of an electric / gas hybrid vehicle), and in other embodiments, multiple motors (e.g., dedicated to a single wheel or axle) may be used as prime mover 1106. In the case of a hydrogen fuel cell embodiment, prime mover 1106 may include one or more motors, and energy source 1104 may include a fuel cell system powered by hydrogen fuel.
[0087] Directional control 1112 may include one or more actuators or sensors for controlling and receiving feedback from directional or steering components to enable vehicle 1100 to follow a desired trajectory. Powertrain control 1114 may be configured to control the speed and / or direction of vehicle 1100 by controlling the output of drivetrain 1102 (e.g., controlling the output power of prime mover 1106, controlling driveline gears in drivetrain 1108, etc.). Brake control 1116 may be configured to control one or more brakes, e.g., disc or drum brakes coupled to the wheels of the vehicle, to slow or stop vehicle 1100.
[0088] Other vehicle types (including, but not limited to, all-terrain or tracked vehicles, and construction equipment) may use different power systems, drive trains, energy sources, directional control, power system control, and brake control. Also, in some embodiments, some components may be combined, for example, directional control of a vehicle is handled primarily by modifying the power output of one or more prime movers. Thus, the embodiments disclosed herein are not limited to the specific application of the techniques described herein in autonomous, wheeled, or land vehicles.
[0089] In the illustrated embodiment, full or semi-automatic control of the vehicle 1100 is realized in a main vehicle control system 1118. The main vehicle control system 1118 may include one or more processors 1122 configured to execute program code commands 1126 stored in a memory 1124, and one or more memories 1124. The processor 1122 may include, for example, a graphic processing unit(s) (GPU) and / or a central processing unit (CPU). The processor 1122 may further include an application specific integrated circuit (ASIC), other chipset, logic circuitry, and / or data processing device. The memory 1124 is used, for example, to load and store data and / or commands for the control system 1118. The memory 1124 may include any combination of suitable volatile memory (e.g., read only memory (ROM), dynamic random access memory (DRAM), random access memory (RAM), non-volatile memory (e.g., flash memory, memory cards, storage media), and / or other storage devices. When an embodiment is implemented in software, the techniques described herein may be implemented with modules, processes, functions, entities, etc. that perform the functions described herein. Modules may be stored in memory and executed by a processor. Memory may be implemented within or external to the processor and may be communicatively coupled to the processor via various means known in the art.
[0090] The sensors 1130 may include various sensors suitable for collecting information from the vehicle's surrounding environment to control the operation of the vehicle 1100. For example, the sensors 1130 may include one or more detection and ranging sensors (e.g., RADAR sensor 1134, LIDAR sensor 1136, or both), satellite navigation (SATNAV) sensor 1132, such as those compatible with any of the various satellite navigation systems (e.g., Global Positioning System (GPS), Global Navigation Satellite System (GLONASS), Beidou Navigation Satellite System (BDS), Galileo, Compass), etc. The radio detection and ranging (RADAR) 1134, light detection and ranging (LIDAR) sensor 1136, and digital camera 1138 (which may include various types of image capture devices capable of capturing still and / or video images) are used to sense stationary and moving objects within the vehicle's immediate area. The camera 1138 may be a monochrome camera or a stereo camera and may record still and / or video images. The SATNAV sensor 1132 is used to determine the vehicle's position on Earth using satellite signals. The sensors 1130 may optionally include an inertial measurement unit (IMU) 1140. The IMU 1140 may include multiple gyroscopes and accelerometers capable of detecting linear and rotational motion in three directions of the vehicle 1100. One or more other types of sensors (e.g., wheel rotation sensor / encoder 1142) are used to monitor the rotation of one or more wheels of the vehicle 1100.
[0091] In various embodiments, the removable hardware pod is unknown to the vehicle and can be attached to a variety of non-autonomous vehicles, including cars, buses, vans, trucks, mopeds, tractor trailers, exercise vehicles, and the like. While autonomous vehicles typically include a full sensor suite, in many embodiments the removable hardware pod may include a dedicated sensor suite. This dedicated sensor suite typically has fewer sensors than a fully autonomous vehicle sensor suite and may include an IMU, a 3D positioning sensor, one or more cameras, a LIDAR unit, and the like. Additionally or alternatively, the hardware pod may collect data from the non-autonomous vehicle itself, such as by integrating with the vehicle's CAN bus to collect various vehicle data including vehicle speed data, braking data, steering control data, and the like. In some embodiments, the removable hardware pod may include computing facilities that aggregate data collected by the removable pod sensor suite and vehicle data collected from the CAN bus and upload the collected data to a computing system for further processing (e.g., uploading the data to the cloud). In many embodiments, the computing equipment in the removable pod can apply a timestamp to each instance of the data before uploading the data for further processing. Additionally or alternatively, one or more sensors in the removable hardware pod can apply a timestamp when the data is collected (e.g., a laser radar unit can provide its own timestamp). Similarly, the computing equipment in the autonomous vehicle can apply a timestamp to data collected by the autonomous vehicle's sensor suite and upload the time-stamped autonomous vehicle data to a computer system for additional processing.
[0092] The output of the sensors 1130 may be provided to a set of main control subsystems 1120, including, for example, a positioning subsystem, a perception subsystem, a planning subsystem, and a control subsystem. The positioning subsystem is primarily responsible for accurately determining the position and orientation (sometimes referred to as "attitude" or "attitude estimate") of the vehicle 1100 within its surrounding environment, typically within some frame of reference. In some embodiments, the attitude is stored in memory 1124 as positioning data. In some embodiments, a surface model is generated from a high-resolution map and stored in memory 1124 as surface model data. In some embodiments, detection and ranging sensors store their sensor data in memory 1124 (e.g., radar data point clouds are stored as radar data). In some embodiments, calibration data is stored in memory 1124. The perception subsystem is primarily responsible for detecting, tracking, and / or identifying objects within the environment surrounding the vehicle 1100. According to some embodiments, machine learning models such as those described above are used to plan the vehicle's trajectory. The control subsystem 1120 is primarily responsible for generating appropriate control signals to control various controls in the vehicle control system 1118 to achieve the planned trajectory of the vehicle 1100. Similarly, machine learning models are used to generate one or more signals to control the autonomous vehicle 1100 to achieve the planned trajectory.
[0093] It should be understood that the collection of components for the vehicle control system 1118 shown in FIG. 11 is merely an example. In some embodiments, individual sensors may be omitted. Additionally or alternatively, in some embodiments, multiple sensors of the same type shown in FIG. 11 are used for redundancy and / or coverage of different areas around the vehicle. Also, there may be additional sensors of other types to provide actual sensor data related to the operation and environment of the wheeled land vehicle, in addition to the types described above. Similarly, in other embodiments, different types of control subsystems and / or combinations of control subsystems may be used. Additionally, while the primary control subsystem 1120 is shown as separate from the processor 1122 and memory 1124, it will be understood that in some embodiments, some or all of the functionality of the primary control subsystem 1120 may be implemented using program code commands 1126 resident in one or more memories 1124 and executed by one or more processors 1122, and in some cases, the primary control subsystem 1120 may be implemented using the same processor and / or memory. The subsystems may be implemented, at least in part, using various special purpose circuit logic, various processors, various field programmable gate arrays (FPGAs), various application specific integrated circuits (ASICs), various real time controllers, etc. As discussed above, the multiple subsystems may utilize circuits, processors, sensors, and / or other components. Additionally, the various components within the vehicle control system 1118 may be networked in various ways.
[0094] For example, vehicle 1100 may include one or more network interfaces, e.g., network interface 1154 adapted to communicate with one or more network interfaces 1150 (e.g., a LAN, a WAN, a wireless network, and / or the Internet, etc.) to enable communication of information with other vehicles, computers and / or electronic devices, including a central service, such as a cloud service, from which vehicle 1100 receives environmental and other data for automatic control.
[0095] For additional storage, the vehicle 1100 may also include one or more mass storage devices, such as a floppy disk or other removable disk drive, a hard disk drive, a direct access storage device (DASD), an optical drive (e.g., CD drive, DVD drive, etc.), a solid state storage drive (SSD), network attached storage, a storage area network, and / or a tape drive. The vehicle 1100 may also include a user interface 1152 such that the vehicle 1100 can receive inputs from a user or operator and generate outputs for the user or operator. This user interface 1152 may be, for example, one or more displays, a touch screen, a voice and / or gesture interface, buttons and other tactile controls, etc. Alternatively, the user inputs are received from, for example, a remote operator, via another computer or electronic device, for example, via an application on a mobile device or a web interface.
[0096] Disclosed herein are systems and methods for object detection and detection confidence. The disclosed methods may be suitable for autonomous driving, but may also be used in other applications, such as robotics, video analytics, weather forecasting, medical imaging, and the like. The disclosure may describe an exemplary autonomous vehicle 1100. Although the disclosure primarily provides examples using autonomous vehicles, other types of devices may be used to implement the various methods described herein, such as robots, camera systems, weather forecasting devices, medical imaging devices, and the like. Additionally, the methods may be used to control an autonomous vehicle, or for other purposes, such as, but not limited to, video surveillance, video or image editing, video or image research or retrieval, object tracking, weather forecasting (e.g., using radar data), and / or medical imaging (e.g., using ultrasound or magnetic resonance imaging (MRI) data).
[0097] Those skilled in the art can understand that each of the units, algorithms, and steps described and disclosed in the embodiments of the present disclosure can be realized using electronic hardware or a combination of computer software and electronic hardware. Whether a function is implemented in hardware or software depends on the application conditions and the design requirements of the technical solution. Those skilled in the art can realize the functions for each specific application using different methods. And such realization should not exceed the scope of the present disclosure. Those skilled in the art can understand that the operation processes of the above systems, devices, and units are basically the same, so it can refer to the operation processes of the systems, devices, and units in the above embodiments. For ease and simplicity of description, these operation processes will not be described in detail.
[0098] The software functional unit may be stored in a computer-readable storage medium when embodied as a product and used and sold. Based on this understanding, the technical proposal proposed in the present disclosure may be substantially or partially implemented in the form of a software product. Alternatively, some of the technical proposals useful in the prior art may be realized in the form of a software product. The software product in the computer is stored in a storage medium containing a plurality of commands for a computing device, such as a personal computer, a server, or a network device, to execute all or part of the steps disclosed in the embodiments of the present disclosure. The storage medium includes a USB disk, a mobile hard disk, a ROM, a RAM, a floppy disk, or other types of media capable of storing program code. Although the present disclosure has been described in connection with what is considered to be the most practical and preferred embodiment, it should be understood that the present disclosure is not limited to the disclosed embodiment, but is intended to encompass various configurations made without departing from the broadest interpretation of the scope of the appended claims.
[0099] However, other modifications, variations, and substitutions are possible. Thus, the specification and figures are to be regarded as illustrative, not restrictive. In the claims, graphic symbols in parentheses are not to be construed as limiting the claims. The word "comprehensive" means that it does not exclude the presence of other parts or steps than those recited in the claims. Also, the terms "a" and "one" as used in this text define one or more. Furthermore, the introductory phrases "at least one" and "one or more" used in the claims should not be construed as limiting another claim element introduced by the indefinite article "a" or "one" to an invention containing only one of the introduced claim element. This is true even if the same claim includes the introductory phrases "one or more" and the indefinite article "a" or "one". This also applies to the use of definite articles. Unless otherwise stated, terms such as "first" and "second" are used to arbitrarily distinguish between the elements described by these terms and are not meant to be used to indicate a temporal or other priority of these elements. The fact that certain features are recited in different claims does not mean that a combination of these features cannot be used to advantage. While certain features of the invention have been illustrated and described herein, those skilled in the art may envision many modifications, substitutions, changes, and equivalents. It is therefore to be understood that the appended claims are intended to cover all such modifications and variations that fall within the true spirit of the invention.
[0100] It should be understood that various features of the embodiments of the present disclosure that are described for clarity in the context of individual embodiments may also be provided in combination in a single embodiment. Conversely, various features of the embodiments of the present disclosure that are described for brevity in the context of a single embodiment may also be provided alone or in any suitable subcombination. It should be understood by those skilled in the art that the embodiments of the present disclosure are not limited by what has been particularly shown and described above. Rather, the scope of the embodiments of the present disclosure is defined by the appended claims and their equivalents.
[0101] The foregoing description of the disclosed embodiments is provided to enable others to make or use the disclosed subject matter. Various modifications to these embodiments will be readily apparent, and the general principles defined herein may be applied to other embodiments without departing from the spirit or scope of the foregoing. Thus, the foregoing description is not intended to be limited to the embodiments set forth herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein. Thus, the claims are not intended to be limited to the embodiments set forth herein, but are to be accorded the fullest scope consistent with the language of the claims, and references to elements in the singular do not mean "one and only," unless expressly stated as such, but rather "one or more." Unless otherwise stated, the term "some" refers to one or more. All structural and functional equivalents of the elements of the various embodiments described in the foregoing description, whether known or later known, are expressly incorporated herein by reference and are intended to be encompassed by the claims. Moreover, the subject matter disclosed herein is not intended to be publicly specific, regardless of whether such disclosure is expressly set forth in the claims. No claim element should be construed as an apparatus functional unless the element is expressly recited using the phrase "apparatus for." It is understood that the specific order or hierarchy of blocks in the disclosed processes is an example of an exemplary method. Based on design preferences, it is understood that the specific order or hierarchy of blocks in the processes can be rearranged while remaining within the foregoing scope. The accompanying method claims present the elements of the individual blocks in a sample order, and are not meant to be limited to the specific order or hierarchy presented.
[0102] The various examples shown and described are provided as examples only to illustrate various features of the claims. However, features shown and described with respect to any given example are not necessarily limited to the associated example, and may be combined with or in any combination with other examples shown and described. Furthermore, the claims are not intended to be defined by any examples. The above method descriptions and process flow diagrams are provided as illustrative examples only, and are not intended to require or imply that the blocks of the various examples must be performed in the order presented. As will be understood, the order of blocks in the above examples may be performed in any order. Terms such as "then," "then," and "next" are not intended to limit the order of the blocks, and these words are merely used to guide the reader through the method description. Furthermore, any reference to a claim element in the singular, for example using the articles "a," "an," or "the," should not be construed as limiting the element to the singular. The various illustrative logic blocks, modules, circuits, and algorithm blocks described with respect to the examples disclosed herein may be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, various exemplary components, blocks, modules, circuits, and blocks have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in various ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure. Hardware for implementing the various exemplary logic, logic blocks, modules, and circuits described in connection with the examples disclosed herein may be implemented or performed using a general-purpose processor, a DSP, an ASIC, an FPGA or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein.A general purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Alternatively, some blocks or methods may be performed by circuitry that is specific to a given function. EXAMPLES
[0103] A method for visualizing neurons of an AI model for autonomous driving includes obtaining one or more neurons from a plurality of neurons for a task of the AI model; determining, for each neuron of the one or more neurons, a corresponding ROI of inputs related to the task, where the corresponding ROI is encoded by the one or more neurons for the task; and applying a first operation including an LRP to generate, for at least a portion of the one or more neurons, a human-explainable representation of the determined corresponding ROI of inputs. EXAMPLES
[0104] In the method according to Example 1, the input is collected for the task by sensors, a recorded human driving database, and / or cloud storage. EXAMPLES
[0105] In the method according to example 1 or 2, the input is a processed image frame, and the corresponding ROI for each neuron of the one or more neurons includes a set of pixels of the processed image frame corresponding to the corresponding ROI. EXAMPLES
[0106] In a method according to any one of Examples 1 to 3, the input is a sequence of processed image frames, and the corresponding ROI for each neuron of the one or more neurons includes a union of each pixel set of the sequence of processed image frames, each pixel set corresponding respectively to a sub-ROI of each processed image frame of the sequence of processed image frames. EXAMPLES
[0107] In the method according to any one of Examples 1 to 4, generating the human-explainable representation includes applying a second operation including VBP to identify, from the determined corresponding ROIs, a specific ROI that contributed most to the prediction made by the AI model to complete the task, thereby improving computational efficiency and enhancing model accuracy. EXAMPLES
[0108] In the method according to any one of Examples 1 to 5, the VBP is sequentially applied after the LRP, and the LRP is used to identify the one or more neurons from the plurality of neurons that most contributed to the prediction made by the AI model to complete the task, thereby improving computational efficiency and enhancing the accuracy of the model. EXAMPLES
[0109] The method according to any one of Examples 1 to 6, wherein the AI model includes a mixing block and a model backbone. EXAMPLES
[0110] In the method according to any one of Examples 1 to 7, applying the first operation and the second operation includes applying the LRP through the mixing block to obtain a weight mask for a feature map of the one or more neurons, weighting the feature map of the one or more neurons using the weight mask to obtain a weighted feature map of the one or more neurons, and applying the VBP to back-transmit the weighted feature map of the one or more neurons through the model backbone. EXAMPLES
[0111] The method according to any one of the first to eighth embodiments, wherein the input is a spectrogram of an audio segment. EXAMPLES
[0112] A non-transitory computer-readable storage medium having stored thereon commands that, when executed by one or more processors, cause the one or more processors to: obtain one or more neurons from a plurality of neurons for a task of an AI model; determine, for each neuron of the one or more neurons, a corresponding ROI of an input related to the task, where the corresponding ROI is encoded by the one or more neurons for the task; and apply a first operation including an LRP to generate, for at least a portion of the one or more neurons, a human-explainable representation of the determined corresponding ROI of the input. EXAMPLES
[0113] In the non-transitory computer-readable storage medium of Example 10, the input is collected for the task by sensors, and / or a recorded human driving database, and / or cloud storage. EXAMPLES
[0114] In the non-transitory computer-readable storage medium of Example 10 or 11, the input is a processed image frame, and the corresponding ROI for each neuron of the one or more neurons includes a set of pixels corresponding to the corresponding ROI of the processed image frame. EXAMPLES
[0115] In a non-transitory computer-readable storage medium described in any one of Examples 10 to 12, the input is a sequence of processed image frames, and the corresponding ROI for each neuron of the one or more neurons includes a union of each pixel set of the sequence of processed image frames, each pixel set corresponding respectively to a sub-ROI of each processed image frame of the sequence of processed image frames. EXAMPLES
[0116] Generating the human-explainable representation in the non-transitory computer-readable storage medium described in any one of Examples 10 to 13 includes applying a second operation including VBP to identify, from the determined corresponding ROIs, a specific ROI that contributed most to the prediction made by the AI model to complete the task, thereby improving computational efficiency and enhancing model accuracy. EXAMPLES
[0117] The non-transitory computer-readable storage medium of any one of Examples 10 to 14 is sequentially applied with the VBP after the LRP, and the LRP is used to identify the one or more neurons from the plurality of neurons that contributed most to the prediction made by the AI model to complete the task, thereby improving computational efficiency and enhancing model accuracy. EXAMPLES
[0118] In the non-transitory computer-readable storage medium of any one of Examples 10 to 15, the AI model includes a mixing block and a model backbone. EXAMPLES
[0119] In the non-transitory computer-readable storage medium of any one of Examples 10 to 16, applying the first operation and the second operation includes applying the LRP through the mixing block to obtain a weight mask used for a feature map of the one or more neurons, weighting the feature map of the one or more neurons using the weight mask to obtain a weighted feature map of the one or more neurons, and applying the VBP to back-transmit the weighted feature map of the one or more neurons through the model backbone. EXAMPLES
[0120] In any one of Examples 10 to 17, the input is a spectrogram of an audio segment. EXAMPLES
[0121] A computer-implemented system includes one or more processors and one or more memory devices that store commands that, when executed by the one or more processors, cause the one or more processors to: obtain one or more neurons from a plurality of neurons for a task of an AI model; determine, for each neuron of the one or more neurons, a corresponding ROI of an input related to the task, where the corresponding ROI is encoded by the one or more neurons for the task; and apply a first operation including an LRP to generate a human-explainable representation of the determined corresponding ROI of the input for at least a portion of the one or more neurons. EXAMPLES
[0122] In the system described in Example 19, generating the human-explainable representation includes applying a second operation including VBP to identify, from the determined corresponding ROIs, a specific ROI that contributed most to the prediction made by the AI model to complete the task, thereby improving computational efficiency and enhancing model accuracy.
Claims
1. A method for visualizing neurons of an AI (Artificial Intelligence) model for autonomous driving, comprising: Obtaining one or more neurons from a plurality of neurons for a task of the AI model; determining, for each neuron of the one or more neurons, a corresponding Region of Interest (ROI) of inputs relevant to the task, the corresponding ROI being encoded for the task by the one or more neurons; and applying a first operation including Layer-wise Relevance Propagation (LRP) to generate an explainable artificial intelligence-based representation of the determined input corresponding ROI for at least a portion of the one or more neurons. A method for visualizing neurons of an AI model for autonomous driving, comprising:
2. The explainable AI-based representation is a human-interpretable explainable representation or a machine-interpretable explainable representation; 2. The method of claim 1 .
3. The input is image frames collected and processed for the task by sensors, a recorded human driving database, and / or cloud storage; the corresponding ROI for each neuron of the one or more neurons includes a set of pixels of the processed image frame that correspond to the corresponding ROI; 2. The method of claim 1 .
4. the input is a sequence of processed image frames; the corresponding ROI for each neuron of the one or more neurons comprises a union of respective pixel sets of the sequence of processed image frames; each said set of pixels respectively corresponding to a sub-ROI of a respective processed image frame of said sequence of processed image frames; 2. The method of claim 1 .
5. Generating an explainable AI-based representation Applying a second operation including Visual Back-Propagation (VBP) to identify a specific ROI that contributes most to the prediction made by the AI model to complete the task from the determined corresponding ROIs, thereby improving computational efficiency and model accuracy.
2. The method of claim 1 .
6. The VBP is applied in sequence after the LRP, and the LRP is used to identify the one or more neurons from the plurality of neurons that most contributed to the prediction made by the AI model to complete the task, thereby improving computational efficiency and model accuracy; 6. The method of claim 5 .
7. The AI model includes a mixing block and a model backbone; Applying the first operation and the second operation applying the LRP through the mixing block to obtain a weight mask for use in a feature map of the one or more neurons; weighting the feature map of the one or more neurons using the weight mask to obtain a weighted feature map of the one or more neurons; and applying the VBP to transmit the weight feature map of the one or more neurons back through the model backbone.
7. The method of claim 6.
8. The input is a spectrogram of an audio segment.
2. The method of claim 1 .
9. A non-transitory computer-readable storage medium having commands stored thereon, comprising: A non-transitory computer-readable storage medium, the instructions of which, when executed by one or more processors, cause the one or more processors to perform the method of any one of claims 1 to 8.
10. 1. A computer-implemented system comprising: one or more processors and one or more memory devices for storing commands; A system, wherein the commands, when executed by the one or more processors, cause the one or more processors to perform a method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Method for operating neural network and data classification system
JP2021124979A