Device and method for utilizing attention maps for efficient mixing of depths or mixing of experts

By using transformer attention maps to determine token importance, the computational and training issues of MoD models are addressed, achieving faster and more stable processing with reduced resource consumption.

DE102024208985A1Pending Publication Date: 2026-03-19ROBERT BOSCH GMBH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
DE102024208985
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-19
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Existing MoD models based on transformer architectures face computational overhead and training instability due to external router networks, leading to increased model complexity and performance issues.

Method used

Utilize attention maps from previous layers within the transformer model to determine token importance, eliminating the need for separate router networks and improving training stability.

Benefits of technology

This approach reduces computational overhead, conserves memory, and enhances training stability by deriving token weights directly from attention maps, resulting in faster runtime and more efficient resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A computer-implemented method for operating a neural network to process its input data, in particular to classify input signals received from a sensor, wherein the neural network comprises at least one MoE (Mixture of Experts) layer comprising multiple experts, and wherein the method comprises the following steps: propagating input data through the neural network comprising the at least one MoE layer, determining an expert selection based on activations of a preceding layer of the Mixture-of-Expert layer of the neural network, and processing the input data using the selected expert within the at least one MoE layer.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method for improving MoD (Mixture of Depths) layers in depth learning models, in particular those based on a transformer architecture. State of the art

[0002] The MoD (Mixture of Depths) concept is known from the publication by Raposo, David, et al. “Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.” arXiv-Preprint arXiv:2404.02258 (2024).

[0003] Models of Dependence (MoDs) aim to improve the computational efficiency of deep learning models at scale, particularly those based on a transformer. MoDs achieve this by selectively processing input tokens based on their importance using dedicated "router" networks. The routers analyze input data and assign importance scores to individual tokens in each layer, dictating which tokens receive full or partial processing.

[0004] However, this approach often suffers from two main drawbacks: computational overhead due to the additional router networks themselves, and potential training instability during the router optimization process. These limitations often lead to increased model complexity and can negatively impact overall performance. Advantages of the invention

[0005] Instead of using external routers, this invention directly utilizes the outputs, such as attention maps, from previous layers in the model to determine token importance. Since the outputs inherently contain information about token importance, this eliminates the need for dedicated routing networks.

[0006] The advantage of this approach is faster runtime, as no overhead from separate router calculations is required. Less memory is used during inference and training because the absence of additional routers results in fewer parameters in the model. Furthermore, training stability is improved because token weight is derived directly from attention, leading to more stable training and avoiding router-related problems. Disclosure of the invention

[0007] In a first aspect, the invention relates to a method for operating a neural network for processing its input data, in particular for classifying input signals received from a sensor, wherein the neural network comprises at least one MoE (Mixture of Experts) layer comprising several experts.

[0008] The processing of the input data is performed by propagating the input data across the neural network. For the at least one MoE layer, a step is performed to determine an expert selection based on activations of a previous layer of the mixture-of-experts layer of the neural network, specifically before the MoE layer is applied. The activations can be neuron activations from a layer preceding the at least one MoE layer. More precisely, the previous layer of the MoE layer is the immediately preceding layer that is (directly) connected to the input of the MoE layer. After the experts have been selected, the propagated input data is further propagated by the selected experts within the at least one MoE layer, and preferably, a feature for the next layer or an output of the neural network is then output.The output can be a classification with respect to several classes of input data.

[0009] An expert in a Mixture of Experts (MoE) layer can be understood as a specialized subnetwork designed to handle a specific subset of the MoE layer's input data. In other words, the experts themselves can be considered separate small networks within the neural network.

[0010] Typically, at the beginning of the propagation step through the neural network, general features can be determined by the first layers of the neural network. The intermediate features are used by the present invention to select a subset of experts. These features are then passed to the selected experts. The experts' output can be the final output or be further processed.

[0011] Preferably, the step of determining an expert selection is performed by a selection function for selecting at least one expert of the MoE layer, wherein the selection function inputs are the activations and the output of the selection function characterizes which expert or experts of the MoE layer must be selected.

[0012] It is proposed that the preceding layer of the MoE layer is a transformer layer, where the input activations of the selection function constitute an attention map of the transformer layer. In this embodiment, the selection function determines importance ratings for input tokens of the transformer layer and selects a subset of tokens from the input data based on the determined importance ratings, with tokens with higher importance ratings being more likely to be selected.

[0013] Surprisingly, the inventors discovered that the attention map contains relevant information for selecting experts.

[0014] A transformer layer can be understood as a building block of deep learning models that analyzes sequences of data by understanding relationships between elements. It uses an "attentional mechanism" to determine which parts of the sequence are most relevant to each other, enabling it to process information with a strong sense of context.

[0015] It should be noted that the invention is particularly relevant for applications requiring real-time processing and limited computing resources. It can be applied to various tasks, such as image recognition, audio signal processing, anomaly detection, and control of technical systems, including robotics and autonomous vehicles.

[0016] Preferably, the invention can be applied to video and audio analysis, particularly for tasks involving the prediction of continuous values ​​– this is known as regression analysis. For example, it can estimate values ​​such as distance, speed, acceleration, or even track the movement of an object within the data. This is achieved by analyzing low-level features that are fundamental elements, such as edges or individual pixel characteristics in images or similar fundamental components in audio data.

[0017] Further aspects of the invention consider using the neural network of the first aspect of the invention as a classifier, by a method comprising the following steps: - Receiving a sensor signal, including data from a sensor, - Determining an input signal that depends on the sensor signal, and - Feeding the input signal into the classifier to obtain an output signal that characterizes a classification of the input signal.

[0018] The classifier, e.g., a neural network, can be equipped with a structure that is trainable to identify and differentiate, for example, pedestrians and / or vehicles and / or traffic signs and / or traffic lights and / or road surfaces and / or human faces and / or medical anomalies in imaging sensor images. Alternatively, the classifier, e.g., a neural network, can be equipped with a structure that is trainable to identify spoken commands in audio sensor signals.

[0019] Embodiments of the invention are discussed in more detail with reference to the following figures. The figures show: Fig. 1 a schematic embodiment of the architecture of the invention; Fig. 2 a schematic flowchart of an embodiment of the invention; Description of the embodiments

[0020] The field of "Mixture of Experts" investigates models for efficient inference and training of neural networks. More specifically, "Mixture of Experts" focuses on larger networks with only a slight increase in the number of computational operations (e.g., FLOPs or MACs). To achieve this increase in model size without increasing the computational effort, only a portion of the network (the so-called expert) is active during inference. This subnetwork is selected by a router (also known as a gating mechanism or dispatcher). Typically, the experts are fixed blocks from which the routing network makes its selection.

[0021] It is important to note that the router is an additional small network. Although the computational overhead of this small network is usually not too large, it can have an impact in the field of embedded and computationally constrained networks, as it increases the inference time (e.g., autonomous driving).

[0022] The present invention aims to eliminate the aforementioned overhead of the routing network. Instead of having two networks, one for routing and a main network for the actual task at hand (e.g., image recognition), we propose integrating the routing network into the main network. The resulting dual use of parts of the network allows us to further reduce the inference overhead.

[0023] Known mixture-of-expert models consist of an external router and a main network. The main network has one or more layers, each composed of a mixture of experts (often called mixture-of-expert layers). These experts are themselves separate, small networks. Different experts are active depending on the input. It should be noted that the neural network can be applied to computer vision tasks where the input type is images. However, the neural network can also be used in applications with any input data type, such as text, audio, etc.

[0024] In a first aspect of the invention, it is proposed to discard the external router and reuse internal features 12 of the main neural network 10 to decide which expert 11 should be activated.

[0025] For a concrete, non-restrictive example, we consider a network (e.g., a fully connected MLP) with an integrated router, as in Fig. 1 shown. Let's say we have an input x and a network f. To create a router within this network, we take a layer L=2 (as shown in Fig. (1 shown). Here we take 3 neurons N=1,2,3 within this layer. Now we use not only the neuron activations to calculate the output, but also to select the expert 11. The expert can be selected, for example, by applying a selection function: argmaxN(f(L=2,N=1)(x),f(L=2,N=2)(x),f(L=2,N=3)(x))

[0026] Depending on the `argmax` value, we select a single expert for use in the calculation. The other experts are hidden and not used. Since `argmax` also depends on the input, different experts can be selected for different inputs.

[0027] Overall, the above explanation is simplified. More generally, this approach can also be applied to other network architectures such as transformers, CNNs, and RNNs. Neural network 10 can also have multiple expert layers. Furthermore, an expert can be more than just a neuron in the network; it can also be the network itself. In our example, a single expert is active, but this could also be top-k experts.

[0028] In a second exemplary embodiment of the invention, the neural network comprises 10 transformer-based backbones. The inventors have surprisingly discovered that attentional layers within these transformers contain information about relative token importance that is relevant for routing. Instead of obtaining routing decisions from dedicated networks, it is proposed to obtain the routing decisions directly from attentional maps.

[0029] The MoD layer produces initial token Y t from entry token X t based on a transformer block f. Under MoD, only some tokens are processed by the transformer block, depending on the capacity factor c.

[0030] In this example, the token selection is based on attention map A. t-1from the previous layer. The aggregation of attention maps to calculate importance ratings can be done in various ways, such as by calculating averages across a row or column. The higher the ratings, the more likely the corresponding tokens are to be selected.

[0031] Overall, our algorithm can be written as pseudocode as follows:

[0032] Fig. Figure 2 shows a schematic flowchart of an embodiment of the invention for the more computationally efficient operation of a neural network with a MoE layer.

[0033] The process begins at step S21. In this step, neural network 10 is deployed. Its architecture includes at least a Mixture-of-Experts (MoE) layer, which itself comprises a collection of individual "experts"—smaller subnetworks specialized in processing different types of data and patterns.

[0034] Next, a step is performed to receive and propagate input data S22. Data that needs to be processed, such as images, audio, or sensor readings, is propagated via the neural network.

[0035] In the subsequent step S23, the data is propagated to the MoE layer, and the experts are then selected. The input data is initially processed through the layers of the neural network until it reaches the MoE layer. Instead of using a separate router network, neuron activations are extracted by a specific layer within the main network. This layer is strategically positioned before the MoE layer to capture relevant information from the processed input. The extracted neuron activations are then provided as input for the expert, for example, according to Equation 1 above.

[0036] Once the expert is selected, the input data bypasses the other experts and is processed only by the selected expert within the MoE layer. This targeted processing conserves computational resources and enables specialized handling of the input.

[0037] Finally, the output is passed from the selected expert along the remaining layers of the main neural network for further processing and output (S24).

[0038] After all relevant layers have been processed, the neural network produces the final processed output. This can be a classification, a prediction, or a transformed representation of the input data, depending on the network's goal.

[0039] In an optional step S25, the output of S24 can be used to control a technical system accordingly.

[0040] Fig. Figure 2 also shows a computer 30, comprising a memory element 40 on which a computer program is stored. When the computer program is executed on the computer 30, it causes the computer 30 to execute the procedure 20. QUOTES INCLUDED IN THE DESCRIPTION

[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited non-patent literature

[0000] Raposo, David, et al. “Mixture-of-Depths: Dynamically allocating computation in transformer-based language models.” arXiv preprint arXiv:2404.02258 (2024

[0002]

Claims

[1] Computer-implemented method (20) for operating a neural network (10) for processing its input data, in particular for classifying input signals received from a sensor, wherein the neural network comprises at least one MoE (Mixture of Experts) layer comprising multiple experts, the method comprising the following steps: Propagation (S22) of input data through the neural network (10), comprising, for at least one MoE layer, determining an expert selection based on activations of a preceding layer of the mixture-of-expert layer of the neural network; and Processing the input data using the selected expert (11) within at least one MoE layer. [2] Method according to claim 1, wherein the step of determining an expert selection is performed by a selection function for selecting at least one expert of the MoE layer, wherein the selection function inputs are the activations and the output of the selection function characterizes which expert of the MoE layer must be selected. [3] Method according to claim 2, wherein the selection function comprises an argmax function or a top-k selection function. [4] Method according to claim 2, wherein the preceding layer of the MoE layer is a transformer layer, wherein the input activations of the selection function are an attention map of the transformer layer. [5] Method according to claim 4, wherein the selection function determines importance ratings for input tokens of the transformation layer, selecting a subset of tokens from the input data based on the determined importance ratings. [6] Method according to claim 5, wherein determining the importance ratings comprises aggregating values ​​in the attention map according to the individual tokens. [7] Method according to claim 6, wherein aggregating values ​​in the attention map comprises using values ​​across rows and / or columns of the attention map. [8] Computer-implemented method for using the neural network (10) according to any one of the preceding claims as a classifier for classifying sensor signals, in particular for computer vision, image recognition, audio signal processing, anomaly detection or control of a technical system, wherein the classifier is trained using the method according to any one of claims 1 to 7, comprising the following steps: - Receiving a sensor signal, comprising data from a sensor (30), - Determining an input signal that depends on the sensor signal, and - Feeding the input signal (x) into the classifier to obtain an output signal that characterizes a classification of the input signal. [9] Computer-implemented method for using the classifier of claim 8 to provide an actuator control signal for controlling an actuator, comprising all steps of the method of claim 8 and further comprising the following step: - Determining the actuator control signal as a function of the output signal. [10] Method according to claim 9, wherein the actuator controls an at least partially autonomous robot and / or a manufacturing machine and / or an access control system and / or an edge computing device and / or an autonomous vehicle, a household appliance, an electric tool. [11] Computer program designed to cause a computer to execute the method according to any one of claims 1 to 10, including all its steps, when the computer program is executed by a processor (30). [12] Machine-readable storage medium (40) on which the computer program according to claim 11 is stored. [13] System (30) designed to perform the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Asymmetrical robustness for classification in hostile environments

    DE102020215485A1