Processing of measurement data in neural networks with enhanced memory

DE102024202129A1Pending Publication Date: 2025-09-11ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102024202129
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-07
Publication Date
2025-09-11

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Method (100) for processing measurement data (2) to an output (3) with regard to a given task with a neural network architecture (1) comprising a sequence of an input layer (11), several intermediate layers (12-14) and an output layer (15), comprising the steps: • the measurement data (2) are fed (110) to the input layer (11); • the output of the input layer (11) is processed (120) successively in the intermediate layers (12-14) and finally in the output layer (15), wherein the output of a layer (11-14) is fed to the next layer (12-15) as input; • the desired output (3) is output (160) from the output layer (15); where • a rule (6) for retrieving and / or aggregating contents (8) from a memory (7) is determined (130) on the basis of a predetermined functional approach (4) in accordance with an intermediate product (5) supplied as an output by at least one predetermined intermediate layer (12-14); • the contents (8) retrieved and / or aggregated in accordance with this provision (6) are merged with the intermediate product (5) in accordance with a further specified provision (9) (140) and • the intermediate product (5#) modified in this way replaces the original intermediate product (5) during further processing in the neural network architecture (1) (150).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present invention relates to the processing of measurement data in neural networks, which is, for example, a basic building block for environmental monitoring of vehicles or robots. State of the art

[0002] Neural network architectures are increasingly being used to process measurement data, especially images. After training with a sufficiently large set of training examples that exhibit a variability appropriate for the respective application, neural networks can be expected to generalize well to measurement data unseen during training. This is especially true for neural networks used as image classifiers or for semantic segmentation of images.

[0003] The information learned during training is stored in the parameters that characterize the behavior of the neural network. These parameters are optimized during training to ensure the best possible overall processing of the training examples. These parameters can include, for example, weights with which the inputs fed to each neuron in the network are summed to activate that neuron. Disclosure of the invention

[0004] The invention provides a method for processing measurement data into an output with respect to a given task. This method uses a neural network architecture comprising a sequence of an input layer, several intermediate layers, and an output layer.

[0005] In this process, the measurement data is fed to the input layer. The output of the input layer is processed successively in the intermediate layers and finally in the output layer, with the output of each layer being fed as input to the next layer. This does not preclude the possibility of connections between layers that skip one or more layers in the sequence. The desired output of the neural network architecture as a whole is finally output from the output layer.

[0006] This process is extended to the extent that, based on a predefined functional approach and in accordance with an intermediate product delivered as output by at least one predefined intermediate layer, a rule for retrieving and / or aggregating content from a memory is determined. The content retrieved and / or aggregated according to this rule is merged with the intermediate product according to another predefined rule. During further processing in the neural network architecture, the intermediate product modified in this way replaces the original intermediate product.

[0007] This means that the intermediate product is not always subjected to the same change, but rather this change depends on the intermediate product itself. In the broadest sense, the process can be understood as parsing the intermediate product and deriving a query for the contents stored in the memory. This is roughly comparable to a customer entering a tailor's shop, placing a piece of clothing on the table, and asking which accessories from the tailor's stock can be used to enhance this particular piece of clothing. For example, a patch with a colorful motif can be requested for a pair of trousers, and Velcro shoulder pads for a sweater.

[0008] A stock of accessories in the tailor's shop creates scope for creative garment design. For example, the tailor can attach the patches to a dress or use the shoulder pads on T-shirts or polo shirts. The tailor comes up with such creative ideas because they have a rough overview of their inventory, enabling them to apply previously acquired knowledge in new contexts.

[0009] Exactly the same capability is imparted to the neural network architecture according to the method proposed here. Important features can be stored in the memory during training and then adaptively applied during the inference phase of the neural network architecture. If the given functional approach still has free parameters, these can also be optimized during training.

[0010] Experiments have shown that the memory capacity of the neural network architecture improved in this way noticeably increases its performance on a variety of tasks. This performance can be measured, for example, using test inputs for which the target outputs are known, which the network architecture is supposed to determine from these test inputs. The better the actual outputs match the target outputs, the better the performance of the neural network architecture.

[0011] The same quantitative improvement in performance could also be achieved by other means, such as scaling up the neural network architecture. Surprisingly, however, it has been shown that, relative to the quantitative improvement in performance, the improvement in memory proposed here requires much less additional computational effort than, for example, scaling up the network architecture.

[0012] In a particularly advantageous embodiment, the specified functional approach includes aggregating the intermediate product. This allows the essential information to be extracted from the intermediate product.

[0013] In particular, aggregation can, for example, involve reducing the dimensionality of an intermediate product in the form of a tensor by applying a convolution operation. For example, if the intermediate product comprises the output of a convolutional layer in a convolutional neural network (CNN), each convolution kernel applied in this intermediate layer produces its own feature map from the input of this intermediate layer. In neural network architectures for image processing, individual convolutional layers can contain on the order of 1,000 filter kernels, so that the intermediate product accordingly comprises 1,000 feature maps. The information from all these feature maps can be combined, for example, using a 1×1 convolution operation. The result then has only the dimensionality of a single feature map.

[0014] In another particularly advantageous embodiment, the specified functional approach involves using an evaluation network, which is another trained neural network, to determine parameters for weighting multiple contents from the memory from the intermediate product. The contents from the memory are then aggregated according to these parameters. In this way, for example, an intermediate product that still has the dimensionality of an entire feature map even after aggregation can be condensed into a few parameters, which can then be used to assemble contents from the memory. For this purpose, the evaluation network can particularly advantageously be a fully interconnected neural network.

[0015] In particular, the contents from the memory can be combined, for example, into a weighted sum. The coefficients in this sum can be, for example, the parameters determined using the evaluation network, or parameters determined from the intermediate product in any other way. The coefficients thus determine the strength of the contribution of each content in the memory to the change applied to the intermediate product.

[0016] In particular, value ranges for the parameters can be defined by choosing an activation function that translates activations of neurons in the output layer of the evaluation network into outputs of these neurons.

[0017] The various contents in memory, which are combined in the weighted sum, span a space of all possible results of this weighted sum. However, this space is only fully utilized if the coefficients can assume any value. However, if the coefficients are taken from an evaluation network, they are determined there using an activation function. Many such activation functions only deliver results within a limited range of values, which is also useful to prevent individual values ​​from escalating completely upwards or downwards. For example, • the softmax activation function coefficients between 0 and 1, where the sum of all coefficients is 1, • the hyperbolic tangent activation function coefficients in the open interval (-1, +1) and • The sigmoid activation function yields coefficients in the open interval (0,1). In contrast, a linear activation function yields coefficients in the full range (-∞, +∞).

[0018] The content retrieved from the memory and / or aggregated can be linked to the intermediate product in any way to modify it. However, it is particularly advantageous to add this content to the intermediate product. This allows this content to be considered additionally without completely obscuring the basic message of the intermediate product.

[0019] In another particularly advantageous embodiment, at least one predetermined intermediate layer is selected in the neural network architecture that is closer to the output layer than to the input layer. The features detected in these layers are comparatively complex features composed of the features detected in the first layers. To the extent that the neural network architecture serves to classify the input or parts of the input (e.g., by means of semantic segmentation), these deeper layers contain class-specific details, but without yet allowing objects to be clearly recognized, as in the case of images as input. Examples of complex features in deeper layers are textures.

[0020] This means that, for example, when processing images containing similar objects or situations, similar compositions of content from the memory are also merged with the intermediate product. The additional processing of the intermediate product using the specified functional approach and the merging of the subsequently selected content from the memory with the intermediate product thus form a block that can learn to distinguish between objects and situations during training of the neural network architecture. Particularly in neural network architectures designed as some form of classifier, the activations of the output layer of the evaluation network can form a class-specific pattern. This means that the newly added block as a whole enriches the intermediate product with class-specific details.

[0021] As explained above, in a particularly advantageous embodiment, a convolutional neural network with a plurality of convolutional layers, which process their respective inputs by smoothly applying one or more filter kernels to produce one or more feature maps as output, is chosen as the neural network architecture. The intermediate product then comprises at least one feature map. The effect of merging with contents from the memory can then be particularly well interpreted in the context of the respective application.

[0022] As previously explained, the neural network architecture is particularly advantageous as an image classifier and / or as a semantic segmenter for images. By adding class-specific details as explained above, the accuracy of the classification or semantic segmentation is increased with comparatively little effort.

[0023] In a particularly advantageous embodiment, measurement data recorded with at least one sensor is selected. A control signal is generated from the output of the neural network architecture. A vehicle, a driver assistance system, a robot, a quality control system, an area monitoring system, and / or a medical imaging system is controlled with the control signal. Due to the improved performance with which the neural network fulfills its task, the probability is then increased that the action performed by the controlled technical system in response to the control signal is appropriate to the situation embodied in the measurement data.

[0024] In another particularly advantageous embodiment, before the measurement data is fed to the input layer, training examples are processed in the same way to produce outputs. The resulting outputs are evaluated using a predefined cost function. Parameters that characterize the behavior of the neural network architecture, as well as at least the contents of the memory, are optimized with the goal of improving the evaluation by the cost function. In this way, knowledge is stored in the memory during training, which improves accuracy and performance in later operational use (inference).

[0025] In particular, additional parameters that characterize the behavior of the given functional approach can be included in the optimization with regard to the evaluation by the cost function. This allows the neural network architecture to learn to query the appropriate information from memory depending on the situation and then combine it with the intermediate product.

[0026] The method can, in particular, be fully or partially computer-implemented. Therefore, the invention also relates to a computer program with machine-readable instructions that, when executed on one or more computers and / or compute instances, cause the computer(s) and / or compute instances to execute the described method. In this sense, control units for vehicles and embedded systems for technical devices that are also capable of executing machine-readable instructions are also to be regarded as computers. Compute instances can, for example, be virtual machines, containers, or serverless execution environments, which can be provided in a cloud, in particular.

[0027] The invention also relates to a machine-readable data carrier and / or a downloadable product containing the computer program. A downloadable product is a digital product that can be transmitted over a data network, i.e., downloaded by a user of the data network, and which can be offered for immediate download, for example, in an online shop.

[0028] Furthermore, one or more computers and / or compute instances may be equipped with the computer program, the machine-readable data carrier or the download product.

[0029] Further measures improving the invention are presented in more detail below together with the description of the preferred embodiments of the invention with reference to figures. Examples of implementation

[0030] It shows: Fig. 1 embodiment of the method 100 for processing measurement data 2; Fig. 2 Example of the extension of the processing chain in the neural network architecture 1 to implement the method 100; Fig. 3 Example of storing class-specific information in memory 7 during training of an image classifier. Fig. 4 Example of a modification of an intermediate 5 to a modified intermediate 5#.

[0031] Fig. 1 is a schematic flow diagram of an embodiment of the method 100 for processing measurement data 2 to an output 3 with regard to a predetermined task using a neural network architecture 1. This neural network architecture 1 has a sequence of an input layer 11, several intermediate layers 12-14 and an output layer 15.

[0032] According to block 105, in particular, for example, a convolutional neural network having a plurality of convolutional layers that process their respective input by slidingly applying one or more filter kernels to one or more feature maps as output can be selected as neural network architecture 1.

[0033] According to block 106, the neural network architecture 1 can be designed, for example, as an image classifier and / or as a semantic segmenter for images.

[0034] In step 110, the measurement data 2 are fed to the input layer 11.

[0035] According to block 111, measurement data 2 recorded with at least one sensor can be selected. This sensor can, for example, be carried by a vehicle 50 or robot 60 or be mounted in some other way such that it can detect a traffic situation involving at least one vehicle 50 and / or robot 60.

[0036] In step 120, the output of input layer 11 is successively processed in intermediate layers 12-14 and finally in output layer 15. The output of each layer 11-14 is fed as input to the next layer 12-15.

[0037] In step 130, a rule 6 for retrieving and / or aggregating contents 8 from a memory 7 is determined on the basis of a predetermined functional approach 4 in accordance with an intermediate product 5 delivered as output by at least one predetermined intermediate layer 12-14.

[0038] According to block 131, this predetermined functional approach 4 may include aggregating the intermediate product 5.

[0039] According to block 131a, the aggregating may include reducing the dimensionality of an intermediate product in the form of a tensor by applying a convolution operation.

[0040] According to block 132, an evaluation network A, which is another trained neural network, can be used to determine parameters 8a for weighting multiple contents 8 from memory 7 from the intermediate product 5. The contents 8 from memory 7 can then be aggregated according to these parameters 8a according to block 133.

[0041] According to block 132a, value ranges for the parameters 8a can be determined by choosing an activation function that translates activations of neurons in the output layer of the evaluation network A into outputs of these neurons.

[0042] According to block 134, at least one predetermined intermediate layer 12-14 can be selected in the neural network architecture 1, which is closer to the output layer 15 than to the input layer 11.

[0043] Insofar as the neural network architecture 1 according to block 105 is a convolutional neural network that produces feature maps in the intermediate layers 12-14, the intermediate product 5 can comprise at least one feature map according to block 135.

[0044] In step 140, the contents 8 retrieved and / or aggregated according to the predefined rule 6 are merged with the intermediate product 5 according to a further predefined rule 9.

[0045] According to block 141, the further predetermined rule 9 may include adding the contents 8 retrieved and / or aggregated from the memory 7 to the intermediate product 5.

[0046] In step 150, the intermediate product 5# modified in this way replaces the original intermediate product 5 in further processing in the neural network architecture 1. That is, where the original intermediate product 5 was used, its modified version 5# is used.

[0047] In step 160, the desired output 3 is output from the output layer 15.

[0048] In the Fig. In the example shown in Figure 1, a control signal 170a is formed from the output 3 of the neural network architecture 1 in step 170.

[0049] In step 180, a vehicle 50, a driver assistance system 51, a robot 60, a quality control system 70, an area monitoring system 80, and / or a medical imaging system 90 are controlled with the control signal 170a.

[0050] For training the neural network architecture 1 including the extended processing described here, the Fig. In the example shown in Figure 1, in step 210, training examples 2a are processed into outputs 3 in the same way as the measurement data 2 is processed in later real-time operation. This means that the training examples 2a pass through the input layer 11 and the intermediate layers 12-14, with the intermediate product 5 of at least one of the intermediate layers 12-14 being modified to a new version 5# as previously described. The output 3 is then created in the output layer 15.

[0051] The outputs 3 thus obtained are evaluated in step 220 with a predetermined cost function L. For example, the training examples 2a can be labeled with target outputs 3a, and the cost function L can include a comparison of the actually obtained outputs 3 with the target outputs 3a.

[0052] In step 230, parameters 1a, which characterize the behavior of the neural network architecture 1, as well as at least additionally the contents 8 in the memory 7, are optimized with the goal of improving the evaluation by the cost function L. The fully optimized states of the parameters 1a* and the contents 8* are denoted by the reference symbols 1a* and 8*, respectively.

[0053] According to block 231, parameters 4a, which characterize the behavior of the given functional approach 4, can also be included in the optimization 230 with regard to the evaluation by the cost function L. This can, in particular, concern, for example, parameters that characterize the behavior of the evaluation network A. The fully optimized state of these parameters is denoted by the reference symbol 4a*.

[0054] Fig. 2 illustrates by way of example the extension of the processing chain in the neural network architecture 1 for the implementation of the method 100.

[0055] In the Fig. In the example shown in Figure 2, intermediate product 5 is a feature map output by a convolutional layer of a convolutional neural network. This feature map contains a separate two-dimensional channel for each filter kernel used in the convolutional layer. The channels are stacked along the third dimension, which denotes the specific filter kernel.

[0056] According to block 131a of method 100, the intermediate product 5 is reduced to a single two-dimensional channel using a 1x1 convolution operation. This two-dimensional channel is then rewritten into a one-dimensional vector, which, according to block 132, is processed by an evaluation network A configured as a fully interconnected neural network into parameters 8a for retrieving contents 8 from memory 7. These parameters 8a, in combination with the general approach of summing the contents 8 from memory 7 in a weighted sum according to block 133, form the rule 6 for retrieving and aggregating the contents 8 from memory 7. The entire processing chain from the intermediate product 5 to the weighted sum M of the contents 8 forms the specified functional approach 4.

[0057] In the Fig. In the example shown in Figure 2, memory 7 contains four different contents M1 to M4, each of which has the same dimensionality as intermediate product 5. Accordingly, the weighted sum M has the same dimensionality. It can therefore be added to intermediate product 5 according to the further specified rule 9 and step 140. This creates the modified intermediate product 5#, which replaces the original intermediate product 5 during further processing.

[0058] Fig. Figure 3 illustrates, using the example of an image classifier, how class-relevant information 8 is stored in memory 7 during training, and at the same time, the evaluation network A remembers when to retrieve which information 8 from memory 7. In the partial images af, the activation, i.e., the coefficient for the contribution to the weighted sum, is plotted against the index i of the respective content 8 in memory 7. In addition to the activation itself, the interval of the standard deviation is also plotted.

[0059] Panel a shows the average of the activations 8a for images of all possible classes.

[0060] Panel b shows activations 8a for images of the "Goldfish" class. The change compared to panel a, which occurs mainly between indices 3 and 5, is small but noticeable.

[0061] Subimage c shows the activations 8a for images of the "Cliff" class. This subimage c is qualitatively significantly different from subimages a and b.

[0062] Subpanel d shows activations 8a for images of the class "Plane." Compared to subpanel d, for example, activation 8a has been shifted between indices 1 and 2 on the one hand and indices 8 to 10 on the other.

[0063] Panel e shows the activations 8a for images of the class "Pug." Compared to panel d, the activation 8a for index 1 has increased significantly.

[0064] Subimage f shows the activations 8a for images of the class "Church." This subimage most closely resembles subimage d, with less activation 8a for indices 1 to 3 and more activation 8a for indices 8 to 10.

[0065] The class-specific distribution of activations 8a across the contents 8 in memory 7 shows that, firstly, memory 7 contains information 8 after training that is relevant to specific, concrete classes. In the tailor's analogy mentioned at the beginning, a sensible inventory of accessories and haberdashery has been created. Secondly, evaluation network A has also learned to sensibly combine the contents 8 stored in memory 7 in a modular system depending on the class. In the tailor's analogy, the tailor has thus learned exactly what he can pull from the shelf for a specific task and how to combine them sensibly.

[0066] Fig.Figure 4 illustrates how, when processing an image of a church as measurement data 2, an intermediate product 5 is changed into a modified intermediate product 5#. In addition to the before state 5 and the after state 5#, the change 5#-5 is also plotted separately. This change 5#-5 is most pronounced in the center of the church and emphasizes specific features of its architecture. Thus, information 8 has been specifically retrieved and aggregated from memory 7, which will advance the further analysis of an image of this class.

Claims

[1] Method (100) for processing measurement data (2) to an output (3) with regard to a given task with a neural network architecture (1) comprising a sequence of an input layer (11), several intermediate layers (12-14) and an output layer (15), comprising the steps: • the measurement data (2) are fed (110) to the input layer (11); • the output of the input layer (11) is processed (120) successively in the intermediate layers (12-14) and finally in the output layer (15), wherein the output of a layer (11-14) is fed to the next layer (12-15) as input; • the desired output (3) is output (160) from the output layer (15); where • a rule (6) for retrieving and / or aggregating contents (8) from a memory (7) is determined (130) on the basis of a predetermined functional approach (4) in accordance with an intermediate product (5) supplied as an output by at least one predetermined intermediate layer (12-14); • the contents (8) retrieved and / or aggregated in accordance with this provision (6) are merged with the intermediate product (5) in accordance with a further specified provision (9) (140) and • the intermediate product (5#) modified in this way replaces the original intermediate product (5) during further processing in the neural network architecture (1) (150). [2] Method (100) according to claim 1, wherein the predetermined functional approach (4) comprises aggregating the intermediate product (5) (131). [3] The method (100) of claim 2, wherein the aggregating comprises reducing the dimensionality of an intermediate product (5) in the form of a tensor by applying a convolution operation (131a). [4] Method (100) according to one of claims 1 to 3, wherein the predetermined functional approach (4) comprises • using an evaluation network (A), which is another trained neural network, to determine from the intermediate product (5) parameters (8a) for the weighting of several contents (8) from the memory (7) among each other (132) and • to aggregate (133) the contents (8) from the memory (7) according to these parameters (8a). [5] Method (100) according to claim 4, wherein value ranges for the parameters (8a) are determined by the choice of an activation function which translates activations of neurons in the output layer of the evaluation network (A) into outputs of these neurons (132a). [6] Method (100) according to one of claims 1 to 5, wherein the further predetermined rule (9) comprises adding (141) the contents (8) retrieved and / or aggregated from the memory (7) to the intermediate product (5). [7] Method (100) according to one of claims 1 to 6, wherein in the neural network architecture (1) at least one predetermined intermediate layer (12-14) is selected (134) which is closer to the output layer (15) than to the input layer (11). [8] Method (100) according to one of claims 1 to 7, wherein • a convolutional neural network with a plurality of convolutional layers that process their respective input by slidingly applying one or more filter kernels to one or more feature maps as output is selected as the neural network architecture (1) (105) and • the intermediate product (5) comprises at least one feature card (135). [9] Method (100) according to one of claims 1 to 8, wherein the neural network architecture (1) is designed as an image classifier and / or as a semantic segmenter for images (106). [10] Method (100) according to one of claims 1 to 9, wherein • measurement data (2) are selected (111) that were recorded with at least one sensor; • a control signal (170a) is formed (170) from the output (3) of the neural network architecture (1); and • a vehicle (50), a driver assistance system (51), a robot (60), a quality control system (70), a system (80) for monitoring areas, and / or a medical imaging system (90), with which the control signal (170a) is controlled (180). [11] Method (100) according to one of claims 1 to 10, wherein before supplying (110) the measurement data (2) to the input layer (11) • Training examples (2a) are processed in the same way to outputs (3) (210), • the outputs thus obtained (3) are evaluated with a given cost function (L) (220) and • Parameters (1a) that characterize the behavior of the neural network architecture (1), as well as at least additionally the contents (8) in the memory (7) are optimized (230) to the goal of improving the evaluation by the cost function (L). [12] Method (100) according to claim 11, wherein additional parameters (4a) which characterize the behavior of the predetermined functional approach (4) are included (231) in the optimization (230) with regard to the evaluation by the cost function (L). [13] A computer program comprising machine-readable instructions which, when executed on one or more computers and / or compute instances, cause the computer(s) and / or compute instances to carry out the method (100) according to any one of claims 1 to 12. [14] Machine-readable data carrier and / or download product with the computer program according to claim 13. [15] One or more computers and / or compute instances with the computer program according to claim 13, and / or with the machine-readable data carrier and / or download product according to claim 14.