Efficient simultaneous inference computation for multiple neural networks

By merging identical or similar units of multiple neural networks on a hardware platform and merging the outputs in a central working process, the problem of high computational cost of multiple neural networks on mobile devices is solved, achieving computational efficiency and storage savings, and improving the scalability and security of the system.

CN113379028BActive Publication Date: 2026-08-25ROBERT BOSCH GMBH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110253719.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-10
Filing Date
2021-03-09
Publication Date
2026-08-25
Estimated Expiration
2041-03-09

AI Technical Summary

Technical Problem

Simultaneous inference computations of multiple neural networks on mobile devices are computationally and energy-intensive, and existing technologies struggle to effectively combine identical or similar neural network units to reduce redundant work and storage space.

Method used

Security and independence are ensured by identifying and merging identical or similar units in multiple neural networks on a hardware platform, performing unique inference computations, and merging outputs in a central working process.

Benefits of technology

It significantly reduces the computational cost and energy consumption of multiple neural networks performing inference calculations simultaneously, saves storage space, and avoids performance degradation when expanding the classification system, thereby improving the system's scalability and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113379028B_ABST
    Figure CN113379028B_ABST
Patent Text Reader

Abstract

Method for performing inference computations of a plurality of neural networks (1, 2) on a hardware platform, each neural network (1, 2) having a plurality of neurons (11-21) that aggregate inputs (3a-3g) into network inputs with transfer functions characterized by weights and process network inputs into activations (4a-4d) with activation functions, the method comprising: identifying at least one unit (5a-5c) that comprises one or more transfer functions and / or complete neurons (11-21) and that is present in at least two of the networks (1, 2) in the same or similar form according to a pre-given criterion; performing a unique inference computation for the unit (5a-5c) on the hardware platform such that the unit (5a-5c) provides a set of outputs (6); processing the set of outputs (6a-6c) in the respective network (1, 2) as output of the unit (5a-5c). Method for simultaneously performing a plurality of applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to simultaneous inference computation (Inferenzberechnung) for multiple neural networks on a common hardware platform. Background Technology

[0002] Neural networks trained multiple times are used for classification tasks, such as identifying objects in images. Such neural networks possess strong generalization capabilities. For example, after training with a sufficient number of images containing specific objects (e.g., vehicles), they can also identify new variations of these objects (e.g., vehicles placed on the market only after said training). Neural networks for object recognition are disclosed, for example, by WO 2019 / 162241 A1.

[0003] Meanwhile, applications using neural networks have also entered mobile devices such as smartphones. Therefore, photo management apps, for example, are already equipped with neural networks in a standard manner not only in the Apple iOS ecosystem but also in the Google Android ecosystem. These neural networks categorize photos based on the objects contained within them, as stored on the smartphone. For instance, you can enter "license plate number" into the search bar and it will display all photos where you should see the license plate number.

[0004] Such inference calculations require high computational resources, and this high computational resource consumption, especially in the case of mobile devices, comes at the cost of battery life. Summary of the Invention

[0005] Within the scope of this invention, a method for performing inference computation of multiple neural networks on a hardware platform has been developed. Each of these neural networks has multiple neurons. These neurons respectively aggregate one or more inputs into network input using a transfer function characterized by weights. Then, an activation function processes the network input to activate the corresponding neurons.

[0006] Generally, the neural network can be configured, for example, as a classifier to assign observational data, such as camera images, thermal images, radar data, LiDAR data, or ultrasonic data, to one or more pre-given categories. These categories can, for example, represent objects or states that need to be detected in the observed area. The observational data can, for example, originate from one or more sensors mounted on a vehicle. The actions of a driver assistance system or a system for at least partially automated vehicles, which are matched to specific traffic conditions, can then be derived based on the category assignments provided by the neural network. The neural network can, for example, be a convolutional neural network (CNN) subdivided into multiple layers.

[0007] The method identifies at least one unit, which comprises one or more transfer functions and / or complete neurons and exists in at least two of these networks in the same form or in a similar form according to pre-given criteria. Unique inference computation is performed for the unit on the hardware platform, such that the unit provides an output set. This output set is further processed in the corresponding network to become the output of the unit.

[0008] It has been shown that this approach can significantly reduce the computational cost and energy consumption of inference computations performed simultaneously by multiple neural networks. Furthermore, it saves storage space. This is particularly applicable when the neural networks operate using the same or similar input data or perform similar tasks.

[0009] Therefore, for example, on smartphones, the aforementioned photo management applications, installed in a standard-compliant manner, are also combined with other applications that similarly perform inference calculations on image data. For example, there are applications that can be used to search a collection of photos based on the face of a specific person, or applications that can be used to calculate how a face has looked in the past or how it will look in the future based on a facial image.

[0010] Furthermore, when developing complex systems for classifying objects from images or noise from audio data, distributing the task among multiple parallel neural networks is an effective strategy. Especially in the subsequent construction of such systems, this ensures that the system can only be improved through further training, and that further training on one aspect does not lead to side effects such as deteriorating performance on another aspect. In this way, when scaling such classification systems, it also minimizes the need to modify the debugged and tested code again.

[0011] Therefore, classification systems for audio data, for example, could include applications dedicated to speech recognition, applications dedicated to identifying gasoline engine noise, and applications dedicated to identifying acoustic fire alarms. In particular, the first layers can operate very similarly, with these first layers used to extract basic features from the audio data. Performing inference computations separately for all applications would incur a great deal of unnecessary duplication. The method described above can save most of this additional cost. Although the computer-based identification of similar units that can be merged through inference computation within the network consumes computation time once for each new configuration (Konstellation) of the neural network to be evaluated simultaneously, this cost is quickly offset by avoiding duplication of work.

[0012] The classification system can also be extended with less overhead for programming and training. For example, if it is desired to extend to recognizing diesel engine noise, a neural network currently used for gasoline engine noise can be replicated as a template, adapted in terms of architecture if necessary, and then trained using diesel engine noise. Here, the ability to recognize gasoline engine noise already acquired through training remains unimpaired because the new network is designed independently for diesel engines. However, by utilizing the commonalities (Gemeinsamkeit) with networks currently used for gasoline engines within the scope of the method, the additional overhead of separate implementations for diesel and gasoline engines is minimal compared to networks trained jointly for both diesel and gasoline engines from the outset.

[0013] When regions in two different neural networks can be reasonably identified as two occurrences of the same unit depends on the specific application. Below are some examples of standards (Kriterium) that can be used individually or in combination when identifying these units in a computer-based manner.

[0014] For example, the pre-given criterion may include: the unit receives the same or similar inputs in the at least two networks. The degree to which the two sets of inputs fed to the two neural networks in their respective inference computations can be considered "similar" to each other can, in particular, depend, for example, on the degree to which these inputs relate to physical observations of the same scene using one or more sensors.

[0015] In the examples mentioned above where different types of noise should be identified, all the audio data used could originate from the same microphone arrangement. However, a specialized microphone could also be used to identify fire alarms, specifically one that is particularly sensitive in a frequency range common to fire alarms.

[0016] Similarly, even if, for example, images of a scene recorded by multiple cameras are recorded from different angles, these images can still be "similar" to each other enough to enable the merging of inference computations.

[0017] Alternatively or in combination, the similarity between two input sets can also depend on the extent to which they originate from the same, similar, and / or overlapping physical sensors. Thus, for example, even if one image was recorded using the front-facing camera of a smartphone and another image was recorded using the rear-facing camera of the same smartphone, the basic steps for extracting raw features from the images are similar.

[0018] Alternatively or in combination, the pre-given criteria may also include: characterizing the units in the at least two networks by the same or similar weights of the transfer functions or neurons. These weights reflect the "knowledge" of the network acquired during training. Here, the weights do not necessarily need to be numerically similar. The similarity of two sets of weights may also depend, for example, on the degree to which the distributions formed by these sets of weights are similar to each other.

[0019] For example, a neuron in the first network might receive the same weighted input from all four neurons in the previous layer, while another neuron in the second network might receive the same weighted input from all three neurons in the previous layer. Thus, when comparing the two networks, these weights differ not only in their quantity but also in their numerical value (Zahlenwert). However, a common pattern remains: using inputs from all the neurons available in the previous layer, and weighting these inputs equally among themselves.

[0020] In a particularly advantageous design, the multiple input sets obtained by the unit in the respective network are combined into a single input set. In this way, even if the inputs used in the networks differ slightly, these inference computations for the unit can still be merged into a single inference computation. Thus, the same image data can be processed, for example, at a resolution of 8 bits per pixel in an application using a first neural network and at a resolution of 16 bits per pixel in a second application using a second neural network. This difference does not become an obstacle to merging the inference computations. The results of the inference computation are generally, to a certain extent, in the case of neural networks, insensitive to small changes in the input, provided that these small changes are not "adversarial examples," which are deliberately constructed to cause misclassification, for example, through optically insignificant changes in the image data. The extent to which the results of the inference computation are insensitive to small changes in the input or even the weights is demonstrated during the training of a specific neural network.

[0021] Similarly, multiple sets of weights representing the units in the corresponding network can be combined into a unique set of weights. Even if the same neural network is trained twice using the same training data, the weights will not be exactly the same. This is therefore unpredictable, since the training typically starts with random initial values.

[0022] Combining multiple input sets and / or weight sets into a unique set for inference computation can include, for example, forming a summary statistic, such as the mean or median, from multiple different sets on an element-by-element basis.

[0023] In a specific configuration of neural networks to be evaluated simultaneously, one or more common units can be identified, and the inference computations of these units can be combined separately. For example, the first network may share a unit with the second network, while the second network may share another unit with the third network.

[0024] In another particularly advantageous design, the neural networks are fused into a single, unified neural network such that each unit appears only once within that network. Inference computation is then performed on the fused neural network on the hardware platform. This fusion makes the overall architecture of these neural networks more common. Furthermore, it reduces the overhead of branching to common inference computations at each inference iteration and the overhead of retransmitting results back to participating networks.

[0025] As described above, in a particularly advantageous design, the input to the neural network can include the same or similar audio data. Thus, the neural network can be trained separately to classify different types of noise. As mentioned above, the noise classification system can then be easily extended to recognize other types of noise.

[0026] This also applies to another particularly advantageous design where the input to the neural network includes the same or similar image data, thermal image data, video data, radar data, ultrasonic data, and / or LiDAR data, and where the neural network is trained separately to classify different objects. The scalability already mentioned, without risking degradation in understanding previously learned objects, is particularly important in the context of at least partially automated driving vehicles in road traffic. If multiple neural networks are evaluated in parallel, for example, it is possible to retrospectively equip the vehicle with the recognition of newly introduced traffic signs by legislators without compromising the recognition of traffic signs known to date. This can be a powerful argument when obtaining official permits for such vehicles.

[0027] The measurement data can be obtained through a physical measurement process and / or through partial or complete simulation of such a measurement process and / or through partial or complete simulation of a technical system capable of being observed using such a measurement process. For example, realistic images of the situation can be generated by means of computational tracking of light beams (“raytracing”) or by using neural generator networks (e.g., Generative Adversarial Networks, GANs). Here, cognition from the simulation of the technical system, such as the location of a specific object, can also be introduced as auxiliary conditions. The generator network (e.g., conditional GAN, cGAN) can be trained with regard to the targeted generation of images that satisfy these auxiliary conditions.

[0028] Typically, control signals can be formed from the inference calculations of one or more neural networks. These control signals can then be used to control vehicles and / or systems for quality control of mass-produced products and / or systems for medical imaging and / or access control systems.

[0029] In the applications on the mobile devices mentioned at the beginning, different applications are typically allocated separate, abschotten storage areas to prevent them from accessing each other. This is to avoid the applications from interfering with each other, or, for example, malicious applications potentially spying on or corrupting the databases of other applications. On the other hand, this security mechanism also prevents the discovery of identical or similar units within the neural networks running in the applications.

[0030] Therefore, the present invention also relates to another method for simultaneously executing multiple applications on a hardware platform. This method proceeds by allocating a storage area to each application on the hardware platform, the storage area being protected from access by other applications. Within each application, at least one inference computation is performed for at least one neural network.

[0031] Within the scope of this method, each application requests the inference computations it requires from a central worker process. Here, the application specifies both the neural network to be evaluated and the inputs to be processed relative to the worker process. Within the central worker process, all requested inference computations are performed using the method described above. The worker process then returns the outputs of these inference computations to the requesting applications, respectively.

[0032] Therefore, from the perspective of each application, a request for inference computation is similar to a regular call to a subroutine or library. The worker process is responsible for finding matching common units in the neural networks to be evaluated simultaneously at runtime and merging these inference computations accordingly.

[0033] Within the storage area used by the worker process, inference computations requested by different applications are performed simultaneously. However, this does not violate or even completely abolish the security model that separates these applications from each other. The applications have no opportunity to execute potentially malicious binary code within the storage area. The worker process only accepts specifications (Spezifikations) for the neural network and its input. Attached Figure Description

[0034] Other measures to improve the invention are further illustrated below, together with the description of preferred embodiments of the invention based on the accompanying drawings.

[0035] Figure 1 An embodiment of a method 100 for inference computation of multiple neural networks 1, 2 is shown; Figure 2 Exemplary neural networks 1 and 2 with common units 5a-5c are shown. Figure 3 It shows the result of Figure 2 Network 7 is a fusion of networks 1 and 2 shown. Figure 4 An embodiment of a method 200 for simultaneously executing multiple applications A and B is shown; Figure 5 An exemplary hardware platform 30 is shown during the execution of method 200. Detailed Implementation

[0036] Figure 1This is a schematic flowchart of an embodiment of a method 100 for inference computation of multiple neural networks 1 and 2. In step 110, identical or similar units 5a-5c in networks 1 and 2 are identified. In step 120, unique inference computations are performed on each of these units 5a-5c on the hardware platform 30, thereby providing outputs 6a-6c for each unit. In step 130, these outputs 6a-6c are processed in networks 1 and 2 as outputs of units 5a-5c. This means that the outputs 6a-6c are further computed in networks 1 and 2 as if inference computations were performed independently for units 5a-5c in networks 1 and 2 respectively. Therefore, inference results 1* and 2* of networks 1 and 2 are ultimately generated. In step 140, control signals 140a can be formed from these inference results. In step 150, the control signals are used to control vehicle 50 and / or classification system 60 and / or system 70 for quality control of mass-produced products and / or system 80 for medical imaging and / or access control system 90.

[0037] Figure 2 Two exemplary neural networks, 1 and 2, are shown. Network 1 consists of neurons 11-16, and processes inputs 3a-3d into activations 4a and 4b of neurons 15 and 16, which form the output 1* of network 1. Network 2 consists of neurons 17-21, and processes inputs 3e-3g into activations 4c and 4d of neurons 20 and 21, which form the output 2* of network 2.

[0038] Since the input 3e of neuron 17 is the same as the input 3c of neuron 13, neurons 13 and 17 are considered as similar units 5a in the two networks 1 and 2, respectively.

[0039] Since the input 3f of neuron 18 is the same as the input 3d of neuron 14, neurons 14 and 18 are considered as another unit 5b that is similar in the two networks 1 and 2, respectively.

[0040] Furthermore, neurons 15 and 20 obtained similar data and were therefore considered as another unit 5c that was similar in the two networks 1 and 2.

[0041] For units 5a, 5b, and 5c, only one inference calculation needs to be performed each.

[0042] like Figure 3 As shown, networks 1 and 2 can then be merged with each other, such that neurons 13 and 17, 14 and 18, or 16 and 20 are combined (zusammenlegen). For network 7 merged in this way, the inference computation can be performed on hardware platform 30. Together, this provides the outputs 1* and 2* of networks 1 and 2.

[0043] Figure 4 This is a schematic flowchart of an embodiment of a method 200 for simultaneously executing multiple applications A and B, which require inference computations from neural networks 1 and 2, respectively.

[0044] In step 210, each application A and B requests the inference computation required by its application from the central worker process W, where networks 1 and 2, and the inputs 3a-3d and 3e-3g to be processed are respectively passed (übergeben). In step 220, worker process W performs the inference computation using method 100 described above, and generates outputs 1* and 2*. In step 230, these outputs 1* and 2* are returned to applications A and B.

[0045] Figure 5 An exemplary hardware platform 30 is shown during the execution of method 200. Separate storage regions 31 and 32 are assigned to applications A and B, respectively, and storage region 33 is assigned to worker process W.

[0046] exist Figure 5 In the example shown, application A makes a request. Figure 2 The inference computation of network 1 is shown, while application B simultaneously requests... Figure 2 The inference computation of network 2 shown in the figure. The workflow uses... Figure 3 The diagram shows a network 7 formed by fusing two networks, 1 and 2. The fused network 7 processes the inputs 3a-3d of network 1 into output 1*, which can be used as the output of network 1 by application A. Similarly, the fused network 7 processes the inputs 3e-3g of network 2 into output 2*, which can be used as the output of network 2 by application B.

Claims

1. A method (100) for performing inference computation of multiple neural networks (1, 2) on a hardware platform (30), wherein each of the neural networks (1, 2) has multiple neurons (11-21), the neurons respectively aggregating inputs (3a-3g) into network inputs using a transfer function characterized by weights, and processing the network inputs into activations (4a-4d) using an activation function, the method comprising the following steps: • Identify (110) at least one unit (5a-5c), the at least one unit comprising one or more transfer functions and / or complete neurons (11-21), and existing in the same form or in a similar form according to a pre-given criterion in at least two different networks (1, 2) of the network; • Perform (120) unique inference computation on the hardware platform (30) for the units (5a-5c) in the at least two different networks (1, 2) such that the units (5a-5c) provide an output set (6). • The output set (6a-6c) is processed (130) into the output of the unit (5a-5c) in the corresponding network (1, 2). in, The inputs to the neural networks (1, 2) include the same or similar audio data, and the neural networks (1, 2) are respectively trained to classify different noises, and / or The inputs to the neural networks (1, 2) include the same or similar image data, thermal image data, video data, radar data, ultrasound data and / or LIDAR data, and the neural networks (1, 2) are respectively trained to classify different objects.

2. The method (100) according to claim 1, wherein, The pre-given criteria include: the units (5a-5c) obtain the same or similar inputs (3a-3g) in at least two networks (1, 2).

3. The method (100) according to claim 2, wherein, The similarity between the two input sets (3a-3g) depends on the extent to which the inputs relate to physical observations of the same scene using one or more sensors.

4. The method (100) according to claim 2 or 3, wherein, The similarity between two input sets (3a-3g) depends on the extent to which the inputs originate from the same, similar and / or overlapping physical sensors.

5. The method (100) according to any one of claims 1 to 3, wherein, The pre-given criteria include: characterizing the unit (5a-5c) in the at least two networks (1, 2) by the same or similar weights of the transfer function or neurons (11-21).

6. The method (100) according to claim 5, wherein, The similarity between two weight sets depends on the degree to which the distributions formed by the weight sets are similar to each other.

7. The method (100) according to any one of claims 1 to 3, wherein, The multiple input sets (3a-3d, 3e-3g) obtained by the units (5a-5c) in the corresponding networks (1, 2) are combined (121) into a unique input set (3a-3d, 3e-3g). ).

8. The method (100) according to any one of claims 1 to 3, wherein, The multiple weight sets representing the units (5a-5c) in the corresponding networks (1, 2) are combined (122) into a unique weight set.

9. The method (100) according to any one of claims 1 to 3, wherein, The neural networks (1, 2) are thus fused (123) into a single neural network (7) such that the units (5a-5c) appear only once in the neural network, and wherein the inference computation (124) is performed on the fused neural network (7) on the hardware platform (30).

10. The method (100) according to any one of claims 1 to 3, wherein, The results of inference calculations from one or more neural networks form (140) control signals (140a), and the control signals (140a) are used to control (150) a vehicle (50) and / or a system (70) for quality control of mass-produced products and / or a system (80) for medical imaging and / or an access control system (90).

11. A method (200) for simultaneously executing multiple applications (A, B) on a hardware platform (30), wherein at least one inference computation is performed for at least one neural network (1, 2) within each application (A, B), wherein a storage region (31, 32) is allocated to each application (A, B) on the hardware platform (30), the storage region being protected from access by other applications (A, B), the method comprising the steps of: • Each application (A, B) requests (210) the inference computation required by the application from the central working process (W) given the specified neural network (1, 2) and the input to be processed (3a-3d, 3e-3g); • Within the central working process (W), the inference calculation (220) is performed using the method (100) according to any one of claims 1 to 10; • The working process (W) will output the inference calculation (1) ,2 ) respectively return (230) to the application (A, B) that made the request.

12. A computer program containing machine-readable instructions that, when executed on one or more computers, cause the one or more computers to perform the method (100, 200) according to any one of claims 1 to 11.

13. A machine-readable data carrier having a computer program according to claim 12.

14. A computer equipped with a computer program according to claim 12 and / or equipped with a machine-readable data carrier according to claim 13.

Citation Information

Patent Citations

  • Real-time object detection using depth sensors

    WO2019162241A1

  • Neural Network Processor Incorporating Multi-Level Hierarchical Aggregated Computing And Memory Elements

    US20180285718A1