Methods for processing sensor data
By dividing the machine learning algorithm into front-end and back-end parts for feature extraction and further processing, the computational and energy challenges of autonomous driving systems are addressed, achieving efficient data processing and adaptable sensor integration with reduced energy consumption.
Patent Information
- Application Number
- DE102024200417
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-17
- Publication Date
- 2025-07-17
AI Technical Summary
The high computational demands and energy consumption associated with processing sensor data from multiple cameras and sensors for autonomous driving, particularly due to the large data volumes and bandwidth requirements, pose significant challenges for central processing units, especially in level 4 autonomous systems.
The machine learning algorithm is divided into a front-end part for feature extraction and a back-end part for further processing, with the front-end part executed on a first computing unit within an environment detection unit, reducing data transmission and energy consumption by preprocessing and early fusion of sensor data before transmission to a central processing unit.
This approach drastically reduces data transmission and energy requirements, allowing for efficient processing and early fusion of sensor data, while enabling scalable and adaptable systems that can handle various sensor types and suppliers with minimal retraining effort.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The present invention relates to methods of sensor data using a machine learning algorithm, in particular a deep neural network, a computing unit and a computer program for carrying out the method. Background of the invention
[0002] In the field of autonomous or automated driving, sensors such as cameras, lidar, or radar sensors can be used. These sensors can capture sensor data or, in particular, image data, which can then be processed to obtain information for automated driving, including navigation. Disclosure of the invention
[0003] According to the invention, methods for processing sensor data as well as a computing unit and a computer program for implementing the methods are proposed, having the features of the independent patent claims. Advantageous embodiments are the subject of the subclaims and the following description.
[0004] The invention generally relates to sensors and environment detection units that have such sensors. As mentioned, such a sensor can be or have a camera or another image sensor. In addition, a sensor can also have or be a lidar sensor or a radar sensor. An ultrasound sensor or an audio sensor are also conceivable, for example. In general, such a sensor can be used to detect image and / or depth information about the environment. An image sensor can, for example, detect image data as sensor data. In the case of a camera, for example, this can be referred to directly as an image; in the case of a lidar or radar sensor, such sensor data generally also includes visual and / or spatial information, e.g. depth information, about the environment, but not, for example, in high resolution and in color, but instead, for example, in a so-called point cloud with distance information.Although the invention is described below particularly with reference to image sensors and image data, this applies generally to sensors and sensor data.
[0005] For autonomous or automated driving, a large number of cameras and other sensors (e.g., the aforementioned lidar and / or radar sensors) can be used, even for a single vehicle. Such image sensors can transmit the image data, usually as so-called RAW data, for further processing, e.g., to a central processing unit. Such a central processing unit typically receives the image data from multiple image sensors. Processing can involve various tasks or stages.
[0006] Preprocessing of the sensor or image data (or, in the case of the camera, the image) can be performed. The sensor data or image is cropped, for example, to match the regions of interest. The sensor data or image can then be rectified, for example, depending on the properties of the image sensor or camera. This can, for example, correct optical distortion. Furthermore, the sensor data, or in particular the image data, can be converted, for example, from the RGB to the YUV color space. It is also conceivable that an optical flow is calculated.
[0007] The possibly preprocessed sensor data or images can then be fed into a machine learning algorithm, specifically a so-called deep neural network (DNN). A deep neural network is an artificial neural network (ANN) with multiple layers or processing levels between the input and output layers.
[0008] The machine learning algorithm or DNN can then perform feature extraction in a first step. This involves extracting shapes or other basic features at a very low level from the sensor data or image. Further steps can then follow, particularly object recognition, semantic segmentation, or object classification. Classified objects, in the specific example of automated driving, are traffic signs, other vehicles, or people, for example. Optical flow can also be used to determine the trajectory of objects, for example.
[0009] Based on the detected objects, a vehicle trajectory can then be planned or calculated. In general, the output data of the machine learning algorithm or DNN can be used for automated driving.
[0010] Given the typically high number of cameras or image sensors—or sensors in general—and the resulting data rate at which sensor data must be transmitted (from the sensors to the central processing unit), this presents an enormous computational challenge. Data often needs to be transmitted in gigabits; the computing power of the central processing unit for, for example, Level 4 autonomous driving is expected to peak at 4,000 TOPs (Terra operations per second). This also presents enormous energy consumption challenges.
[0011] These challenges have been shown to be mitigated by in-memory computing. This uses a special form of memory that allows the DNN's MAC (multiplication and accumulation) operations to be computed without loading the coefficients. This significantly reduces memory bandwidth, the main power consumer.
[0012] Nevertheless, the remaining bandwidth required for the necessary data rate is still very high and accounts for the largest part of the energy consumption.
[0013] As has been further demonstrated, so-called chiplets can be used instead of large monolithic SoCs (SoC stands for System on Chip). A chiplet is a very small integrated circuit (IC) that contains a precisely defined subset of functions. It is designed to be combined with other chiplets on a so-called interposer in a single package. Multiple chiplets can also be combined with each other, for example.
[0014] This trend is driven by the need for even more computing power, but also by the effect of declining yields on very large chips. For example, the larger a silicon chip, the greater the likelihood of a manufacturing defect.
[0015] By combining multiple chiplets into a larger system, the best semiconductor process can be selected for a specific chiplet's task. Systems become scalable because chiplet designs can be reused.
[0016] However, communication between these chiplets is still required. This is easy for, say, a power supply or timing device, but connecting an AI or GPU chiplet to the system and having it process multiple camera streams is challenging.
[0017] However, a large processing unit and the transmission of raw data are preferred because they not only allow the sensor data or image data itself to be processed, but also allow the information or sensor data from multiple sensors to be combined to obtain a more consistent perception (by fusing the image data from multiple image sensors). This usually occurs before the sensor data is incorporated into the machine learning algorithm or DNN. In this respect, this is also referred to as early fusion.
[0018] Another aspect is that a typical system design allows replacing the sensor or camera of one type with another type and / or from another supplier by simply adjusting the parameters in the mentioned pre-processing step.
[0019] Against this background, it is proposed to divide the machine learning algorithm or DNN, or more precisely, its multiple processing levels, into at least two parts. A first or so-called front-end part, which then carries out, for example, feature extraction, and a second or so-called back-end part, which then takes care of everything else. The first part, or the front-end, is then outsourced to a first computing unit, which can, for example, be part of an environmental detection unit (and which receives the sensor data acquired by the actual sensor). This brings with it various advantages. The amount of data to be transmitted can thus sometimes be drastically reduced and the energy budget can be distributed more evenly across the system, which, for example, reduces the cooling requirements for the main computing unit.
[0020] Thus, sensor data, which has been acquired by a sensor (which, for example, may also be part of the environment detection unit), is initially provided in a first computing unit. This sensor data is then provided as first input data for the machine learning algorithm or DNN. In one embodiment, however, the aforementioned preprocessing may also be provided; here, the sensor data is then preprocessed on the first computing unit to obtain the first input data, which is then provided.
[0021] The first input data is then processed by executing at least some of the processing levels of the machine learning algorithm or DNN on the first processing unit. In particular, only the aforementioned first part of the processing levels is executed on the first processing unit. However, as has been shown, all processing levels can also be executed on the first processing unit, for example, in the case of not particularly extensive machine learning algorithms or DNNs.
[0022] Initial output data obtained by executing at least part of the processing levels of the machine learning algorithm, i.e. in particular only the first part, are then provided for further use.
[0023] Processing the first input data by executing at least part, or only the first part, of the processing levels of the machine learning algorithm on the first computing unit includes, for example, feature extraction based on the first input data (which in turn corresponds to or is based on the image data). It is also conceivable that (simple) object recognition or object classification could also be performed based on the first input data.
[0024] Depending on the type of environmental detection unit, this may mean, for example, that the first computing unit mentioned would have to be adapted, replaced or added.
[0025] In one embodiment, the first output data are transmitted to a first computing unit, e.g., the aforementioned central computing unit, for further processing by the aforementioned second part of the processing levels of the machine learning algorithm. There, the first output data obtained by the first computing unit can then be provided as second input data. These second input data are then processed by executing the second part of the processing levels of the machine learning algorithm on the second computing unit. Second output data obtained by executing the second part of the processing levels of the machine learning algorithm are then provided for further use, in particular for automated driving.
[0026] The processing by executing the second part of the processing levels of the machine learning algorithm then includes, for example, object recognition or object classification based on the second input data. In this case, in particular, no object recognition or object classification is performed on the first computing unit, or, for example, only partial or preliminary object recognition or object classification is performed.
[0027] In summary, such a first processing unit (e.g., as part of the environment detection unit) can, in one embodiment, perform preprocessing and, if necessary, (simple) object detection and / or (simple) object classification. This turns the environment detection unit into an intelligent environment detection unit, or a camera into an intelligent camera that can, for example, directly detect a pedestrian or an obstacle and forward this information to a second processing unit, such as the aforementioned central processing unit. This would make this camera an "intelligent" camera for high-volume L2 input systems.
[0028] Likewise, in one embodiment, such a first processing unit can execute the preprocessing as well as some processing levels (layers) of the machine learning algorithm or DNN. The (complete) processing previously performed on the central processing unit is thus split up, so that, for example, the feature extraction step is separated from the rest of the machine learning algorithm.
[0029] Feature extraction on the first processing unit also makes it possible, for example, for the first output data (i.e., after feature extraction) from multiple environmental detection units or their first processing units to be merged on the second processing unit. Although feature extraction is a particularly useful point at which the resulting first output data can be merged, this is also conceivable with other first output data.
[0030] Another aspect that may need to be considered is the amount of data that passes from one processing level or layer of the machine learning algorithm or DNN to the next. This allows the first and second parts to be selected such that the amount of data for the first output data is smaller than the amount of data for the first input data. In other words, the amount of data to be transmitted from the first computing unit to the second computing unit is smaller than the amount of data initially present due to the successful execution of at least part of the processing levels.
[0031] In the aforementioned embodiments, the environment detection unit - or specifically, for example, an image detection unit or camera - can still have a conventional image processor (ISP) that performs the preprocessing, but, for example, in a smaller and / or slower version compared to a central ISP in, for example, a central processing unit. This makes it possible to make all necessary adjustments if an environment detection unit or image detection unit is to be or must be replaced, for example, by a different brand. It is also possible to remove the ISP and retrain the machine learning algorithm or DNN if a new type or new supplier of the environment detection unit is used. With a smaller machine learning algorithm, as is particularly possible within the scope of the present invention, this effort is much lower than retraining the entire machine learning algorithm or DNN.A small DNN could be defined as one that would manage with 20 TOP / s and 20 MB weights. Conveniently, the small DNN fits on the camera or the environment detection unit, so only this part would need to be retrained.
[0032] The estimated computing power for such a step, for example, is only about 20 TOPs. With an in-memory computing power of about 40 TOPs / W (W stands for watt), this means that the computing power outsourced to the image acquisition unit increases the energy budget of the environment or image acquisition unit by 0.5 watts, which roughly corresponds to the amount of energy that a powerful transmitter in a network would require instead. Ideally, the amount of data is already drastically reduced in the camera or the environment acquisition unit, and therefore no longer needs to be transmitted to a central control unit. The transmitter is an interface (power) driver that requires energy to transmit the data. If this energy is saved, calculations can be performed instead, and much less data can be transmitted.
[0033] An advantage of the above-mentioned embodiments compared to known variants is in particular that the amount of data that must be transmitted to the central processing unit is drastically reduced and yet early fusion is possible.
[0034] A concrete example of this is a 12-megapixel (MP) camera, which can be used, for example, for Level 4 autonomous driving. At 30 frames per second and 16 bits / pixel, this results in a bandwidth of 12*16*30 Mbit / s or 5.7 Gb / s. Therefore, it is particularly important to note that for best results, before the preprocessing or feature extraction, the processing level (layer) after which the splitting takes place should be optimized for a small data output (small output feature map). This can be done, for example, within the framework of so-called neural architecture search (NAS).
[0035] A system for data processing according to the invention or a computing unit is set up, in particular in terms of programming, to carry out a method according to the invention, e.g. in one of the described embodiments, in particular insofar as the steps are carried out by the first or the second computing unit.
[0036] The invention also relates to an image capture unit with such a computing unit, in particular the said first computing unit, and with an image sensor.
[0037] The implementation of a method according to the invention in the form of a computer program or computer program product with program code for carrying out all method steps is also advantageous, since this entails particularly low costs, in particular if an executing control unit is also used for other tasks and is therefore already present. Finally, a machine-readable storage medium is provided with a computer program stored thereon as described above. Suitable storage media or data carriers for providing the computer program are, in particular, magnetic, optical and electrical memories, such as hard disks, flash memories, EEPROMs, DVDs, and others. Downloading a program via computer networks (Internet, intranet, etc.) is also possible. Such a download can be wired or cable-based or wireless (e.g. via a WLAN network, a 3G, 4G, 5G or 6G connection, etc.).
[0038] Further advantages and embodiments of the invention will become apparent from the description and the accompanying drawings.
[0039] The invention is illustrated schematically in the drawing using an embodiment and is described below with reference to the drawing. Short description of the drawings Fig. 1 schematically shows a vehicle in an environment for explaining the invention. Fig. 2 shows schematically an arrangement for explaining the invention Fig. 3a shows schematically an arrangement for explaining the invention in one embodiment Fig. 3b schematically shows a computing unit in one embodiment. Fig. 4 shows schematically a sequence of a method for explaining the invention in one embodiment. Embodiment(s) of the invention
[0040] In Fig. Figure 1 schematically illustrates, by way of example, a vehicle 100 in an environment for explaining the invention. The vehicle 100 is traveling, for example, on a road 140. A person 142 is supposed to be on the road, and a traffic sign 144 is supposed to be located at the side of the road, for example.
[0041] The vehicle 100 has, for example, a control or regulating unit 102 and a drive unit 104 (with wheels) for moving the vehicle 100, e.g., along a movement path 108, which here runs, for example, along the road 140 or a lane of the road.
[0042] Furthermore, the vehicle 100 has a central processing unit or a system 110 for data processing, e.g. a control unit, by means of which data can be exchanged with a higher-level system, e.g. via an indicated radio connection.
[0043] As already mentioned, environmental sensors such as cameras or radar sensors can be used in autonomous or automated driving. For example, the vehicle has an environmental detection unit 120, which has a camera as a sensor, and an environmental detection unit 130, which has a radar sensor as a sensor. An environmental detection unit can, in particular, have the actual sensor as well as a computing unit; nevertheless, the term "camera" is often used, for example, when referring to the image detection unit with the actual image sensor (camera) and the computing unit.
[0044] By means of such environment detection units, information about the environment can be recorded, for example also about the person 142 or the traffic sign 144. Within the scope of, for example, object recognition, the person 142 or the traffic sign 144 can then be identified or recognized, based on which the vehicle then navigates, for example, but in simple cases, for example, also simply brakes.
[0045] In Fig. Figure 2 schematically shows an arrangement with multiple environment detection units, such as can be used in a vehicle. Three environment detection units 220a, 220b, 220c, each with a camera as a sensor, and three environment detection units 230a, 230b, 230c, each with a radar sensor as a sensor, are shown by way of example. A central processing unit 210 is also shown, on which, for example, a machine learning algorithm 254, in particular a deep neural network (DNN), can be executed. The environment detection units are all connected to the central processing unit 210 via suitable data connections 250, e.g., a data bus.
[0046] First, a conventional way of processing sensor data will be explained.
[0047] For example, the sensor data (or, in the case of the camera, the image) can first be preprocessed. The sensor data or image can be cropped, for example, according to the regions of interest. The sensor data or image can then be rectified, for example, depending on the properties of the sensor or camera. This can, for example, correct optical distortion. Furthermore, the image data can be converted, for example, from the RGB to the YUV color space. It is also conceivable that an optical flow can be calculated.
[0048] This is done, for example, by means of a so-called ISP or, for example, a computing unit in the relevant environmental detection unit.
[0049] The preprocessed sensor data or images are then transmitted via data connections 250 to the central processing unit 210, where they can then be fed to the machine learning algorithm or DNN 254. As already mentioned, the machine learning algorithm has several layers or processing levels between the input and output layers. This is represented here by various parallel, vertical lines of varying lengths.
[0050] The machine learning algorithm or DNN can then perform feature extraction in a first step. This involves extracting shapes or other basic features at a very low level from the sensor data or image. Further steps can then follow, particularly object recognition, semantic segmentation, or object classification. Classified objects, in the specific example of automated driving, are traffic signs, other vehicles, or people, for example. Optical flow can also be used to determine the trajectory of objects, for example.
[0051] In Fig. Figure 3a also schematically illustrates an arrangement with multiple environment detection units, such as can be used in a vehicle, specifically to explain one embodiment of the invention. Three environment detection units 320a, 320b, 320c, each with a camera as a sensor, are shown by way of example.
[0052] For the environment detection unit 320a, the sensor 322a (here a camera) and a computing unit 324a are shown as examples, hereinafter also referred to as the first computing unit. The same applies to the other environment detection units. Furthermore, a central computing unit 310 is shown, hereinafter referred to as the second computing unit.
[0053] As mentioned, the multiple processing levels of the machine learning algorithm or DNN can be divided into a first part 354.1 and a second part 354.2. Overall, the multiple processing levels of the first and second parts can be similar to those of the machine learning algorithm 254 according to Fig. 2 correspond.
[0054] In the specific example, each of the environment detection units is assigned such a first part, i.e. the first part of the processing levels is executed on the respective first processing unit, a second part is assigned to the central processing unit, i.e. the second part of the processing levels is executed on the second processing unit.
[0055] The second part 354.2 again comprises several separate sections, each of one or more processing levels, which will be discussed in more detail later.
[0056] In Fig. 3b, the environment detection unit 320a is shown again with functional units, namely the actual camera 322a, the first part 354.1 of the DNN, an ISP 326a and an interface (e.g., I / O) 328a. The first computing unit is not explicitly shown here. However, the ISP can, for example, be part of the first computing unit. Via the interface 328a, data from the environment detection unit 320a can be transmitted via a data connection (not shown here) (e.g., the data connection 250 according to Fig. 2) to the second processing unit (there then, for example, also via a corresponding interface).
[0057] As already mentioned, it is conceivable, for example, with simple DNNs, to execute the entire DNN on the first processing unit, i.e., the environment detection unit. In this case, the second part 354.2 would not be present in the second processing unit.
[0058] In Fig. Figure 4 schematically illustrates a process flow for explaining one embodiment of the invention. This will be explained in particular with reference to Figures 3a and 3b.
[0059] In step 400, sensor data 402 are provided in the first computing unit, which have been acquired by means of a sensor, as is the case, for example, in Fig. 3a for the environment detection unit 320a.
[0060] In step 404, for example, a preprocessing of the sensor data 402 can be carried out on the first computing unit in order to obtain first input data 406. The preprocessing can be carried out, for example, on or by means of the ISP as in Fig. 3b. The first input data thus obtained are then provided in step 408.
[0061] In step 410, the first input data is then processed by executing the first part of the processing levels of the machine learning algorithm on the first computing unit. This can also be done as in, for example, Fig. 3a for the environment detection unit 320a. This then includes, for example, a feature extraction 412. The data obtained in this way are provided as first output data 414, step 416.
[0062] In a step 418, these first output data 414 are then transmitted to the second computing unit, as is the case, for example, in Fig. 3a. There, the first output data, step 420, is then provided as second input data 422.
[0063] In step 424, these second input data are then processed, in particular by executing the second part of the processing levels on the second computing unit, as described, for example, in Fig. 3a is shown.
[0064] In particular, however, as already mentioned and as also in Fig. As shown in Figure 3a, the first output data are received from multiple environmental detection units or first computing units in the second computing unit and then provided. For this purpose, the first output data or second input data can then be fed to various processing levels of the second part 354.2 in order to merge the data.
[0065] Fusion, or more specifically sensor data fusion, can generally be described as the linking of the output data from several sensors. The goal is usually to obtain information of better quality. This improvement in quality is achieved, for example, by gaining new insights through linking, e.g., the temporal correlation of individual sensor data or sensor signals. It is advantageous if the individual sensor signals are not pre-processed too heavily. Otherwise, it can happen, for example, that a DNN_A that evaluates a sensor_A discards information that is not relevant for the result_A, but would have been of interest for a result_B from a DNN_B (with sensor_B).
[0066] Data processed in this way can then, step 426, be used as second output data 428 for further use, e.g. in the context of automated driving as in relation to Fig. 1 mentioned.
Claims
[1] Method for processing image data using a machine learning algorithm, in particular a deep neural network, wherein the machine learning algorithm has several processing levels, comprising: Providing (400), in a first computing unit (324a), sensor data (402) which have been acquired by means of a sensor (322a), Providing (408) the sensor data as first input data or first input data (406) obtained based on the sensor data, Processing (410) the first input data by executing at least part of the processing levels of the machine learning algorithm on the first computing unit (324a), Providing (416) first output data (414) obtained by executing at least part of the processing levels of the machine learning algorithm for further use. [2] The method of claim 1, further comprising: preprocessing (404) the sensor data on the first computing unit (324a) to obtain the first input data (406). [3] The method of claim 1 or 2, wherein processing the first input data by executing at least the part of the processing levels of the machine learning algorithm on the first computing unit (324a) comprises feature extraction (412) in the first input data. [4] Method according to one of the preceding claims, wherein a data quantity of the first output data is less than a data quantity of the first input data. [5] Method according to one of the preceding claims, wherein the sensor comprises one of the following sensors: a camera, a radar sensor, a lidar sensor, an ultrasonic sensor, an audio sensor. [6] Method according to one of the preceding claims, wherein only a first part (354.1) of the processing levels of the machine learning algorithm is executed on the first computing unit (324a), and wherein the first output data (414) are transmitted (418) to a second computing unit (310) for further processing by a second part (354.2) of the processing levels of the machine learning algorithm. [7] A method for processing image data using a machine learning algorithm, in particular a deep neural network, wherein the machine learning algorithm has several processing levels, comprising: Providing (420), in a second computing unit (310), one or more sets of second input data (422), wherein the second input data have each been obtained as first output data from a respective first computing unit, wherein the respective first output data have each been obtained by processing respective first input data based on respective sensor data acquired by a respective sensor by executing a respective first part of the processing levels of the machine learning algorithm, Processing (424) the one or more sets of second input data, and Providing (426) second output data (428) for further use. [8] The method of claim 7, wherein processing the one or more sets of second input data is performed by executing a second portion of the processing levels of the machine learning algorithm on the second computing unit, and wherein the second output data is obtained by executing the second part of the processing levels of the machine learning algorithm. [9] The method of claim 7 or 8, wherein a plurality of sets of second input data are obtained, and wherein processing the plurality of sets of second input data comprises fusing the plurality of sets of second input data. [10] A method according to any one of claims 7 to 9, wherein the first output data of the one or at least one of the plurality of sets has been obtained according to a method according to any one of claims 1 to 6. [11] Computing unit (324a, 310) configured to execute the steps performed by the first computing unit or the second computing unit of a method according to any one of the preceding claims. [12] Environment detection unit (320a) comprising a computing unit according to claim 11 and a sensor. [13] A computer program comprising instructions which, when executed by a computer, cause the program to carry out the method steps of a method according to any one of claims 1 to 10 when executed on the computer. [14] A computer-readable storage medium on which the computer program according to claim 13 is stored.