Method and neural network for partitioning distributed detection of three-dimensional objects

By partitioning three-dimensional objects on sensor computing nodes and aggregate computing nodes, the data redundancy and scalability problems of multi-sensor environment detection in automated driving are solved, and a higher level of semantic meaning and system scalability are achieved, real-time processing capabilities are improved and bandwidth savings are saved.

CN120564007APending Publication Date: 2025-08-29ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510224481.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-27
Filing Date
2025-02-27
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

In the detection of multi-sensor environments, especially in the field of automated driving, there are problems such as insufficient data redundancy, inflexible scalability and lack of semantic meaning on computing nodes, which leads to difficulties in system design.

Method used

A distributed neural network is adopted to partition three-dimensional objects on sensor computing nodes and aggregation computing nodes, and a neural network is used to aggregate and process data on a unified three-dimensional representation, providing higher-level semantic meanings and system scalability, reducing bandwidth requirements.

Benefits of technology

It realizes a higher level of semantic meaning and system scalability, improves the system's affordability and real-time processing capabilities, and saves bandwidth.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564007A_ABST
    Figure CN120564007A_ABST
Patent Text Reader

Abstract

The invention relates to a method for partitioning the distributed detection of three-dimensional objects from detected data, in which a neural network (200) distributed over at least one sensor computing node (310) and one aggregation computing node (320) is used, at least one sensor which detects data from its environment is associated with each of the sensor computing nodes (310), the sensor computing node (310) forwards the evaluated data to the aggregated computing node (320) and partitions a common or unified three-dimensional representation. The invention also relates to a neural network suitable for carrying out the method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a method for partitioning the distributed detection of three-dimensional objects and to a neural network suitable for carrying out the method. Background Art

[0002] The method presented herein has particular application in the field of automated driving. Automated driving requires the detection or perception of the vehicle's environment. To this end, cameras are typically used as sensors to detect the environment and, from the detected data, particularly image data, to detect relevant objects in the environment. Known methods for detecting three-dimensional objects from multi-sensor / multi-view data transform the abstract representations of the data from each individual sensor into a common or unified three-dimensional representation. These three-dimensional representations are then collected or aggregated and used to predict three-dimensional objects.

[0003] These approaches are usually implemented on a single machine or a single computing node. As the number of participating sensors increases, multi-sensor / multi-view object detection methods on a single computing node have huge limitations due to the large amount of raw sensor data received, lack of redundancy, and inflexible scalability.

[0004] A known approach for distributed implementation across multiple computing nodes, primarily for reducing latency in edge devices and cloud components, is distributed DNN (deep neural network) inference. DNN models are typically partitioned at a low level of abstraction, for example, by splitting sensor-space input data from a single sensor through dynamic load balancing mechanisms and processing each part on other computing nodes.

[0005] This could mean, for example, that an image is cut into different sub-planes or patches, and each sub-plane is processed on a separate compute node, with some communication between intermediate compute nodes for correct results. Furthermore, there are rudimentary distributed inference strategies for classification of multi-sensor / multi-view data, but no rudimentary distributed inference strategies exist for 3D object detection, where low-level sensor spatial features are aggregated across multiple compute nodes before being transformed into a 3D representation.

[0006] The application of DNN partitioning at a low level of abstraction and the distribution of multiple computation nodes known from the prior art enables scalability, however, there is a lack of sufficient semantic meaning in the sensor space where the partitioning occurs. This makes it difficult to achieve convincing semantic redundancy and scalability, making correct system design difficult. Summary of the Invention

[0007] Against this background, a method according to the invention and a neural network according to the invention are proposed. Specific embodiments will follow from the subsequent description.

[0008] A method for partitioning during the distributed detection of three-dimensional objects based on detected sensor data is proposed, wherein a neural network distributed on at least one sensor computing node and one aggregation computing node is used, each sensor computing node is assigned at least one sensor that detects data from its environment, the sensor computing nodes forward the evaluated data to the aggregation computing node, and the partitioning is performed on a common or unified three-dimensional representation.

[0009] A neural network is a network of artificial neurons that has a natural biological prototype. A natural neural network is a network of neurons in the nervous system. Deep neural networks (DNNs) mimic the workings of the human brain.

[0010] The neural network proposed here is designed, for example, as a deep neural network and is provided for carrying out a method of the type described here.

[0011] Therefore, a method for partitioning distributed implementations of 3D object detection is proposed. This method provides a direct relationship between partitions and spatial regions, thereby delivering higher-level, more human-interpretable meaning. This enables better, more targeted, robust, and scalable system-level design based on 3D geometry. Furthermore, this method enables bandwidth savings.

[0012] It is proposed to directly use a common or unified three-dimensional representation of three-dimensional objects ( Figure 1230 in the example, the method for detecting three-dimensional objects is partitioned and distributed, and the individual sensor representations are aggregated in the three-dimensional representation. Therefore, the three-dimensional representation transmitted between the computing nodes is structurally similar to the real three-dimensional world and therefore has a human-understandable geometric meaning at a higher level of abstraction. In actual practice, this results in partitioning into sensor computing nodes, which process sensor data and transform it into a three-dimensional representation. All three-dimensional representations are then sent to an aggregation computing node, which calculates a common or unified three-dimensional representation. Utilizing the geometric and mathematical size or characteristics of the three-dimensional representation provides advantages in partitioning, such as:

[0013] A) Simple system scalability is achieved for both additional sensors per sensor computing node and additional sensor computing nodes connected to the aggregation computing node. In both cases, the remaining distributed system partitioning, including interfaces, is not affected.

[0014] B) If there are multiple overlapping sensors on different computing nodes and the same range in the 3D world is examined, then an explicitly resilient design can be achieved.

[0015] C) Deterministic and fixed partitioning of the algorithm or load and the transmission of fixed-size data blocks between computing nodes, instead of dynamic load balancing methods or variable-size data blocks, enables real-time processing.

[0016] D) Because the 3D representation is smaller than the raw sensor data or low-level abstraction features, or can be constructed smaller than them, implicit bandwidth savings are achieved.

[0017] E) Explicit bandwidth savings are achieved between compute nodes by aggregating the 3D representations of all sensors at (on) each sensor compute node, and / or by transmitting only the portion of the 3D representation that is within the field of view of the contributing sensors.

[0018] F) The DNN-based 3D object detection method can be trained as a whole (end-to-end) without having to consider the proposed partitioning method at training time. Partitioning can be applied afterwards.

[0019] Further advantages and embodiments of the present invention are apparent from the description and the accompanying drawings.

[0020] Of course, the features mentioned above and those yet to be explained below can be used not only in the respectively stated combination but also in other combinations or on their own without departing from the scope of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 A block diagram shows a basic processing chain of a method for detecting three-dimensional objects.

[0022] Figure 2 A block diagram shows the partitioning of a method for detecting three-dimensional objects for a distributed implementation.

[0023] Figure 3 Block diagram illustrating partitioning of a method for detecting three-dimensional objects for distributed implementation, including pre-aggregation of three-dimensional representations at sensor compute nodes.

[0024] Figure 4 An exemplary three-dimensional representation is shown based on a grid from a bird's-eye view. DETAILED DESCRIPTION

[0025] The invention is schematically illustrated with the aid of embodiments in the drawings and will be described in more detail below with reference to the drawings.

[0026] Figure 1 A block diagram illustrating the basic processing chain of a method for detecting three-dimensional objects is shown. The diagram shows sensor input data 110, a three-dimensional object 120, a neural network 200 for detecting three-dimensional objects, a sensor feature encoder 210, a unit for transforming the sensor features into a three-dimensional representation 220, aggregation of three-dimensional representations from multiple sensors 230, and a decoder 240 for the three-dimensional representation.

[0027] From now on, the detailed description will be based on an exemplary application case of automated driving with a sensor configuration consisting of multiple cameras that detect the world from different perspectives or lines of sight, wherein a DNN-based three-dimensional object detector is used.

[0028] The method is not limited to this specific application. It can also be applied to sensors other than cameras, such as lidar, radar, ultrasound, and the like. It can also be applied to sensor configurations designed for other applications where the goal is environmental monitoring, such as in robotics more generally, or in monitoring systems with sensors distributed over a wider area. The method is particularly applicable to sensors whose fields of view overlap.

[0029] Furthermore, it must be taken into account that the transmitted data blocks may have a fixed size.

[0030] The signal chain and processing chain underlying the non-distributed 3D object detection method are Figure 1One or more camera images are input as sensor input data 100 to a network 200 for detecting three-dimensional objects, which predicts a three-dimensional object 120 as output, for example as a bounding box. The sensor input data 110 is provided by a sensor, for example a camera.

[0031] Each camera image is encoded into low-level image space features using a sensor feature encoder 210 and subsequently transformed from image space into a three-dimensional representation using unit 220. The three-dimensional representations of each camera from multiple cameras are then aggregated into a common or unified 3D representation 230, which is then decoded into a unified 3D object list using a decoder 240. In addition to the 3D object list, the 3D representation can also serve as the basis for other detection tasks and object representations, such as segmentation, instance masking, or spatial occupancy, for which the proposed method can also be applied.

[0032] The proposed partitioning strategy for the distributed implementation of the 3D object detection method is Figure 2 In display.

[0033] Figure 2 A neural network 150 for detecting three-dimensional objects is shown, which is distributed across sensor computing nodes 310, namely computing node A 310a and computing node B 310b, and aggregator computing node 320. Optionally, sensor computing node 310 can also simultaneously assume the role of aggregator computing node 320. More than one aggregator computing node 320 can also be present, for example to provide redundancy. The number of sensor computing nodes 310 is arbitrary and can be freely varied. On the hardware side, the computing nodes can be represented, for example, by chips, dedicated computing accelerators, or complete computer systems. It is proposed that after the transformation from sensor space to a three-dimensional representation by means of unit 220, partitioning is performed and that the three-dimensional representation data 115 is sent from sensor computing node 310 to aggregator computing node 320.

[0034] If multiple cameras are assigned to a sensor computing node 310b, as in Figure 33D representations 230 for these cameras can be pre-aggregated at sensor compute node 310b before being sent to aggregation node 320, as shown in FIG. 3 (compute node B), to save bandwidth. The number of sensors per sensor compute node 310 is arbitrary and can vary from one sensor compute node 310 to another. Sensor types / modalities can also be arbitrarily mixed, even within a sensor compute node 310. From a mathematical perspective, pre-aggregation only requires an associative aggregation function, such as sum, average, minimum, or maximum, which are common choices.

[0035] The unified three-dimensional representation 230 may have smaller memory requirements than the raw sensor data 110 or features at a lower level of abstraction.

[0036] Furthermore, it may be provided that only parts of the unified three-dimensional representation 230 are transmitted to the aggregate computing node 230 .

[0037] In order to effectively use some of the advantages of the proposed partitioning strategy, other extensions and characteristics of the three-dimensional representation 230 and the determination of the aggregation function become important. The following description and details exemplarily consider the three-dimensional space representation by the unit 220 as a projection from a bird's-eye view onto a two-dimensional Euclidean grid, such as Figure 4 However, the concept can be transferred to other 3D spatial representations, such as various projections, voxel, and polar coordinate representations.

[0038] Figure 4 An exemplary three-dimensional representation 400 of a vehicle environment is shown from a bird's-eye view, having a number of cells 402. The scale, number of cells, and objects shown, in this case a vehicle 404, are arbitrary.

[0039] Let’s explore the previously mentioned advantages again in more detail:

[0040] A) The numerical range of the individual elements of the grid, viewed from above, is constant relative to the number n of participating cameras. How this is achieved depends on the aggregation function used. For example, the minimum or maximum function is essentially constant here. If these representations R i The average value of the number of participating sensors and the grid in the bird's-eye view can be expressed as i With each R i Send together, making applicable:

[0041] R aggregiert =∑ i R i / ∑i w i

[0042] The scaling is invariant with respect to n. All effective R i The same can be added there. This aggregation is beneficial for sensor / computing node failure tolerance (B) and also for system scalability if more computing nodes or sensors are to be added.

[0043] B) Because the geometry of the 3D representation resembles the relevant parts of the 3D world, system robustness and redundancy can be directly configured to meet specific criteria in a clear and understandable manner. By processing all sensor data at a sensor computing node, field of view and redundancy considerations are directly associated with that sensor computing node. Combined with a suitable aggregation function, such as the one described in point A, redundancy can be added to the system, thereby increasing robustness, without changing the overall partitioning, interfaces, and system design. For example, if the area in front of the ego vehicle in an automated driving system is to be redundantly covered by independent cameras and computing nodes, this criterion can be easily reflected in the system design using the proposed partitioning strategy: two groups, each with an appropriate number of cameras, are created, each covering the relevant field of view, thereby providing sensor redundancy. Each group is assigned its own sensor computing node, whose outputs are connected to one or more aggregation computing nodes. This ensures coverage in the event of a sensor or computing node failure.

[0044] Since the sensor data from each individual sensor is distributed across the compute nodes, low-level partitioning mechanisms do not provide such simple traceability. Therefore, a compute node error may affect not only a known three-dimensional region, but also any set of outputs or all outputs. Also note that the proposed partitioning enables multiple instances of the aggregate compute node to be distributed across independent hardware to increase throughput when needed. Sensor compute node outputs must simply be sent to all aggregate compute nodes in parallel.

[0045] C) Known distributed DNN inference methods for applications in edge and cloud computing, such as those consisting of mobile apps and cloud servers, attempt to reduce latency through adaptive strategies, where variable portions of data are sent from edge devices to the cloud. While this provides flexibility for dynamically changing systems, such as those caused by changes in data bandwidth or computing power, it is not feasible for real-time applications. Therefore, the proposed partitioning is based on fixed-size data blocks, where the size of the data to be transferred and the number of computational steps to be performed on a specific compute node are known and determined in advance. This is fundamental for real-time applications, where predictability is essential.

[0046] D) In ​​automated driving applications, raw sensor data, especially camera data, tends to be significantly larger due to its high resolution. Consequently, low-level features are also quite large, as their size increases linearly with sensor resolution. Consequently, known partitioning strategies for transmitting raw sensor data or low-level features at high resolution require significant bandwidth between compute nodes. Conversely, the size of the bird's-eye view grid is independent of sensor resolution and can be designed based on the needs of the object detection function, making it smaller than the raw or low-level data. This reduces bandwidth requirements between compute nodes.

[0047] E) If a suitable aggregation function is used, such as that described in A, all bird's-eye view grids of the cameras assigned to a sensor computing node can be aggregated in advance at the sensor computing node. This reduces the amount of data transmitted to the aggregation computing node and makes it constant, for example, independent of the number of cameras assigned to the sensor computing node. Furthermore, the bird's-eye view representation in the device memory typically enables efficient access to specific image blocks, such as rectangular regions representing a specific area of ​​the world.

[0048] Therefore, only the image blocks covered by the field of view of the participating cameras need to be transmitted to the aggregation node, along with a small amount of metadata describing the location of the image block within the overall bird's-eye view grid. Regions that are more complex and more closely cropped than rectangular are also possible, but the required amount of metadata describing their shape must be considered. If real-time capabilities are important, for sensor configurations with a fixed field of view in the 3D representation, the metadata can be transmitted to the aggregation node before the online implementation to avoid additional bandwidth usage.

[0049] F) Model training for DNN-based 3D object detection methods is typically a non-trivial task. The proposed partitioning strategy does not adversely affect DNN training, thereby preventing increased complexity.

Claims

1. A method for partitioning when performing distributed detection of a three-dimensional object based on detected data, wherein: Using a neural network (200) distributed on at least one sensor computing node (310) and one aggregation computing node (320), Each sensor computing node (310) is equipped with at least one sensor that detects data from its environment. The sensor computing node (310) forwards the evaluated data to the aggregation computing node (320), and Partitioning is performed on a common three-dimensional representation.

2. The method according to claim 1, wherein A plurality of sensors are connected to at least one sensor computing node (310).

3. The method according to claim 1 or 2, wherein: The method is carried out in a motor vehicle which is configured for automated operation.

4. The method according to claim 1 or 2, wherein: The method is implemented in a field selected from the group consisting of a robot and a monitoring system.

5. The method according to any one of claims 1 to 4, wherein: The sensor is selected from the group consisting of a camera, a lidar sensor, a radar sensor, and an ultrasonic sensor.

6. The method according to any one of claims 1 to 5, wherein: Use sensors with overlapping fields of view.

7. The method according to any one of claims 1 to 6, wherein: Probing is performed corresponding to the partitioning performed.

8. The method according to any one of claims 1 to 7, wherein: Probing is performed in real time.

9. The method according to any one of claims 1 to 8, wherein: The detected objects are displayed as a representation (400) from a bird's-eye view.

10. The method according to any one of claims 1 to 9, wherein The method is implemented in a deep neural network.

11. A neural network, wherein The neural network is distributed on at least one sensor computing node (310) and one aggregation computing node (320), and is configured to execute the method according to any one of claims 1 to 10.

12. The neural network according to claim 11, wherein The neural network is constructed as a deep neural network.