Identification of target objects in the field of automated driving

The neural network with a transformer encoder addresses the challenges of non-linear error propagation and temporal dependencies in automated systems by efficiently processing sensor data to identify relevant objects for vehicle control, improving safety and functionality.

DE102024200723A1Pending Publication Date: 2025-07-31ROBERT BOSCH GMBH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
DE102024200723
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-26
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Existing automated systems face challenges in processing sensor data for fully or partially automated vehicles due to non-linear error propagation when transforming 2D data to 3D and inadequate modeling of temporal dependencies, limiting their effectiveness in dynamic environments.

Method used

A method using a neural network with a transformer encoder and multi-head self-attention layer processes sensor data to form an input tensor with variable dimensions, determining a probability distribution for object relevance through a classifier, enabling efficient identification of target objects for control functions.

Benefits of technology

This approach allows for flexible and resource-efficient processing of variable input data, providing reliable probability distributions for object relevance, enhancing the control of automated systems like vehicles by prioritizing relevant objects for functions such as lane keeping and emergency braking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to identifying relevant target objects in the environment of an at least partially automated system. For this purpose, a situation analysis is used to identify properties corresponding to the objects. These object properties are fed into a model, which determines a probability distribution for the relevance of the objects using a neural network with a transformer encoder.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present invention relates to a method for identifying target objects, a device for identifying target objects, a driver assistance system with such a device, a vehicle with such a driver assistance system as well as a computer program product and a computer-readable storage medium. State of the art

[0002] One goal of developing new systems, particularly in the field of automated driving, is to improve safety. The present invention will preferably be described below in connection with a fully or partially automated vehicle—however, it is not limited to this. Rather, the basic principle of the invention can also be applied to any other fully or at least partially automated systems, such as a robotic system.

[0003] Systems for fully or at least partially automated vehicle operation are becoming increasingly important. In addition to the goal of fully automated vehicles, numerous driver assistance system concepts already exist that support the driver in one or more functions or at least temporarily assume these functions completely. For example, lane keeping systems are known that can perform longitudinal vehicle control. Data from video sensors and / or radar sensors are preferably used for this purpose. From the data from these sensors, a suitable trajectory for the vehicle can be determined, and relevant objects in the vehicle's surroundings can be identified.

[0004] The paper by R. Kanjee, AK Bachoo, and J. Carroll, "Vision-based Adaptive Cruise Control using pattern matching," 2013 6th Robotics and Mechatronics Conference (RobMech), Durban, South Africa, 2013, pp. 93-98, doi: 10.1109 / RoboMech.2013.6685498, describes a camera-based ACC system that uses a single camera to determine the distance between the preceding vehicle and the ACC vehicle. Pattern matching with lane detection is used for vehicle detection. The vehicle and range detection algorithms are validated using real-world data.

[0005] The paper "Attention Is All You Need" by Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin (arXiv:1706.03762) describes a network architecture called a transformer. This network architecture is based exclusively on attention mechanisms and completely dispenses with repetitions and convolutions. Disclosure of the invention

[0006] The present invention provides a method and a device for identifying target objects with the features of the independent patent claims, as well as a driver assistance system, a vehicle with such a driver assistance system, and the computer-related objects mentioned in the other independent claims. Further advantageous embodiments are the subject of the dependent patent claims.

[0007] According to a first aspect, the present invention provides a method for identifying target objects for an at least partially automated system. The method comprises a step for providing sensor data from at least one sensor system. In particular, sensor data can be provided from any suitable sensor systems for environment detection. This sensor data can include, for example, radar data, image data, map data, or the like. The sensor data can be aggregated from sensors of different types. Furthermore, the method comprises a step for identifying objects and optionally structures in the environment of the system. The objects and / or structures can be identified in particular using the provided sensor data. Furthermore, the method comprises a step for determining relations between the identified objects. Furthermore, the method comprises a step for forming an input tensor.The input tensor can, in particular, comprise a predetermined number of object properties for a variable number of objects. The object properties can be formed according to the specific relations between individual objects or between objects and structures. The input tensor can thus, for example, be formed from a matrix, wherein the dimension for the number of identified objects is variable. Furthermore, one or more dimensions for the object properties can be constant. The method further comprises a step for determining an output tensor. The output tensor can, in particular, be determined using the input tensor. The output tensor specifies a probability distribution for the relevance of the identified objects.In particular, the output tensor is determined using a classifier, in particular a neural network, further in particular a neural network with a transformer encoder. Finally, the method comprises a step of identifying at least one target object. The target objects can be prioritized with regard to their relevance, in particular using the output tensor and / or a prioritization algorithm. For example, criteria can be implemented in the prioritization algorithm that represent a predeterminable set of conditions, such as collision-free operation by maintaining safety distances, etc.

[0008] According to a further aspect, the present invention provides a device for identifying target objects for an at least partially automated system. The device is particularly designed to carry out the method according to the invention. The device comprises a receiving device and a plurality of processing devices. The receiving device is designed to receive sensor data from at least one sensor system and to provide it. This can in particular be sensor data from an optical system, such as from one or more cameras and / or sensor data from a radar system. A first processing device is designed to identify objects and optionally also structures in the surroundings of the vehicle using the provided sensor data.A second processing device is configured to determine relations between the identified objects and / or between identified objects and structures. A third processing device is configured to form an input tensor. The input tensor can comprise a predetermined number of object properties for each identified object. The object properties can be formed, in particular, from the relations between the previously determined objects or, optionally, between objects and structures. A fourth processing device is configured to determine an output tensor using the input tensor. The output tensor specifies a probability distribution for the relevance of the identified objects.The output tensor can be determined in particular using a classifier, in particular a neural network, further in particular a neural network with a transformer encoder. A fifth processing device is designed to identify at least one target object. One of the target objects can be specified as relevant, in particular by means of a prioritization algorithm and / or by means of the output tensor.

[0009] According to yet another aspect, the present invention provides a computer-readable storage medium. The storage medium comprises instructions that, when executed by a computer, cause the computer to carry out the method according to the invention.

[0010] According to yet another aspect, the present invention provides a computer program product. The computer program product comprises a computer program with instructions that, when executed by a computer, cause the computer to carry out the method according to the invention.

[0011] According to a further aspect, the present invention provides a driver assistance system comprising a device according to the invention for identifying target objects and a control device. The control device is designed to control at least one driver assistance function of the at least partially automated system (e.g., vehicle, robot) using the identified target object. In particular, the at least one driver assistance function of the vehicle can comprise a lane keeping function, a lane change function, a distance keeping function, a cruise control function, and / or a braking function, in particular an emergency braking function.

[0012] According to a further aspect, the present invention provides an at least partially automated vehicle with a driver assistance system according to the invention.

[0013] The present invention is based on the realization that the evaluation of sensor data for fully or at least partially automated systems, such as systems for a fully or at least partially automated vehicle, presents numerous challenges. On the one hand, the transformation of sensor data from a sensor measuring in 2D space, such as a video sensor, into a three-dimensional space can lead to non-linear error propagation. On the other hand, the temporal dependencies of the provided sensor data must be adequately modeled. Conventional systems can quickly reach their limits in this regard.

[0014] Therefore, one idea of ​​the present invention is to take this finding into account and to create a concept for processing sensor data for use in a fully or at least partially automated system, such as a fully or partially automated vehicle. To this end, one idea of ​​the present invention is to provide a concept based on a classifier, in particular a neural network, which delivers reliable and meaningful results even for a variable number of input variables, in particular for the objects to be evaluated. In particular, the present invention creates a concept that can be implemented in a resource-efficient manner.This provides the possibility of accessing a multitude of input data, such as the individual objects and their properties, as well as output data, such as a probability distribution for the relevance of the individual objects. This distinguishes the inventive concept significantly from conventional systems, in which a complex neural network, which is difficult to implement and acts as a kind of black box, generates a final control variable, such as a steering angle or similar, from raw sensor data, or from classic classification approaches that are limited to a predetermined number of input data. According to the invention, the number of input data, also referred to as objects, is not limited but variable.

[0015] The data processing using the inventive concept can be implemented with relatively low hardware requirements. Furthermore, the available input and output data can be used simultaneously for multiple application areas (e.g., lane keeping systems, emergency braking assistance, evasive maneuvering, etc.) if necessary.

[0016] The sensor data for identifying target objects can, in principle, originate from any suitable sensors as part of at least one sensor system. For example, the sensor data can be provided by an optical system, such as a camera or LiDAR. In particular, special cameras for stereo recordings and / or cameras for non-visible light, such as infrared, are also possible. Furthermore, the sensor data can originate from a radar system, for example. Depending on the application, any other suitable systems, such as ultrasound or similar, are also possible. This sensor data can be provided via any suitable interface, such as a bus system or other suitable communication technologies.

[0017] Different suitable approaches can be used to identify objects and, optionally, structures in the environment. Objects can be considered to be elements in the system's environment - particularly dynamic or moving ones - which may be relevant for controlling the system. For example, in a system for a fully or at least partially autonomous vehicle, the objects can be other road users, such as vehicles, cyclists, pedestrians, or similar. However, other dynamic or static elements such as obstacles, traffic lights, or similar can also be considered to be objects. Features in the system's environment - particularly stationary or static ones - which may be relevant for the behavior of the (own) system and / or the other elements in the system's environment can be considered to be structures. For vehicles, for example, road markings, roadside developments, traffic signs, or similar.be viewed as structures. From such structures, for example, possible movement paths or similar for the respective elements can be derived. However, it is understood that, depending on the application, other elements or features may also be relevant as objects or structures. The objects and / or structures can be identified, for example, using suitable detection methods, for example, using a pattern or object recognition algorithm. In principle, any other suitable methods are also possible, for example, methods using a neural network or similar.

[0018] Through a suitable situation analysis using the information from objects and structures, relations between the identified objects and / or between the objects and the structures can then be determined. For example, it is possible to determine which structures are relevant for a specific object and may be important for a future direction of movement. The determination of relations can be carried out with reference to the own system (e.g., ego vehicle). Thus, the determination of relations can include the determination of relations between the own at least partially automated system (e.g., ego vehicle) and other road users in the traffic scene. In addition, the method can include the determination of relations between the identified objects and / or identified structures.

[0019] Since the environment and the objects and / or structures therein can change continuously, especially in a dynamic environment, the number of identified objects in the system's environment can also vary. Consequently, the input tensor for determining the probability distribution can also specify a variable number of objects. Accordingly, the neural network for processing the input tensor is designed to process such an input tensor with a variable number of objects and associated object properties. In particular, with a smaller number of objects, the input tensor does not have to be enlarged to a tensor of fixed size, e.g. by adding fictitious values ​​or zero values. This enables flexible, efficient and thus resource-saving and fast processing.

[0020] For such processing of input tensors with variable sizes and, in particular, a variable number of objects, neural networks with a transformer encoder, in particular a transformer encoder with a multi-head self-attention layer and, preferably, a fully connected feedforward network, have proven advantageous. Such types of neural networks are currently primarily used in the field of text and speech processing. However, such types of neural networks have not yet been established in the field of controlling automated systems, in particular for controlling fully or at least partially autonomous vehicles.

[0021] For a final determination of the probability distribution for the relevance of the objects, an additional instance can be placed downstream of the transformer encoder. For example, this additional downstream instance could be an instance with a softmax function. This instance for determining the probability distribution can also include one or more (fully connected) feedforward networks.

[0022] As a result of processing the input tensor, a probability distribution for the relevance of the identified objects is determined. In particular, this makes it possible to determine which of the identified objects are of high relevance for controlling the at least partially automatically operating system. For controlling a vehicle, for example, objects, in particular road users, can be identified that are of high relevance in the vehicle's own lane (of the vehicle's own system, i.e., the vehicle) and, if applicable, in an adjacent lane. This could, for example, be a vehicle driving directly ahead or a vehicle in an adjacent lane that could merge into the vehicle's own lane or that could be on a potential trajectory when the vehicle changes lanes.However, it should be understood that, depending on the application, any other objects could also be considered relevant. For example, from the probability values ​​obtained by the output tensor, one object or a predetermined number of objects can be selected that have the highest probability values.

[0023] According to one embodiment, the classifier comprises a neural network (NN). The NN may comprise a transformer encoder, preferably with a multi-head self-attention layer. According to a specific embodiment, the output tensor is formed using feed-forward neural networks, in particular fully connected feed-forward neural networks, which are connected downstream of the transformer encoder.

[0024] According to one embodiment, a softmax component is connected downstream of the feed-forward neural networks. In particular, such a softmax component can be implemented together with other feed-forward networks and, for example, form a prediction head.

[0025] According to one embodiment, the relationships between the identified objects and optionally structures are determined using a classifier, in particular a neural network. Such a neural network is positioned upstream of the transformer encoder. For this purpose, a neural network with multiple linear layers can be used, for example. Alternatively, image processing algorithms can be used to preprocess the received or provided sensor data.

[0026] According to one embodiment, the input tensor is formed for a predetermined or predeterminable number of time steps. In this way, the temporal progression of the scene to be monitored in the system's environment can also be taken into account.

[0027] According to one embodiment, the objects include stationary or moving objects. Furthermore, identifying objects in the system's environment can include identifying other relevant objects, in particular road users, other vehicles, people, or units. Furthermore, depending on the application, any other types of objects can of course also be considered.

[0028] According to one embodiment, identifying structures in the environment of the system includes identifying lines, in particular roadway boundaries, peripheral structures, and / or elements characteristic of an operational function, in particular a control function, of the system. Furthermore, depending on the application, any other types of structures in the environment of the system can of course also be identified and considered.

[0029] According to one embodiment, the method comprises a step of executing at least one function, in particular a control function. This control function can be executed, in particular, using the probability distribution of the output tensor.

[0030] For example, the control function can select one or more objects, in particular a predetermined number, based on the probability distribution and consider them for the control function. The function can also be configured as an output function to output outputs to an output unit, for example, to output warnings as a vibration signal (e.g., on the steering wheel), as an acoustic and / or visual signal, or as signals on a display.

[0031] The above embodiments and further developments can be combined with one another as desired, where appropriate. Further embodiments, further developments, and implementations of the invention also include combinations of features of the invention not explicitly mentioned above or described below with respect to the exemplary embodiments. In particular, those skilled in the art will also add individual aspects as improvements or additions to the respective basic forms of the invention. Short description of the drawings

[0032] Further features and advantages of the invention are explained below with reference to the figures. These show: Fig. 1: a schematic representation of an exemplary scene to explain the method for identifying target objects; Fig. 2: a flowchart of how a method for identifying target objects according to an embodiment may be based; Fig. 3: a schematic representation of an architecture that may underlie a determination of target objects according to one embodiment; and Fig. 4: a schematic representation of a block diagram for a device for identifying target objects according to an embodiment. Description of embodiments

[0033] The present invention is described below using exemplary embodiments. The exemplary embodiments preferably relate to fully or at least partially autonomous vehicles.

[0034] In principle, however, the present invention is also applicable to any other fully or at least partially autonomously operating systems. For example, the inventive concept can also be applied to industrial systems, such as manufacturing systems, production systems, or household appliances such as household robots (vacuum and / or floor-mopping robots). Therefore, the description of vehicles is not intended to represent a limitation of the present invention.

[0035] Fig. 1 shows a schematic representation of a scene with an at least partially autonomously driving vehicle 10, in which a method for identifying target objects can be implemented according to one embodiment. The vehicle 10 can, for example, comprise one or more sensors 11 that are part of a sensor system. The sensor system can be arranged in and / or on the vehicle and / or outside (for example on an infrastructure). The sensor system can be distributed. These sensors 11 can, for example, be used to sense an environment of the vehicle 10. For example, the sensors 11 can be one or more cameras, radar sensors, LiDAR, ultrasonic sensors, or any other suitable sensors. The sensors can, for example, detect objects 20-i in the environment of the vehicle 10.Depending on the type of sensor 11, an object 20-i can be detected one-dimensionally (e.g., only distance), two-dimensionally, or three-dimensionally. If necessary, a relative velocity of objects 20-1 can also be determined, e.g., using Doppler radar or similar.

[0036] The objects 20-i can, for example, be other road users (other here means in particular others in relation to the own / ego system) such as other vehicles, pedestrians, bicycles, motorcycles, etc. The objects 20-i can, in particular, be moving or dynamic objects that can move in the traffic scene (movable) or that can change their properties in the traffic scene over time (dynamic), such as variable obstacles, traffic signs, in particular traffic lights, or the like.

[0037] In addition to the objects 20-i, the sensors 11 can also detect structures 30-i in the vicinity of the vehicle 10. Such structures 30-i can, for example, be predominantly stationary structures, such as road markings, guard rails, peripheral buildings, or the like.

[0038] For example, such structures can characterize possible lanes, road markings, or the like. Furthermore, the sensors 11 may also detect other information, such as the illumination of a brake light, a turn signal of another vehicle, the status of a traffic light, information from traffic signs, or the like.

[0039] The sensors 11 can provide their sensor data to a device 1, in particular a receiving device E for receiving the sensor data. This receiving device E is part of the device 1 for identifying target objects, also referred to below as the device for short. The device 1 can evaluate the data from the sensor system, in particular from its sensors 11, and then generate information for controlling the vehicle 10. In particular, the device 1 can use the data from the sensors 11 to identify objects 20-i in the environment of the vehicle 10 and determine their relevance for further processing for controlling the vehicle 10 or for controlling functions of the vehicle (e.g., output and / or warning functions).

[0040] Fig. 2 shows a schematic representation of a flow chart for a method for identifying target objects, as may be the basis of an embodiment.

[0041] It is understood that, in the description of the embodiments, the individual features can be combined with one another as desired, where appropriate. In particular, embodiments of methods can also include any steps and features that are described in connection with corresponding elements of devices. Conversely, any components suitable for executing the corresponding method steps can also be implemented in the described devices. The functional features of the method-based solution correspond to the respective hardware modules, and vice versa.

[0042] In step S 10, sensor data from at least one sensor 11 is provided. The sensor(s) 11 may, for example, be the previously described sensors 11 for detecting in an environment, in particular in the environment of a vehicle 10. Accordingly, the sensor data may, for example, be image data from one or more cameras, radar data, LiDAR data, or the like.

[0043] In step S20, a situation analysis is then performed using the provided sensor data. For example, in a step S21, objects 20-i and optionally structures 30-i can be identified. For example, the identification of objects 20-i or structures 30-i can be carried out using pattern recognition of image data or by evaluating radar data. Depending on the type of available sensor data and the application, any other methods for identifying objects 20-i or structures 30-i can also be used. For example, the identification of objects 20-i or structures 30-i can in principle also be carried out using a suitable neural network.

[0044] Furthermore, for example, in a step S 22, relations between the identified objects 20-i among themselves and / or between objects 20-i and structures 30-i can be determined. It is also possible to determine relations between the objects 20-i or structures 30-i and the own system, also called the ego system. For example, relative spatial positions (distances) between objects 20-i or between an object 20-i and relevant structures 30-i can be determined. Furthermore, for example, relative speeds between objects 20-i or to structures 30-i can be determined. In addition, any further properties or relations (such as differences and / or similarities with regard to certain attributes, such as vehicle types, engine types, etc.) between objects 20-i among themselves or between objects 20-i and structures 30-i can of course also be determined. In principle, any suitable methods can be used for this purpose.

[0045] Furthermore, in a step S 23, an input tensor x can be formed. Such an input tensor x can be formed, for example, by forming a two-dimensional or multi-dimensional matrix. In this case, for example, one dimension of the input tensor x can correspond to the number of objects 20-i to be evaluated. This number of objects 20-i as well as the corresponding dimension of the input tensor x can be variable. Furthermore, for each object 20-i in the input tensor x, one or more dimensions with object properties can be specified. This number of object properties is generally constant and predetermined, so that the corresponding dimensions of the input tensor x also always remain the same. The object properties can be determined in particular on the basis of the previously determined relations between the objects 20-i among themselves or between objects 20-i and structures 30-i.The object properties can, for example, be numerical values, each of which corresponds to a previously defined property. Object properties can specify, for example, relative velocities, distances, spatial relationships between objects or structures, and any other properties related to an object. Since the dimensions for the object properties are usually constant and fixed, either a zero value or a value indicating an unavailable property can be inserted at the corresponding position in the input tensor x for unavailable object properties.

[0046] The input tensor x can then be processed, for example, in a step S 30 to determine an output tensor y. This output tensor y can specify a probability distribution for the relevance of the identified objects 20-i. For this purpose, a neural network, for example, can be used in this step S 30, as will be explained in more detail below. Alternatively or additionally, another classification method can be used.

[0047] Finally, in a step S40, one or more relevant target objects, in particular a predetermined number of relevant target objects, can be identified. The relevant target objects can be identified, in particular, using the previously determined output tensor y.

[0048] Based on the identified relevant target objects, at least one control function can then be executed in a further step, if necessary. For example, a trajectory for vehicle 10 can be planned based on the identified relevant objects, and a corresponding steering function for vehicle 10 can then be controlled.

[0049] Fig. Figure 3 shows a schematic representation of an exemplary architecture for detecting target objects according to one embodiment. Sensor data, in particular, for example, the previously described data from sensors 11, can be provided as input data 100.

[0050] This input data can be processed, for example, in a first module 200, for example, as part of a situation analysis. In this module 200, referred to as "object embedding," several linear layers 210 can be provided for a linear transformation. In this way, for example, the previously described input tensor x can be formed. Due to dynamic environmental conditions, the number of identified objects and thus also the size of the input tensor x can vary.

[0051] An essential component of the architecture for identifying the target objects is a classification module, in particular the neural network 300, which receives and processes the data from the upstream object embedding module 200. As in Fig. 3, this neural network can be formed from a transformer encoder 300, in particular a transformer encoder with multi-head self-attention.

[0052] A layer 310 of the transformer encoder 300 can, for example, be composed of a multi-head self-attention module 310 with a downstream addition and normalization component (add and normalize) 311, as well as several parallel, particularly fully connected, feed-forward neural networks (FFN) 313 and a further downstream addition and normalization component (add and normalize) 314. The entire transformer encoder 300 can comprise several such layers 310, 320, 330.

[0053] For the final determination of the probability distribution, the output of the transformer encoder 300 can be fed to a downstream prediction head 400. Such a prediction head 400 can, for example, be formed from several parallel fully-connected feedforward neural networks (FFN) 410 and a downstream softmax function 420.

[0054] The output is thus an output tensor y, which contains a probability value for relevance in the current scene of the environment for each object 20-i of the input tensor y. For example, the probability distribution can contain 500 values ​​between 0 and 1, with the sum of all probability values ​​equal to 1, for example.

[0055] Based on this probability distribution 500, relevant objects 20-i in the environment can then be selected. The selection of at least one relevant target object can be carried out, for example, using a prioritization algorithm or a neural network. Downstream processing can, for example, focus on these selected relevant objects 20-i and, if necessary, disregard the remaining objects 20-i. At least, a relevant control function can be limited to the identified relevant objects 20-i.

[0056] Fig. Figure 4 shows a schematic representation of a device 1 according to one embodiment. Such a device 1 can be used, for example, for a fully or at least partially automated system, such as a fully or at least partially automated vehicle, in particular a driver assistance system.

[0057] The device 1 comprises a receiving device E and a plurality of processing devices 3 to 6. The receiving device E can receive sensor data from a sensor system with one or more sensors 11. For this purpose, the receiving device E can be communicatively coupled to the sensors 11 in any desired manner, in particular via a wired communication connection. For example, the sensor data can be transmitted via suitable data bus connections.

[0058] A first processing device 2 of the device 1 can evaluate the sensor data received by the receiving device E to identify objects 20-i and structures 30-i. For this purpose, the approaches already described above can be used, in particular.

[0059] A second processing device 3 of the apparatus 1 can then determine relations between the identified object 20-i and optionally the structures 30-i. A third processing device 4 can then form an input tensor x. A dimension of the input tensor x can vary depending on the number of identified objects 20-i. However, further dimensions of the input tensor x can also be predefined, so that a predetermined number of object properties are specified for each identified object 20-i. These object properties can be formed, in particular, using the previously determined relations between objects 20-i and structures 30-i. If one or more object properties are not present for an object 20-i, a correspondingly predefined value can be used to indicate that this object property is not present.

[0060] The input tensor x can be fed to a fourth processing device 5, which, using this input tensor x, determines an output tensor y by means of a classifier, in particular a neural network. In particular, the previously described approaches for the neural network with a transformer encoder can be implemented for this purpose. The output tensor y with the probability distribution is then fed to a fifth processing device 6, which, based on this probability distribution, identifies one or more target objects and selects a relevant one.

[0061] This identified and selected at least one relevant target object can then be used by downstream components to execute automated functions.

[0062] In summary, the present invention relates to identifying relevant target objects in the environment of an at least partially automated system. For this purpose, it is provided to identify properties corresponding to the objects using a situation analysis. These object properties are fed into a model, which determines a probability distribution for the relevance of the objects using a classifier, in particular a neural network, preferably with a transformer encoder. QUOTES CONTAINED IN THE DESCRIPTION

[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited non-patent literature

[0000] http: / / dx.doi.org / 10.1037 / 0033-295X.101.1.103 R. Kanjee, AK Bachoo, and J. Carroll, „Vision-based Adaptive Cruise Control Using Pattern Matching,“ 2013 6th Robotics and Mechatronics Conference (RobMech), Durban, South Africa, 2013, pp. 10-11. 93–98, doi:10.1109 / RoboMech.2013.6685498

[0004] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin, „Attention Is All You Need“ , arXiv:1706.03762

[0005]

Claims

[1] Method for identifying target objects for an at least partially automated system, comprising the steps: Providing (S 10) sensor data from at least one sensor system for environmental detection of the at least partially automated system; Identifying (S 21) objects (20-i) and optionally structures (30-i) in the environment of the system using the provided sensor data; Determining (S 22) relations between the identified objects (20-i) and optionally between objects (20-i) and structures (30-i); Forming (S 23) an input tensor (x), wherein the input tensor (x) comprises a predetermined number of object properties for each identified object (20-i) of a variable number of objects, wherein the object properties are formed according to the previously determined relations; Determining (S 30) an output tensor (y) using the input tensor (x), wherein the output tensor (y) specifies a probability distribution for a relevance of the identified objects (20-i), wherein the output tensor (y) is determined using a classifier, in particular a neural network; Identifying (S 40) at least one target object that has been specified as relevant using the output tensor (y). [2] Method according to claim 1, wherein the neural network is formed with a transformer encoder and in particular comprises a multi-head self-attention layer. [3] Method according to claim 1 or 2, wherein the output tensor (y) is formed using feed-forward neural networks, in particular simply convolved or fully-connected feed-forward neural networks, which are connected downstream of the transformer encoder. [4] Method according to claim 3, wherein the feed-forward neural networks are followed by a softmax component or a sigmoid component. [5] Method according to one of claims 1 to 4, wherein the determination of relations between the identified objects (20-i) and optionally structures (30-i) takes place in a pre-processing of the provided sensor data by means of an image processing algorithm or in particular by using a neural network which is arranged upstream of the transformer encoder. [6] Method according to one of claims 1 to 5, wherein the input tensor (x) is further formed for a predeterminable number of time steps. [7] Method according to one of the preceding claims 1 to 6, wherein the object (20-i) is a stationary or a moving object (20-i) and / or wherein the identification of objects (20-i) in the environment of the system comprises an identification of further relevant objects (20-i), in particular road users, further vehicles, persons and / or units. [8] Method according to one of the preceding claims 1 to 7, wherein the structure (30-i) is a stationary structure, in particular an infrastructure, and / or wherein the identification of structures (30-i) in the environment of the system comprises identifying lines, in particular roadway boundaries, peripheral buildings and / or elements characteristic of a function of the system. [9] Method according to one of the preceding claims 1 to 8, comprising a step of executing at least one function, in particular a control function of the system and / or an output function, using the probability distribution of the output tensor. [10] A computer program product comprising a computer program comprising instructions which, when the computer program is executed by a computer, cause the computer to carry out the method according to any one of claims 1 to 9. [11] A computer-readable storage medium comprising instructions which, when executed by a computer, cause the computer to carry out the method according to any one of claims 1 to 9. [12] Device (1) for identifying target objects for an at least partially automated system, wherein the device is designed to carry out a method according to one of claims 1 to 10, comprising: a receiving device (E) designed to receive and provide sensor data from at least one sensor system; a first processing device (2) designed to identify objects (20-i) and optionally structures (30-i) in the environment of the system using the provided sensor data; a second processing device (3) designed to determine relations between the identified objects (20-i) and optionally structures (30-i); a third processing device (4) designed to form an input tensor (x), wherein the input tensor (x) comprises a predetermined number of object properties for each identified object (20-i), wherein the object properties are formed according to the previously determined relations; a fourth processing device (5) designed to determine an output tensor (y) using the input tensor (x), wherein the output tensor (y) specifies a probability distribution for a relevance of the identified objects (20-i), and wherein the output tensor (y) is determined using a classifier, in particular a neural network; and a fifth processing device (6) designed to identify at least one target object that has been specified as relevant using the output tensor (y). [13] Driver assistance system, with a device (1) for identifying target objects according to claim 12; a control device which is designed to control at least one function, in particular a driver assistance function, of an at least partially automated system using the at least one identified target object. [14] Driver assistance system according to claim 13, wherein the at least one driver assistance function of the vehicle comprises a lane keeping function, a lane change function, a distance keeping function, a cruise control function and / or a braking function, in particular an emergency braking function. [15] At least partially automated vehicle with a driver assistance system according to claim 13 or 14.

Citation Information

Patent Citations

  • Method and assistance system for detecting objects in the vicinity of a vehicle

    DE102009009211A1

  • Method and processor circuit for operating an automated driving function with object classifier in a motor vehicle, as well as motor vehicle

    DE102021119871A1