Method for carrying out a semantic segmentation of an environment of a technical system

WO2026166786A1PCT designated stage Publication Date: 2026-08-13ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-01-23
Publication Date
2026-08-13

Smart Images

  • Figure EP2026051702_13082026_PF_FP_ABST
    Figure EP2026051702_13082026_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method (100) for carrying out a semantic segmentation of an environment of a technical system (1), comprising the following steps: - providing (101) sensor data, the sensor data resulting from sensing by at least one sensor (2) of the technical system (1), - extracting (102) features from the provided sensor data, - providing (103) environment specifications (11), each of which represents a region in the environment of the technical system (1) in the form of vectors, - updating (104) the environment specifications (11) on the basis of the extracted features, - determining (105) a temporal propagation of the environment specifications (11) taking into account a movement of the technical system (1), - carrying out (106) the semantic segmentation of the environment of the technical system (1) on the basis of the updated (104) environment specifications (11) and / or on the basis of the temporally propagated environment specifications (11).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] R.416430

[0002] - 1 -

[0003] Description

[0004] title

[0005] Methods for performing a semantic segmentation of the environment of a technical system

[0006] The invention relates to a method for performing a semantic segmentation of the environment of a technical system. Furthermore, the invention relates to a computer program, a device, and a storage medium for this purpose.

[0007] State of the art

[0008] Autonomous systems, such as automated vehicles, regularly rely on perception systems to understand their environment and plan their next actions accordingly. Perception involves the detection of other (dynamic) objects and the segmentation of the (static) environment into multiple segments that share a specific characteristic. For many applications, such systems require information about whether certain areas around them are visible or hidden to the perception system, and whether an object is located (occupied) there or whether they are clear for traversal. Typically, this information is displayed as a segmentation mask on a bird's-eye view (BEV) grid around the vehicle in question.

[0009] A deep fusion approach for object detection based on lidar, camera, and radar sensors is presented in “F. Drews, D. Feng, F. Faion, L. Rosenbaum, M. Ulrich, and C. Gläser, “Deepfusion: A robust and modular 3d object detector for lidars, cameras and radars,” in 2022 IEEE / RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 560–567, 2022.” ThisR.416430

[0010] - 2 -

[0011] The model architecture does not yet cover the temporal propagation or tracking of detected objects over time. A multi-sensor fusion framework is described in "X. Chen, T. Zhang, Y. Wang, Y. Wang, and H. Zhao, “FUTR3D: A unified sensor fusion framework for 3D detection,” arXiv preprint arXiv:2203.10642, 2022." While the latter avoids the need for a BEV grid and propagates latent information across time steps, it still assumes synchronized data input from all sensors. This problem can be circumvented, for example, by providing an object detection framework that integrates asynchronous multimodal multi-sensor inputs (see "Y. Miron, F. Drews, F. Faion, DD Castro, and I. Klein, “RCF-TP: Radar-camera fusion with temporal priors for 3D object detection,” IEEE Access, 2024).However, no method for deriving environment segmentation masks, as required for occupancy or visibility estimation, is discussed.

[0012] In recent years, numerous approaches for predicting occupancy in automated driving have been developed in research. Papers such as "Y. Zhang, Z. Zhu, and D. Du, “OccFormer: Dual-path transformer for vision-based 3d semantic occupancy prediction,” arXiv preprint arXiv:2304.05316, 2023.” consider the 3D case of predicting occupied voxels in a 3D space around the vehicle. However, this approach only considers one sensor modality (camera) and not the temporal propagation of features.

[0013] In “Y. Huang, W. Zheng, Y. Zhang, J. Zhou, and J. Lu, “GaussianFormer: Scene as gaussians for vision-based 3d semantic occupancy prediction,” arXiv preprint arXiv:2405.17429, 2024.” a similar concept is presented that is based on learning parameters of 3D Gaussian distributions for occupancy prediction.

[0014] Downstream components such as planning do not currently benefit much from 3D compared to 2D representations, while the computational load increases considerably. Therefore, 3D approaches generally require some form of sparsity handling to keep the computational effort manageable. R.416430

[0015] - 3 -

[0016] Other works, such as “Y. Huang, W. Zheng, Y. Zhang, J. Zhou, and J. Lu, “Tri-per-spective view for vision-based 3d semantic occupancy prediction,” arXiv preprint arXiv:2302.07817, 2023,” consider a so-called tri-perspective view, which allows for the inference of 3D occupancy prediction from three orthogonal planes. This work also only considers camera sensors and not temporal propagation. With native 3D sensors (Lidar, radar), the representation of the extracted features in a BEV raster is straightforward. In this setting, semantic segmentation approaches such as “B. Cheng, I.

[0017] Misra, AG Schwing, A. Kirillov, and R. Girdhar, “Masked-attention mask transformer for universal image segmentation,” CVPR, 2022. This approach can be applied to obtain a set of masks corresponding to different classes. For the current application, the classes could be, for example, occupied and free, or visible and invisible. Using a semantic segmentation approach for this task enables the processing of merged features in BEV space. However, multi-view camera setups require pre-projection of the camera features into BEV space, and all sensors must provide synchronized measurements, which is not feasible in practice.

[0018] Disclosure of the invention

[0019] The invention relates to a method with the features of claim 1, a computer program with the features of claim 8, a device with the features of claim 9, and a computer-readable storage medium with the features of claim 10. Further features and details of the invention will become apparent from the respective dependent claims, the description, and the drawings. Features and details described in connection with the method according to the invention naturally also apply in connection with the computer program, the device, and the computer-readable storage medium according to the invention, and vice versa, so that a reciprocal reference is always possible with regard to the disclosure of the invention.

[0020] The invention relates in particular to a method for performing a semantic segmentation of an environment of a technical system, for example a vehicle or a robot, comprising the following R.416430

[0021] - 4 -

[0022] Steps, whereby the steps can be repeated and / or performed in a specific order:

[0023] Providing sensor data, wherein the sensor data results from the acquisition of at least one sensor of the technical system, for example a camera, radar, lidar, infrared and / or ultrasonic sensor; extracting features from the provided sensor data, for example using a machine learning model;

[0024] Providing environmental specifications, each representing an area in the environment of the technical system in the form of vectors, wherein the environmental specifications within the scope of the present invention can also be referred to and understood as mathematical hypotheses or query features and in particular each comprise a list of vectors,

[0025] Updating the environment specifications, or individual values ​​for the vectors of the environment specifications, based on the extracted features,

[0026] Determining a temporal propagation of the environmental specifications taking into account a movement of the technical system, in other words, in particular predicting the environmental specifications, or (the) individual^) values ​​for the vectors of the environmental specifications, for a future journal,

[0027] Performing semantic segmentation of the technical system's environment based on updated environment specifications and / or on temporally propagated environment specifications, for example using a machine learning model, in particular a segmentation head of the machine learning model.

[0028] As part of extracting features from sensor data, the sensor data can be preprocessed, for example, by noise reduction, normalization, or smoothing, to improve data quality. Subsequently, time-based features (e.g., mean, maximum, variance) and frequency-based features (e.g., via Fourier or wavelet transforms) can be calculated to identify characteristic patterns in the time or frequency domain. Further methods such as signal decomposition, statistical analyses, or R.416430

[0029] - 5 -

[0030] Furthermore, machine learning algorithms can be used to enable the identification of more complex features specific to a given application, such as the detection of anomalies or the extraction of latent structures.

[0031] To obtain initial environmental specifications, mathematical hypotheses, or query features, several methods exist, including random sampling, regular grids, and / or probable query features learned from the training data. For example, a straight road alignment is highly probable and can be assumed as a first hypothesis.

[0032] Determining the temporal propagation of environmental specifications, taking into account the motion of the technical system, can involve the following steps. A dynamic model can be formulated that describes the temporal evolution of the environmental specifications. This can be done using differential equations (e.g., ordinary or partial differential equations) or discrete state transition models (e.g., finite difference schemes). The motion of the technical system can be integrated into the formulated dynamic model as an influencing factor, for example, by modifying the coefficients or as an external disturbance. Subsequently, numerical methods such as Euler's method, Runge-Kutta's method, or other time integration methods can be used to calculate the temporal propagation.

[0033] Semantic segmentation is a technique used in machine vision where each pixel of an image is assigned to a specific class or category to enable detailed analysis of the image content. The primary goal is to divide an entire scene—that is, the environment of the technical system—into meaningful areas, such as sky, road, other vehicles, buildings, or people.

[0034] The method according to the invention advantageously allows for the creation of a precise and dynamic representation of the environment of the technical system. The integration of the environmental specifications enables flexible adaptation. R.416430

[0035] - 6 -

[0036] Considering various scenarios and the movement of the technical system ensures a realistic representation of the environment. Semantic segmentation allows objects in the environment to be classified, which in turn enables applications in areas such as autonomous driving.

[0037] Optionally, each of the environment specifications may include a learned embedding and a position encoding, with temporal propagation being performed using the learned embeddings and position encodings of the environment specifications.

[0038] Within the scope of the present invention, an embedding describes, in particular, a concept for transforming data (e.g., words, objects, or features) into a vector space in order to mathematically represent their relationship or meaning. The learned embedding is thus, in particular, a representation that can be learned by a machine learning model during training. This allows, for example, the semantic meaning of the environment specification or the mathematical hypothesis to be encoded in a multidimensional space. The machine learning model is trained, in particular, on the basis of training data comprising various sensor data. Furthermore, at least one loss function can be minimized during training, which reflects how precisely the environment is mapped using the environment specifications.

[0039] Position encoding is added specifically to provide information about the relative or absolute position of elements within the environment specifications or in the environment of the technical system. This is particularly important for machine learning models such as transformers, as these cannot process inherent sequence information themselves.

[0040] Thus, each environmental specification can represent a semantic context and a position in space. Through temporal propagation, or spread, dynamic changes in the environment can be mapped. R.416430

[0041] - 7 -

[0042] Furthermore, it is advantageous if the invention includes the performance of semantic segmentation:

[0043] Deriving a mask area of ​​each environment specification that corresponds to the area represented in the environment of the technical system in the form of vectors,

[0044] Merging all derived mask areas to obtain a segmentation mask for the semantic segmentation of the technical system's environment.

[0045] One form of derived mask area can, for example, correspond to a lane segment in front of or beside the technical system. In other words, each represented area is delimited by a corresponding mask area. The individual mask areas can then be combined to obtain a comprehensive segmentation mask for performing semantic segmentation. This enables a precise representation of the environment and provides an understanding of spatial relationships between objects within it.

[0046] A further advantage can be achieved within the scope of the invention if, during the extraction process, at least one extrinsic parameter of the at least one sensor is taken into account, in particular a position and / or orientation of the sensor with respect to the technical system. This can improve the accuracy of the semantic segmentation, since the sensor data can be correlated with additional information about the sensor position and / or orientation. This enables, in particular, a more accurate interpretation of the sensor data and can contribute to more precise segmentation.

[0047] Furthermore, it may be possible that performing semantic segmentation includes:

[0048] Creating a bird's eye view (BEV) of the environment of the technical system based on semantic segmentation.

[0049] This allows for a meaningful and clear representation of the environment. This offers a particularly advantageous perspective on R.416430.

[0050] - 8 -

[0051] environment and can be useful for various applications, e.g. for better planning of routes of the technical system.

[0052] Furthermore, it is conceivable that the procedure may also include:

[0053] Determining the occupancy of a space in the vicinity of the technical system based on semantic segmentation.

[0054] This enables, in particular, improved environmental perception and decision-making by the technical system. For example, the technical system could use room occupancy to plan how to move in order to avoid obstacles or find the optimal route.

[0055] Preferably, the invention may provide that the method further comprises:

[0056] Detecting at least one object, particularly a dynamic one, in the vicinity of the technical system, for example, another technical system, a vehicle, or a pedestrian, based on the provided sensor data and / or the extracted features; analyzing the movement of the at least one detected object, whereby the determination of temporal propagation is further carried out taking into account the analyzed movement of the at least one detected object. This achieves the particular advantage of generating an even more accurate and realistic representation of the environment. By considering the movement of objects in the vicinity of the technical system, the method can better capture dynamic scenes and objects and thus generate a more realistic picture of the environment.Integrating object analysis into the temporal propagation process can increase the accuracy of segmentation, as the movements of the detected objects are taken into account.

[0057] It is possible that the method according to the invention is used in a vehicle. The vehicle can be, for example, a motor vehicle and / or passenger vehicle and / or at least partially automated / autonomous vehicle. The vehicle can have vehicle equipment, e.g., for providing an autonomous driving function and / or a driver assistance system.

[0058] - 9 -

[0059] The vehicle equipment may be designed to control, accelerate, brake, and / or steer the vehicle at least partially automatically.

[0060] The invention also relates to a computer program, in particular a computer program product, comprising instructions which, when executed by at least one computer, cause it to execute the method according to the invention. Thus, the computer program according to the invention offers the same advantages as those described in detail with reference to a method according to the invention.

[0061] The invention also relates to a data processing device configured to execute the method according to the invention. The device can, for example, comprise at least one computer which executes the computer program according to the invention. The computer can have at least one processor for executing the computer program. A non-volatile data storage device can also be provided in which the computer program is stored and from which the computer program can be read by the processor for execution.

[0062] The invention may also relate to a computer-readable storage medium which contains the computer program according to the invention and / or includes instructions which, when executed by at least one computer, cause it to execute the method according to the invention. The storage medium is, for example, designed as a data storage device such as a hard drive and / or non-volatile memory and / or a memory card. The storage medium can, for example, be integrated into the computer.

[0063] Furthermore, the method according to the invention can also be implemented as a computer-implemented method. Alternatively or additionally, at least one of the disclosed method steps can be computer-implemented and / or carried out automatically.

[0064] Further advantages, features and details of the invention will become apparent from the following description, with reference to the drawings R.416430.

[0065] - 10 -

[0066] Exemplary embodiments of the invention are described in detail. The features mentioned in the claims and in the description can each be essential to the invention individually or in any combination. The following are shown:

[0067] Fig. 1 shows a schematic visualization of a method, a device, a storage medium and a computer program according to exemplary embodiments of the invention.

[0068] Fig. 2 shows a schematic representation of a technical system with a sensor and an object according to exemplary embodiments of the invention.

[0069] Fig. 3 shows a schematic representation of a method according to exemplary embodiments of the invention.

[0070] Fig. 1 schematically shows a method 100, a device 10, a storage medium 15 and a computer program 20 according to exemplary embodiments of the invention.

[0071] Fig. 1 shows, in particular, an embodiment of a method 100 for performing a semantic segmentation of the environment of a technical system 1. In a first step 101, sensor data are provided, the sensor data resulting from the acquisition of at least one sensor 2 of the technical system 1. In a second step 102, features are extracted from the provided sensor data. In a third step 103, environment specifications 11 are provided, each representing an area in the environment of the technical system 1 in the form of vectors. In a fourth step 104, the environment specifications 11 are updated based on the extracted features. In a fifth step 105, a temporal propagation of the environment specifications 11 is determined, taking into account any movement of the technical system 1.In a sixth step 106, the semantic segmentation of the environment of the technical system 1 is carried out on the basis of the updated environment specifications 11 and / or on the basis of the temporally propagated environment specifications 11.R.416430.

[0072] - 11 -

[0073] Fig. 2 shows a schematic representation of a technical system 1, in particular a vehicle or robot, with a sensor 2 and an object 5, which can be detected, for example, using the sensor 2, according to exemplary embodiments of the invention.

[0074] According to exemplary embodiments, the invention relates to a method for estimating properties of an environment of a technical system 1, such as a (at least partially automated) vehicle or robot, particularly using a machine learning model. In particular, an output module for a machine learning model, preferably a deep learning model, is provided, which is based on an attention mechanism and can thus be combined with a multimodal sensor fusion and object tracking model.

[0075] According to exemplary embodiments of the invention, an output module of a machine learning model, a so-called segmentation head, is provided to derive an environment segmentation from temporally propagated latent information. The latent information is provided, in particular, in the form of environment specifications 11, so-called query features, which can be used to direct the attention of the machine learning model to various sensor inputs. The query features can then be updated. Preferably, in alternating steps, a representation of the query features in the next journal is predicted based on a movement of the technical system 1, in particular the (ego) vehicle, before this representation is updated by the next sensor measurement.The exemplary segmentation masks 4 in the lower row can be considered as an occupancy prediction around the position of the technical system 1. Over time, the technical system 1 moves, in particular, through an environment that, according to the invention, can be represented by the predicted segmentation mask 4. Alternatively, a visible space can be estimated, or semantic segmentation can be performed in which each grid cell is assigned a class (e.g., street, sidewalk, building, or dynamic objects 5 such as other vehicles or pedestrians, etc.). R.416430.

[0076] - 12 -

[0077] The advantages of the invention include, in particular, the following. Specifically, it provides an input of temporally propagated query features that enable asynchronous sensor measurements. It is possible to include all necessary data for deriving the output in these query feature vectors. In particular, no latent bird's-eye view is required, making the approach highly flexible and allowing sensor inputs from multiple perspectives without explicit projection onto the bird's-eye view. The output of the machine learning model preferably provides a mask that can segment the environment of the technical system 1 within a flexible range.

[0078] The detailed description of the invention according to exemplary embodiments is described with reference to Fig. 3. The upper path, in particular, represents a temporal propagation of environmental specifications 11, or query feature vectors, according to the invention. There are several ways to obtain initial environmental specifications 11, or mathematical hypotheses or query features, including random sampling, regular grids, and more complex methods. Fig. 3 shows four environmental specifications 11 by way of example; however, in practice, there can be many more, for example, on the order of several hundred. Each environmental specification 11, or each query feature vector, preferably comprises a learned embedding and a positional encoding, which are required, for example, for temporal propagation.Figure 3 illustrates in particular a case of updating the environmental specifications 11 from journal t-1 to t by incorporating sensor data (e.g., lidar, radar, or camera data or features). Features are extracted from the sensor data, for example, using a feature extraction unit 6 such as a machine learning model. In this way, a feature map 7 can be obtained. It is particularly important to note that the feature map 7 does not have to be a bird's-eye view. In the case of a multi-view camera as sensor 2, the environmental specifications 11 can, for example, consider image feature maps by including the camera extrinsic information, i.e., for example, the position and / or orientation of a (camera) sensor 2, in the feature scanning process. R.416430.

[0079] - 13 -

[0080] In Fig. 3, a mask head, or the execution 106 of semantic segmentation, is attached by way of example to an output of update step 104. However, it can just as easily be performed based on the output of prediction step 105, i.e., the determination of temporal propagation. Within the scope of the execution 106 of semantic segmentation, preferably each environment specification 11, or each query feature vector, is used to derive a region 3 (English: "patch"), which can initially simply be any 2D shape representing part of the environment of the technical system 1. Such a shape can, for example, correspond to a lane segment in front of or beside the technical system 1, or vehicle. By granting the machine learning model freedom with regard to the appearance of these shapes, it can learn basic shapes of typical road layouts, e.g., straight lane segments, curves, intersections, sidewalks, etc.To derive a representation of such forms from a machine learning model, it can be advantageous to first derive a feature map 7 that represents the environment of the technical system 1. Therefore, the environmental specifications 11 can be scattered onto a bird's-eye view feature map near the coded position of the query. Furthermore, there are several alternative ways to derive such a representation from the queries (see below).

[0081] The method up to this point provides, in particular, a bird's-eye view of features that can be used directly to derive the desired segmentation mask 4, which, for example, identifies free and occupied areas. According to exemplary embodiments, the invention also includes the possibility of adding additional refinement and / or post-processing steps to this bird's-eye view. As a simple solution, a fully convolutional machine learning model or neural network can be added, which can refine the segmentation result, for example, by improving the smoothness of a contour. It is also conceivable to add a special semantic segmentation network, such as "B. Cheng, I. Misra, AG Schwing, A. Kirillov, and R. Girdhar, “Masked-attention mask transformer for universal image segmentation,” CVPR, 2022," which can also enable the processing of bird's-eye view feature maps with different resolutions.416430.

[0082] - 14 -

[0083] In the latter configuration, the invention, according to exemplary embodiments, functions in particular as an interface between the purely query-based architecture according to the invention and typical semantic segmentation networks. Several loss functions suitable for semantic segmentation can be used to train the machine learning model according to the invention, for example, the loss functions Cross Entropy or Dice.

[0084] For the next time step, the environmental specifications 11 are preferably predicted for objects. If only the static environment is to be segmented, as is typically the case with occupancy or visibility estimation, the only movement to be considered is, in particular, the movement of the technical system 1. If dynamic objects 5 are also to be considered in the environmental segmentation, it must preferably be ensured that the object movement is also taken into account in the temporal propagation.

[0085] Apart from the exemplary implementation of the invention described above, several variations are also possible. For example, a fusion of multiple sensor inputs can be performed. In these cases, it may be necessary to receive synchronized sensor inputs. However, the approach according to the invention can also process asynchronous sensor inputs.

[0086] According to exemplary embodiments, the invention particularly supports environmental segmentation based on a single sensor modality input (e.g., by omitting the fusion step) or based on another set of sensor modalities to be fused. It is also particularly possible for the environmental segmentation approach to use raw sensor input data, such as radar spectral data, as inputs.

[0087] Furthermore, the derivation of the bird's-eye view feature map from query feature vectors or environmental specifications can be replaced by alternative approaches. For example, parameters of 2D probability distributions (e.g., Gaussian distributions) can be estimated and then assigned to the global mask output.

[0088] - 15 -

[0089] Furthermore, the initial query positions can be defined on a regular grid, and the features can be upsampled to the respective grid cell. The latter option particularly allows for hierarchical refinement, starting with coarse cells that can be iteratively refined. This configuration can make it possible to utilize the potential sparsity in the environment mask, which arises, for example, from large, populated areas where increased resolution is not required.

[0090] The optional semantic segmentation network, used to derive improved environment segmentation masks from the initial bird's-eye view feature map, can use various types of intermediate networks, such as a pixel decoder, for improved information exchange between feature maps with different resolutions.

[0091] In another variant, the attention mechanism that processes the environmental specifications 11 within the segmentation head can be combined with an object recognition head. This reduces the computational load, as the same process performs two tasks.

[0092] This invention can also be applied in other areas, for example, in automated assembly lines to determine the presence and shape of objects to be processed. Furthermore, the invention can be used for autonomously operating household appliances such as vacuum cleaners and lawnmowers, which require an estimation of occupancy and obstacles. In addition, the invention can be used to check automatic doors and gates for obstructing objects. Another application could be the processing of surveillance video footage for improved monitoring of buildings or goods by detecting, for example, hazardous substances. Similar to automotive perception systems, the invention can also be applied to infrastructure-related perception systems and traffic monitoring. Since assistance systems for two-wheeled vehicles (e.g., motorcycles, bicycles, etc.)As our systems become increasingly advanced, they can benefit from this invention, which enables the segmentation of their environment. R.416430.

[0093] - 16 -

[0094] The preceding explanation of the embodiments describes the present invention solely by way of examples. Naturally, individual features of the embodiments can be freely combined with one another, provided this is technically feasible, without departing from the scope of the present invention.

Claims

R.416430 - 17 - Claims 1. Method (100) for performing a semantic segmentation of an environment of a technical system (1), comprising the following steps: Providing (101) sensor data, wherein the sensor data result from a recording of at least one sensor (2) of the technical system (1), Extracting (102) features from the provided sensor data, Providing (103) environmental specifications (11), each representing an area in the environment of the technical system (1) in the form of vectors, Updating (104) the environment specifications (11) based on the extracted features, Determining (105) a temporal propagation of the environmental specifications (11) taking into account a movement of the technical system (1), Performing (106) the semantic segmentation of the environment of the technical system (1) based on the updated environment specifications (11) and / or on the basis of the temporally propagated environment specifications (11).

2. Method (100) according to claim 1 , characterized by that each of the environment specifications (11) includes a learned embedding and a positional encoding, wherein temporal propagation is performed using the learned embeddings and positional encodings of the environment specifications (11). R.416430 - 18 - 3. Method (100) according to one of the preceding claims, characterized in that that performing (106) semantic segmentation includes: Deriving a mask area (3) of each environment specification (11) corresponding to the area represented in vector form in the environment of the technical system (1), combining all derived mask areas (3) to obtain a segmentation mask (4) for the semantic segmentation of the environment of the technical system (1).

4. Method (100) according to one of the preceding claims, characterized in that that during the extraction (102) at least one extrinsic parameter of the at least one sensor (2) is taken into account, in particular a position and / or an orientation of the sensor (2) in relation to the technical system (1).

5. Method (100) according to any one of the preceding claims, characterized in that that performing (106) semantic segmentation includes: Creating a bird's-eye view representation of the environment of the technical system (1) based on semantic segmentation.

6. Method (100) according to one of the preceding claims, characterized in that that the procedure (100) further includes: Determine (105) the occupancy of a space in the vicinity of the technical system (1) based on a result of semantic segmentation. R.416430 - 19 - 7. Method (100) according to one of the preceding claims, characterized in that that the procedure (100) further includes: Detecting at least one object in the vicinity of the technical system (1) based on the provided sensor data and / or the extracted features, Analyzing a movement of the at least one detected object (5), wherein the determination (105) of the temporal propagation is further carried out taking into account the analyzed movement of the at least one detected object (5).

8. Computer program (20), comprising instructions which, when the computer program (20) is executed by at least one computer (10), cause it to execute the method (100) according to one of the preceding claims.

9. Device (10) for data processing, which is configured to carry out the method (100) according to any one of claims 1 to 7.

10. Computer-readable storage medium (15) comprising instructions which, when executed by at least one computer (10), cause it to execute the steps of the method (100) according to any one of claims 1 to 7.