Semantic segmentation of a sequence of images using a trained neural network

By training neuronal networks to predict future semantic segmentation masks using adapted training strategies and attention mechanisms, the challenges of semantic segmentation in autonomous driving are addressed, resulting in improved accuracy and situational handling.

DE102023210964A1Pending Publication Date: 2025-05-08VOLKSWAGEN AG
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
DE102023210964
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-06
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Current technologies face challenges in achieving effective semantic segmentation of image sequences for autonomous driving, particularly in predicting future semantic segmentation masks with high accuracy.

Method used

A procedure and system for training neuronal networks using a sequence of images, where the network predicts future semantic segmentation masks by adapting the training strategy and input/output modalities, and utilizing attention mechanisms and specific loss functions such as cross-entropy or focal loss.

Benefits of technology

The proposed solution enables the prediction of future semantic segmentation masks with improved accuracy and efficiency, enhancing the handling of dangerous situations in autonomous driving applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present invention relates to a method, a computer program with instructions, and a device for training a neural network for the semantic segmentation of a sequence of images, as well as a correspondingly trained neural network. The invention further relates to a method, a computer program with instructions, and a device for the semantic segmentation of a sequence of images using a trained neural network, as well as a means of locomotion with imaging sensors that uses a method or device according to the invention. In a first step, a sequence of images is received (10). This is processed by the neural network (11) to predict future semantic segmentation masks. If no ground truth data for the semantic segmentation of the sequence of images is available, it can be provided by a teacher network (12).A loss is determined between the predicted future semantic segmentation masks and the ground truth data (13). The neural network is then trained using the determined loss (14).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present invention relates to a method, a computer program with instructions, and a device for training a neural network for the semantic segmentation of a sequence of images, as well as to a correspondingly trained neural network. The invention also relates to a method, a computer program with instructions, and a device for the semantic segmentation of a sequence of images using a trained neural network, as well as to a means of transport with an imaging sensor system that uses a method or device according to the invention.

[0002] In the following description, some English terms are used, as German-language technical terms have not been established in the field of machine learning. The English terms used are familiar to those skilled in the art.

[0003] Automated driving, also known as autonomous driving, automated driving, or guided driving, is the movement of vehicles, mobile robots, and driverless transport systems that are largely autonomous. There are different degrees of automated driving. In Europe, various transport ministries, such as the Federal Highway Research Institute in Germany, have defined the following levels of automation: • Level 0: “Driver only”, the driver drives, steers, accelerates, brakes, etc. • Level 1: Certain assistance systems help with vehicle operation, including a cruise control system such as ACC (Automatic Cruise Control). • Level 2: Partial automation. Automatic parking, lane guidance, general longitudinal guidance, acceleration, deceleration, etc., including collision avoidance, are handled by the assistance systems. • Level 3: High automation. The driver does not need to constantly monitor the system. The vehicle performs functions such as activating the turn signal, changing lanes, and maintaining lane guidance independently. The driver can attend to other tasks but must assume control upon request within a certain warning period. • Level 4: Full automation. The system permanently assumes control of the vehicle. If the system is no longer able to handle the tasks, the driver can be asked to take over. • Level 5: No driver required. Aside from setting the destination and starting the system, no human intervention is required.

[0004] A slightly different definition of steps is provided by the Society of Automotive Engineers (SAE). In this context, reference is made to the SAE J3016 standard. Such definitions can be used as an alternative to the definitions given above.

[0005] Machine learning, as part of the broad field of artificial intelligence, has gained increasing importance in recent years. One potential and promising application of artificial intelligence systems is autonomous driving. A key topic for the development of autonomous vehicles is the detailed understanding of the vehicle's surroundings, particularly using image data. One aspect of image data processing is semantic segmentation, i.e., the pixel-by-pixel classification of the image data. In the context of autonomous driving, semantic segmentation enables, for example, the differentiation between different objects such as other vehicles, motorcycles, people, or even different lanes. This differentiation can be of great importance for planning upcoming actions.

[0006] In this context, US 2021 / 0146963 A1 describes a method for generating motion prediction data for a plurality of actors with respect to an autonomous vehicle. The method can generate an input sequence describing object detection data. The input sequence can be input into an attention model. Attention weights can be received as the output of the attention model.

[0007] It is an object of the invention to provide improved solutions for the semantic segmentation of a sequence of images using a trained neural network.

[0008] This object is achieved by a method having the features of claim 1 or 3, by a computer program with instructions according to claim 6, by a neural network according to claim 7, by a device having the features of claim 8 or 9, and by a means of transport according to claim 10. Preferred embodiments of the invention are the subject of the dependent claims.

[0009] According to a first aspect of the invention, a method for training a neural network for the semantic segmentation of a sequence of images comprises the steps: - Receiving a sequence of images; - Processing the sequence of images using the neural network to predict future semantic segmentation masks; - determining a loss between the predicted future semantic segmentation masks and ground truth data of the semantic segmentation of the sequence of images; and - Train the neural network using the determined loss.

[0010] According to a further aspect of the invention, a computer program includes instructions which, when executed by a computer, cause the computer to perform the following steps for training a neural network for the semantic segmentation of a sequence of images: - Receiving a sequence of images; - Processing the sequence of images using the neural network to predict future semantic segmentation masks; - determining a loss between the predicted future semantic segmentation masks and ground truth data of the semantic segmentation of the sequence of images; and - Train the neural network using the determined loss.

[0011] The term "computer" should be understood broadly. It specifically includes workstations, distributed systems, and other processor-based data processing devices.

[0012] The computer program may, for example, be made available for electronic retrieval or stored on a computer-readable storage medium.

[0013] According to a further aspect of the invention, an apparatus for training a neural network for the semantic segmentation of a sequence of images comprises: - an input for receiving a sequence of images; - a processing module for processing the sequence of images using the neural network to predict future semantic segmentation masks; - a loss module for determining a loss between the predicted future semantic segmentation masks and ground truth data of the semantic segmentation of the sequence of images; and - a training module for training the neural network using the determined loss.

[0014] Typically, semantic segmentation is predicted frame by frame. However, in order to more easily predict and deal with certain dangerous situations, it is useful to know in advance what the semantic segmentation might look like a few seconds later. The inventive solution therefore uses neural networks that were created, for example, for the prediction of RGB video images, for the temporal prediction of future semantic segmentation masks. For this purpose, the training strategy as well as the input and output modalities are adapted. Predicting semantic segmentation masks with these networks enables prediction in a more compact and efficient domain. The neural networks can be trained, for example, based on the BDD100K dataset or the Cityscapes dataset. The loss can be a cross-entropy loss (CEN).Cross Entropy Loss) or Focal Loss, which is a modification of Cross Entropy Loss.

[0015] According to one aspect of the invention, the ground truth data for the semantic segmentation of the image sequence is provided by a teacher network. In order to train a neural network for predicting future semantic segmentation masks, ground truth labels for the semantic segmentation must be available for all sequences of images on which the predictor is to be trained. The term "ground truth" is occasionally used in German as "Grundwahrheit" (ground truth). In practice, most large datasets do not provide complete ground truth data for the semantic segmentation of the video sequences. This is therefore advantageously created using a teacher network that has been previously trained on images with available labels.

[0016] According to a further aspect, a neural network is trained for the semantic segmentation of a sequence of images using a method according to the invention. Such a neural network delivers high-quality predicted segmentation masks that can be used, for example, for improved handling of hazardous situations in the field of autonomous driving.

[0017] According to a further aspect of the invention, a method for semantic segmentation of a sequence of images by means of a neural network trained according to a method according to the invention comprises the steps: - Receiving a sequence of images; - Processing the sequence of images using the neural network to predict future semantic segmentation masks; and - Output the predicted future semantic segmentation masks.

[0018] According to a further aspect of the invention, a computer program contains instructions which, when executed by a computer, cause the computer to carry out the following steps for the semantic segmentation of a sequence of images by means of a neural network trained according to a method according to the invention: - Receiving a sequence of images; - Processing the sequence of images using the neural network to predict future semantic segmentation masks; and - Output the predicted future semantic segmentation masks.

[0019] The term "computer" should be understood broadly. In particular, it also includes control units, embedded systems, and other processor-based data processing devices.

[0020] The computer program may, for example, be made available for electronic retrieval or stored on a computer-readable storage medium.

[0021] According to a further aspect of the invention, a device for semantic segmentation of a sequence of images by means of a neural network trained according to a method according to the invention comprises: - an input for receiving a sequence of images; - a processing module for processing the sequence of images using the neural network to predict future semantic segmentation masks; and - an output module for outputting the predicted future semantic segmentation masks.

[0022] Especially for autonomous driving applications, it is advantageous to know in advance what the semantic segmentation might look like a few seconds later. This allows, for example, better handling of dangerous situations.

[0023] According to one aspect of the invention, the neural network is based on a convolutional network or a transformer, or uses motion estimates. The convolutional network can, for example, use a convolutional LSTM cell (LSTM: Long Short-Term Memory) as a central element for modeling the temporal dependencies between the images of the sequence, as described in X. Shi et al.: "Convolutional LSTM network: A machine learning approach for precipitation nowcasting", Advances in Neural Information Processing Systems Vol. 28 (2015). A suitable transformer-based network is, for example, VPTR (VPTR: Video Prediction Transformer), described in X. Ye et al.: "VPTR: Efficient Transformers for Video Prediction", arXiv:2203.15836 (2022). A suitable network that uses motion estimates is MAU (MAU: Motion-Aware Unit), described in Z. Chang et al.: "MAU: A Motion-Aware Unit for Video Prediction and Beyond," Advances in Neural Information Processing Systems Vol. 34 (2021). The two networks create a latent representation of the segmentation masks using an encoder, make a prediction on this latent representation, and project this prediction back to a semantic segmentation using a decoder.

[0024] According to one aspect of the invention, the neural network utilizes an attention mechanism. Attention is a mathematical function that weights input variables according to their relative importance within a set of input variables. It can therefore be interpreted as a mimic of biological attention, as it dynamically emphasizes some input variables and de-emphasizes others. Attention can be described as the mapping of a query and a set of key-value pairs to an output. Query, key, and value are vectors. The use of attention mechanisms has the advantage that the decoder can flexibly utilize the most relevant input variables.

[0025] Advantageously, a means of transport equipped with an imaging sensor system comprises a device according to the invention or is configured to carry out a method according to the invention for the semantic segmentation of a sequence of images from the imaging sensor system. The means of transport can be any type of transport, e.g., a car, a bus, a motorcycle, a commercial vehicle, in particular a truck, an agricultural machine, a construction machine, a rail vehicle, etc. In general, the invention can be used in all vehicles that must handle image data relating to the environment.

[0026] Further features of the present invention will become apparent from the following description and the appended claims taken in conjunction with the figures. Fig. 1 schematically shows a method for training a neural network for the semantic segmentation of a sequence of images; Fig. 2 shows a first embodiment of an apparatus for training a neural network for the semantic segmentation of a sequence of images; Fig. 3 shows a second embodiment of an apparatus for training a neural network for the semantic segmentation of a sequence of images; Fig. 4 schematically shows a method for semantic segmentation of a sequence of images using a trained neural network; Fig. 5 shows a first embodiment of an apparatus for semantic segmentation of a sequence of images using a trained neural network; Fig. 6 shows a second embodiment of an apparatus for semantic segmentation of a sequence of images using a trained neural network; Fig. Figure 7 schematically represents a means of transport in which a solution according to the invention is implemented; Fig. Figure 8 illustrates the general functioning of the temporal prediction of semantic segmentation masks; Fig. Figure 9 illustrates the separate training of semantic segmentation and prediction; Fig. Figure 10 illustrates the joint training of semantic segmentation and prediction; Fig. Figure 11 illustrates a fused training approach; Fig. 12 shows an architecture of a convolutional network; Fig. 13 illustrates an approach based on a video prediction transformer; and Fig. Figure 14 illustrates an approach based on a motion-aware unit.

[0027] To better understand the principles of the present invention, embodiments of the invention are explained in more detail below with reference to the figures. It is understood that the invention is not limited to these embodiments and that the described features may also be combined or modified without departing from the scope of the invention as defined in the appended claims.

[0028] Fig. Figure 1 schematically shows a method for training a neural network for the semantic segmentation of a sequence of images. The neural network can, for example, be based on a convolutional network or a transformer, or use motion estimation. Preferably, the neural network uses an attention mechanism. In a first step, a sequence of images is received 10. This is processed 11 by the neural network to predict future semantic segmentation masks. If no ground-truth data of the semantic segmentation of the sequence of images is available, this can be provided by a teacher network 12. A loss is determined 13 between the predicted future semantic segmentation masks and the ground-truth data. The neural network is then trained 14 using the determined loss.

[0029] Fig. Figure 2 shows a simplified schematic representation of a first embodiment of a device 20 for training a neural network NN for the semantic segmentation of a sequence of images I. The neural network NN can, for example, be based on a convolutional network or a transformer, or use motion estimation. Preferably, the neural network NN uses an attention mechanism. The device 20 has an input 21 for receiving a sequence of images I. A processing module 22 is configured to process the sequence of images I by means of the neural network NN in order to predict future semantic segmentation masks S. The processing module 22 can also be configured to receive ground truth data G of the semantic segmentation of the sequence of images I by means of a teacher network FT provide, if these are not already available. A loss module 23 is configured to determine a loss L between the predicted future semantic segmentation masks S and the ground truth data G. A training module 24 is configured to train the neural network NN using the determined loss L. The trained neural network NN can be made available via an output 27 of the device 20.

[0030] The processing module 22, the loss module 23, and the training module 24 can be controlled by a control module 25. Settings of the processing module 22, the loss module 23, the training module 24, or the control module 25 can be changed via a user interface 28. The data generated in the device 20 can be stored in a memory 26 of the device 20 if necessary, for example, for later evaluation or for use by the components of the device 20. The processing module 22, the loss module 23, the training module 24, and the control module 25 can be implemented as dedicated hardware, for example, as integrated circuits. Of course, they can also be partially or completely combined or implemented as software running on a suitable processor, for example, a CPU or a GPU.The input 21 and the output 27 can be implemented as separate interfaces or as a combined bidirectional interface.

[0031] Fig. 3 shows a simplified schematic representation of a second embodiment of a device 30 for training a neural network for the semantic segmentation of a sequence of images. The device 30 has a processor 32 and a memory 31. For example, the device 30 is a computer, a workstation, or a distributed system. Instructions are stored in the memory 31 which, when executed by the processor 32, cause the device 30 to perform the steps according to one of the described methods. The instructions stored in the memory 31 thus embody a program executable by the processor 32 which implements the method according to the invention. The device 30 has an input 33 for receiving a sequence of images. Data generated by the processor 32 are provided via an output 34. In addition, data can be stored in the memory 31.The input 33 and the output 34 can be combined to form a bidirectional interface.

[0032] The processor 32 may include one or more processor units, such as microprocessors, digital signal processors, or combinations thereof.

[0033] The memories 26, 31 of the described embodiments can have both volatile and non-volatile memory areas and can comprise a wide variety of storage devices and storage media, for example hard disks, optical storage media or semiconductor memories.

[0034] Fig. 4 schematically shows a method for the semantic segmentation of a sequence of images using a neural network trained according to the invention. The neural network can, for example, be based on a convolutional network or a transformer, or can use motion estimation. The neural network preferably uses an attention mechanism. In a first step, a sequence of images is received 10, e.g., from an imaging sensor system of a means of transport. The sequence of images is processed 11 by the neural network to predict future semantic segmentation masks. The predicted future semantic segmentation masks are then output 15 for further use, e.g., for use by a (partially) automated driving function of a means of transport.

[0035] Fig. Figure 5 shows a simplified schematic representation of a first embodiment of a device 40 for the semantic segmentation of a sequence of images I by means of a neural network NN trained according to the invention. The neural network NN can, for example, be based on a convolutional network or a transformer or use motion estimation. Preferably, the neural network NN uses an attention mechanism. The device 40 has an input 41 for receiving a sequence of images I, e.g., from an imaging sensor system 61 of a means of transport. A processing module 42 is configured to process the sequence of images I by means of the neural network NN in order to predict future semantic segmentation masks S. An output module 43 is configured to output the predicted future semantic segmentation masks S via an output 46 of the device 40 for further use, e.g.,for use by an assistance system 62 of a means of transport within the framework of a (partially) automated driving function.

[0036] The processing module 42 and the output module 43 can be controlled by a control module 44. Settings of the processing module 42, the output module 43, or the control module 44 can be changed via a user interface 47. The data generated in the device 40 can be stored in a memory 45 of the device 40 if necessary, for example, for later evaluation or for use by the components of the device 40. The processing module 42, the output module 43, and the control module 44 can be implemented as dedicated hardware, for example, as integrated circuits. Of course, they can also be partially or completely combined or implemented as software running on a suitable processor, for example, a CPU or a GPU. The input 41 and the output 46 can be implemented as separate interfaces or as a combined bidirectional interface.

[0037] Fig. 6 shows a simplified schematic representation of a second embodiment of a device 50 for the semantic segmentation of a sequence of images using a neural network trained according to the invention. The device 50 has a processor 52 and a memory 51. For example, the device 50 is a computer, a control unit, or an embedded system. Instructions are stored in the memory 51 which, when executed by the processor 52, cause the device 50 to perform the steps according to one of the described methods. The instructions stored in the memory 51 thus embody a program executable by the processor 52 which implements the method according to the invention. The device 50 has an input 53 for receiving a sequence of images. Data generated by the processor 52 are provided via an output 54. In addition, data can be stored in the memory 51.The input 53 and the output 54 can be combined to form a bidirectional interface.

[0038] The processor 52 may include one or more processor units, such as microprocessors, digital signal processors, or combinations thereof.

[0039] The memories 45, 51 of the described embodiments can have both volatile and non-volatile memory areas and can comprise a wide variety of storage devices and storage media, for example hard disks, optical storage media or semiconductor memories.

[0040] Fig. 7 schematically illustrates a means of transport 60 in which a solution according to the invention is implemented. In this example, the means of transport 60 is a motor vehicle. The motor vehicle has an imaging sensor system 61, e.g., one or more video cameras, the images of which are used by an assistance system 62 as part of a (partially) automated driving function. For this purpose, a neural network NN, which can be implemented on a computer 63 of the motor vehicle, predicts future semantic segmentation masks based on the images. For example, the assistance system 62 can output warnings to an operator of the motor vehicle via a display device 64. A connection to a backend can be established by means of a data transmission unit 65, e.g., for transmitting settings or retrieving updated software for the components of the motor vehicle.A memory 66 is provided for storing data. Data exchange between the various components of the motor vehicle takes place via a network 67.

[0041] In the following, further details of the invention will be explained with reference to Fig. 8 to Fig. 14 will be explained.

[0042] To predict future semantic segmentation masks from a sequence of RGB images, two tasks must be performed: a prediction of a future state and a transformation from image space X into the semantic segmentation space Y. . Both tasks should be performed by a general processing module AV, which processes a sequence of n RBG images x t-n+1:t as input variables and a future semantic segmentation ŷ t+1 The general functionality of the temporal prediction of semantic segmentation masks is described in Fig. 8 illustrates.

[0043] To train a neural network to predict future semantic segmentation masks, ground-truth semantic segmentation labels must be available for all image sequences on which the predictor is to be trained. In practice, most large datasets do not provide complete ground-truth semantic segmentation data for the video sequences. Therefore, these are advantageously created using a teacher network previously trained on images with available labels.

[0044] As mentioned above, a general processing module is designed to predict the semantic segmentation masks of future images given a sequence of past RGB video images. One approach is to first use a neural network to predict future images, followed by a semantic segmentation network to generate semantic segmentation masks for the predicted images. However, the RGB image space typically contains much more information than the semantic segmentation space, making it more difficult to predict future images. Therefore, it may be useful to predict future semantic segmentations directly in the semantic segmentation space. A simple approach to this is to split the semantic segmentation and prediction into two sub-modules that are trained independently. This approach is described in Fig. 9 shown.

[0045] The semantic segmentation is first performed on unprocessed RGB images x t = (xt,i,c)∈X=IHX×WX×CX trained, whereby X the image space with I=[0,1] designated. HX and WX are the height and width of the input image, CX is the channel dimension of the image space (ie CX=3 for the normal RGB format). The index t∈T={1,2,...,T} denotes the image index of a single video sequence with a total of T images, the index i∈JX={1,2,...,HX⋅WX} corresponds to the spatial position and c∈CX={1,2,...,CX} the channel index. In the corresponding notation, the results of the semantic segmentation are yt=(yt,i,s)∈Y=IHX×WX×S where Y the semantic segmentation space and I=[0,1] The index s∈S={1,2,…,S} denotes the class index for a total of S classes.

[0046] Semantic segmentation is trained on a dataset for which ground truth labels for semantic segmentation are available. Since the segmentation network not only receives the input variables for the predictor G generated during inference, but also the ground truth labels for the prediction network during training, it serves as a teacher network FT with an encoder ET and a decoder DT and receives the notation FT(xt;θT), where θ T are the semantic segmentation parameters. This allows the creation of the semantic segmentation y t and the semantic segmentation mask m t = (m t , i) from an image x t be expressed as follows: yt=FT(xt;θT) mt,i=argmaxs∈S(yt,i,s)

[0047] Once the training of the semantic segmentation is completed, the predictor G(y¯t;θG) trained with the outputs of semantic segmentation. θG denotes the network parameters of the predictor G. As the arguments of the predictor G As can be seen, its input variable is not the immediate output y t semantic segmentation FT, but a probability distribution y t about all S semantic classes. Each element y t,i,s corresponds to the probability that the pixel belongs to class s. Each representation y¯t=(y¯t,i,s)∈Y is also part of the semantic segmentation space. The probability distribution is calculated using the following softmax function: y¯t,i,s=softmaxs∈S(yt,i,s).

[0048] These representations are calculated and stored for all sequences of the dataset and used for training the predictor G used because they tend to contain more information in the form of uncertainty than the discrete semantic segmentation masks m t . This probability distribution is also referred to as the “softmax representation” in the following.

[0049] In practice, the predictor receives G not just a single softmax representation y t , but a sequence of n elements ranging from time step t - n + 1 to time step t. The sequence of softmax representations can be expressed as y t-n+1:t = {y t-n+1 , y t-n+2 , ... , y t}. Based on this sequence, the predictor G directly the probability distribution or softmax representation y^t+1∈Y the future semantic segmentation as: y^t+1=G(y¯t−n+1:t;θG).

[0050] For all images of the sequence, the segmentation masks and softmax representations have already been obtained from the teacher network FT created. The predictor G can therefore be easily trained by comparing a loss between the predicted semantic segmentation ŷ t+1 and the teacher network FT created ground truth label y t+1 is used. In the example of Fig. 9 a cross-entropy loss JCE used.

[0051] During inference, the present approach uses the same teacher network FT which is first trained and used for label generation for the predictor G was used for the transformation of the RGB images from the image space X into the semantic segmentation space Y used. The generated softmax representations y t are stored and a subsequence of n representations is fed into the predictor G which outputs a class-wise probability distribution for the future semantic segmentation. The entire run of the described approach can be expressed as follows: y^t+1=G(FT(xt−n+1:t;θT);θG).

[0052] The semantic segmentation mask for the future time step can be obtained by applying the argmax operator to the predicted softmax output ŷ t+1 can be calculated analogously to equation (2).

[0053] As an alternative to the separate training approach described above, joint training can be implemented. This approach is Fig. 10. The joint training approach aims to directly predict semantic segmentations based on RGB images without splitting the training of semantic segmentation and prediction into two separate steps. Instead, a joint sequential model is developed. It consists of a semantic student segmentation FS(x¯t;θS) with the same architecture as semantic teacher segmentation FT, followed by a predictor G(FS(x¯t;θS);θG) Semantic student segmentation FS can be converted into a student encoder IT and a student decoder DS The quantities θ S and θG are the network parameters of semantic student segmentation FS or the predictor G Due to the lack of direct monitoring of the output of semantic student segmentation FS the output is not necessarily a semantic segmentation, but can be any representation z t-n+1:t The entire construct of semantic student segmentation FS and predictor G is trained in an end-to-end approach.

[0054] For label generation during training, the same semantic teacher segmentation FT used as already introduced above. In each training iteration, the teacher network creates FT the softmax representation y t+1 at time t + 1 from the RGB image x t+1 , while the joint training network tries to achieve a similar softmax representation ŷ t+1 based on a sequence of n previous RGB images x t-n+1:tpredict. The student network FS and the predictor G are shown in the example of Fig. 10 trained together by applying a cross-entropy loss J CE between the prediction ŷ t+1 and the goal y t+1 is determined. A complete single pass for generating a predicted probability distribution over the semantic classes based on a sequence of n RGB images can be expressed as follows: y^t+1=G(FS(xt−n+1:t;θS);θG).

[0055] During inference, no semantic segmentation is performed by the teacher network FT required because the sequential predictor G works directly on RGB images.

[0056] The idea underlying the joint training approach is that the semantic segmentation space Y may not be the ideal space to perform the actual prediction. There is a possibility that semantic student segmentation FS learns a completely different representation than that of the teacher network FT learned semantic segmentation, which may then be better suited for prediction.

[0057] As a further alternative to the separate training approach described above, fused training can be implemented. This approach is Fig. 11. Similar to the joint training approach described above, the fused training aims to directly predict semantic segmentations based on RGB images, whereby all network elements are trained together. The key difference to the joint training approach is that the prediction is now performed in the latent embedding space. Z between the student encoder IT and the student decoder DS The predictor is again G designated.

[0058] Analogous to the joint training approach, the entire student network gives the predicted softmax representation ŷ during a training iteration at time step t t+1 for the time step t + 1 based on a sequence of n previous RGB images x t-n+1:t A complete single run can be described as y^t+1=DS(G(ES(xt−n+1:t;θenc);θG);θdec), where θ enc and θ dec the network parameters of the student encoder IT or the student decoder DS In parallel, the already trained semantic teacher segmentation generates FT from the RGB image x t+1 the goal y t+1 for a loss criterion, in the example of Fig. 11 again the cross entropy loss J CE .

[0059] Just like the joint training approach, the fused model does not require semantic segmentation by the teacher network during inference FT. Instead, the softmax representation of a future semantic segmentation is predicted directly based on the RGB images.

[0060] In the training approaches described above, the cross-entropy loss was considered. The cross-entropy loss is a standard loss function used to train neural networks for classification tasks or semantic segmentation. For the semantic segmentation of a single pixel at position i with i∈J={1,2,...,H⋅W}, {1, 2, ... , H · W}, where H and W are the height and width of a feature map, respectively, the cross entropy loss is defined as JiCE=−∑s∈Sy¯i,slog(y^i,s),

[0061] The index s stands for the class index of the feature map with s∈S={1,2,…,S} where S represents the set of class indices. In practice, the target pixel y i,s in the form of a one-hot encoding, so that y i,s ∈ {0,1} and ∑i∈Jy¯i,s=1. The issue ŷ i,s of the network for pixel i that match the one-hot target y i,s corresponds to a softmax representation and thus the estimated probability that pixel i belongs to class s. The final loss can be easily determined by summing or averaging over all pixels.

[0062] As an alternative to the cross-entropy loss, the focus loss can be used, which is a modification of the cross-entropy loss. For a single pixel, the focus loss can be formulated as JiFL=−∑s∈S(1−y^i,s)γy¯i,slog(y^i,s).

[0063] The hyperparameter y is used to define the newly introduced term (1 - ŷ i,s ), which focuses on difficult-to-learn examples. The other symbols in equation (9) were already introduced for the cross-entropy loss. During training over several epochs, there will be certain classes that are easier for a neural network to learn than others. Classes such as driver, person, motorcycle, and bicycle naturally have similar characteristics. Therefore, if the output of the network for a pixel i and class s is subject to large uncertainty, ŷ i,s is unlikely to be close to 1, which would be an absolute certainty. As a result, the term (1 - ŷ i,s) large and weight the actual cross-entropy contribution for class s by a factor closer to 1. On the other hand, for a class that the network predicts with high probability, the predicted probability ŷ i,s for pixel i and class s are closer to 1. The consequence is that the term (1 - ŷ i,s ) becomes small and thus the cross-entropy contribution for this specific pixel and class is reduced, since the network is already sure of its choice.

[0064] Furthermore, equation (9) can be modified by adding an additional class-specific weighting factor α s This weighting factor is used to manually assign special weight to underrepresented classes: JiFL=−∑s∈Sαs(1−y^i,s)γy¯i,slog(y^i,s).

[0065] Fig. Figure 12 shows the architecture of a convolutional network. The convolutional network consists of convolutional layers, transposed convolutional layers, and ConvLSTM layers to exploit temporal dependencies between correlated video frames or semantic segmentations.

[0066] A ConvLSTM layer is essentially a fully connected LSTM, where the standard matrix-vector multiplications in the input-hidden and hidden-hidden mappings are replaced by convolutions, while the internal cell structure remains the same. The equations for calculating the input gate I t , the Forget Gate F t and the output gate O t in time step t are given by It=σ(WI,ih∗xt+WI,hh∗Ht−1+BI) Ft=σ(WF,ih∗xt+WF,hh∗Ht−1+BF) Ot=σ(WO,ih∗xt+WO,hh∗Ht−1+BO) and the update rules for cell state C tand the hidden state H t in time step t are given by Ct=Ft⊙Ct−1+It⊙tanh(WC,ih∗xt+WC,hh∗Ht−1+BC) Ht=Ot⊙tanh(Ct).

[0067] The operators ⊙ and * denote the element-wise multiplication and the convolution operator, respectively, the function σ(·) denotes the element-wise applied sigmoid function and the tensor x t denotes the input to the ConvLSTM at time step t instead of an RGB image. The entities W Z,ih and W Z,hh where Z ∈ {I, F, O, C} are the kernel weights for the input-hidden and hidden-hidden mappings. The bias values ​​for the gates are given by B Z symbolized, again with Z ∈ {I, F, O, C}.

[0068] The entire convolutional network can be divided into an encoder and a decoder. The encoder extracts spatiotemporal features from the input data. It consists of a total of five blocks, each consisting of a standard convolutional layer, followed by a leaky ReLU activation function with a slope of -0.2, and a ConvLSTM cell. A group normalization layer is applied to the input gate, the forgetting gate, the cell gate, and the output gate.

[0069] The first block receives a softmax representation y¯t=(y¯tt,i,s)∈Y, where Y again the semantic segmentation space is with the height H x , the width W xand the total number S of classes. The convolution consists of a (3×3) filter with an output dimension of 16 and a stride of 1. The ConvLSTM consists of 64 (5×5) filters with a stride of 1. This first block reduces the spatial sizes H x and W x the input is not.

[0070] The other four blocks each consist of convolutions with (3×3) filters with a stride of 2 and {64, 96, 128, 256} output feature maps. Each layer reduces the spatial dimensions by a factor of two. The ConvLSTM cells of the blocks each use a (5×5) filter with a stride of 1 and also {64, 96, 128, 256} output feature maps. The resulting output dimensions of the five encoder blocks are Hx16×Wx16×256.

[0071] The decoder is constructed inversely to the encoder and also consists of a total of four upsampling blocks, each consisting of a ConvLSTM cell followed by a transposed convolution with a leaky ReLU activation function with a slope of -0.2. A fifth block serves as the output block.

[0072] The ConvLSTM cells of the first four blocks consist of (5×5) filters with a stride of 1 and {256, 128, 96, 96} output feature maps. The transposed convolutions use (4×4) filters with a stride of 2 and {128, 96, 96, 64} output feature maps. Each convolution doubles the spatial dimensions of a feature map. Therefore, after the fourth block, the input height H is again x and the entrance width W xThe fifth block consists of two further transposed convolutions, a first with a (3×3) filter and 32 output feature maps and a second with a simple (1×1) filter and an output dimension of S channels to convert the output back into the original semantic segmentation space Y Both final transposed convolutions have a stride of 1. According to the previously described separate architecture, the output of the network is the predicted softmax representation ŷ t+1 the semantic segmentation with the corresponding segmentation mask m̂ t+1 .

[0073] Additionally, residual connections between corresponding convolutions and transposed convolutions are used in the form of a U-net. This helps overcome the vanishing gradient problem and facilitates gradient flow through the network during training.

[0074] Fig. Figure 13 illustrates an approach based on a video prediction transformer (VPTR). The transformer-based VPTR model was originally developed for predicting RGB videos. In this application, the architecture is adopted for predicting semantic segmentations. The general architecture consists of three subnetworks. The first is a convolutional encoder. EVTPR(y¯t;θEVPTR), which converts the input variable into a latent intermediate representation z¯t∈Z=ℝ+HZ×WZ×CZ transformed. Z denotes the latent space, where ℝ + all positive real numbers including 0, HZ,WZ and CZ are the height, width and channel dimension of the latent space. θεVPTR denotes the encoder parameters. This is followed by a VidHR former (High-Resolution Transformer for video processing) GVPTR(z¯t−n+1:t;θGVPTR), which connects the actual predictor in the VPTR architecture with the network parameters θGVPTR The last subnetwork is a convolutional decoder GVPTR(z^t;θDVPTR) with parameters θDVPTR, which has a latent representation ẑ t projected back into the original semantic segmentation space y.

[0075] The detailed architecture of the networks and the reason why the VidHR former GVPTR directly the entire sequence z t-n+1:t as input variable are explained below.

[0076] The VidHR Former GVPTR is the network responsible for predicting a future representation given a sequence of reference representations. It uses spatial and temporal attention separately. In contrast to the convolutional network of Fig. 12, it doesn't iterate over the input sequence element by element, but rather receives and processes the entire input sequence at once. This feature is common for transformer-based architectures.

[0077] The VidHR former is essentially a stack of several identical VidHR former blocks. The first VidHR former block receives the latent input sequence z t-n+1:t of length n. In practice, the resulting input tensor has the form HZ×WZ×CZ×n. It first undergoes layer normalization, followed by a local spatial MHSA (Multi-Head Self-Attention). For this, each spatio-temporal feature map is divided into P patches, each with a height and width of K pixels. A single patch of a representation z t can be used as p ρ,t be written, where ρ∈P={1,2,...,P} and p ρ,t ∈ pρ,t∈ℝK2×CZ. The result is a redesigned tensor of the form K2×P×CZ×n. The local spatial MHSA is calculated on each patch, with a single head being calculated as headi(pρ,t)=softmax((pρ,tQWiQ)(pρ,tKWiK)CZ / I)pρ,tWiV,

[0078] Where WiQ,WiK and WiV are the weight matrices for the query, the key and the value and i denotes the index of the head with i∈J={1,2,…,I}. The entities pρ,tQ and pρ,tK denote p ρ,t after applying position encoding. The result of the local spatial MHSA for a single patch p ρ is given by the concatenation of the heads Multiheadi(pρ,t)=concat(head1(pρ,t),…,headI(pρ,t))WO, where W Orepresents the weight matrix of a final linear output projection of the local spatial MHSA. A fixed 2D position encoding consisting of channel-dependent sine and cosine functions is used.

[0079] After the local spatial MHSA, the output is returned to HZ×WZ×CZ×n formed and passed through another normalization layer, followed by a stack of forward layers. This forward convolutional neural network (FFN) consists of a 3×3 depthwise convolution between two pointwise multilayer perceptrons (MLPs) and a layer normalization. The reason for the 3×3 convolution is that the local spatial MHSA alone only exchanges information within a single patch, but not across the entire feature map.

[0080] Up to this point, each layer has only processed spatial features, since the local spatial MHSA and the Conv. FFN layers only operate on a spatial feature map. To utilize the information along the temporal axis, a temporal MHSA follows the Conv. FFN. For this purpose, the dimension of an input feature map is HZ×WZ×CZ×n to HZWZ×CZ×n The temporal MHSA is then a standard MHSA, in which the sequence is viewed along the temporal axis rather than the patch axis. A simple fixed 1D positional encoding of time is used. Unlike the local spatial MHSA, no division into patches is required.

[0081] The last part of the VidHR former block is a two-layer MLP whose output is remapped to the original input dimension HX×HX×S×n the latent input sequence z t-n+1:tis converted back. Residual connections are used for the entire block. Since the output dimension is identical to the input dimension, the output ẑ t-n+2:t+1 again a sequence of n representations. In contrast to the input, the output is shifted into the future by a single time step, i.e., the first entry of the output corresponds to time step t - n + 2, while the first entry of the input corresponds to time step t - n + 1.

[0082] The VidHR former calculates a sequence of temporal predictions ẑ t-n+2:t+1 based on a sequence of latent representations z t-n+1:t , which is generated by the encoder EVPTR This creates softmax representations y t-n+1:t from the semantic segmentation space Y on a latent representation, which here is also represented by z t-n+1:t and is part of the latent space Z Accordingly, the decoder DVPTR for the transformation of a sequence from the latent space Z back to the semantic segmentation space Y The combination of EVPTR and DVPTR forms the VPTR autoencoder.

[0083] Both the encoder and decoder are ResNet-based models adapted from the Pix2Pix model. The encoder consists of several convolutional layers, each combined with batch normalization and ReLU activation functions. The first convolutional layer has a kernel size of (7×7), a stride of 1, and generates 64 output feature maps. This layer is followed by three downsampling blocks consisting of convolutions with (3×3) filters, stride 2, and {128, 256, 528} output feature maps. The encoder is completed with a stack of five ResNet blocks, each consisting of two convolutional layers with a (3×3) filter, a stride of 1, and a constant 528 output feature maps. The two convolutions of a block are followed by a batch normalization layer and a ReLU activation function, just like the previous convolutions of the encoder.The key difference from the previous convolutions is that residual connections are used between the five ResNet blocks to facilitate gradient flow and ensure the trainability of the network.

[0084] For the entire encoder, the spatial dimensions HX and WX the input size is reduced by a factor of 8, while the channel dimension S of the semantic segmentation space is extended to 528. If the temporal dimension is omitted for simplicity, this latent representation has the dimensionality Hx8×Wx8×528. 528.

[0085] The decoder, in turn, is the transposed counterpart of the encoder. It consists of three upsampling layers. Each upsampling layer consists of a transposed convolution with {528, 256, 128} filters with a (3×3) kernel and a stride of 2. These transposed convolutions restore the original spatial dimensions of the input. Additionally, as in the encoder, batch normalization and ReLU activation functions are used. However, unlike the encoder, no ResNet blocks are added after the upsampling layers. Instead, a final convolutional layer with a kernel size of (7×7) maps the 128 feature maps back to the channel dimension S of the input, which corresponds to the number of classes in the semantic segmentation.

[0086] Fig. Figure 14 illustrates an approach based on a Motion-Aware Unit (MAU). This approach shows similarities to both the convolutional approach Fig. 12 as well as the transformer-based approach from Fig. 13. Just like the convolutional network, the network operates as a time series network. Therefore, at a time t, it only receives the softmax representation y t as input and predicts the softmax representation ŷ t+1 depending on the input variable and some internal network states, which will be explained later. Since it does not process an entire sequence at once by weighting the sequence elements against each other, it is not a transformer-based network.

[0087] Similar to the VPTR approach, the MAU-based network performs the actual prediction on a latent representation created by a convolutional encoder. The latent prediction is then transformed back into the semantic segmentation space by a convolutional decoder.

[0088] As already mentioned, the input to the network is a single softmax representation y t for the time step t from the same semantic segmentation space Y, described above. It is first processed by the convolutional encoder EMAU(y¯t;θEMAU) led, whereby θEMAU the parameters of the encoder. The output of the encoder is a latent intermediate representation z¯t∈Z=ℝHZ×WZ×CZ the semantic softmax representation, where Z denotes the latent space and HZ,WZ and CZ are the height, width and channel dimension of the latent space. In contrast to the latent space Z defined for VPTR, which consists of all positive real numbers ℝ + including 0, the latent space Z for this architecture over all real numbers.

[0089] The encoder itself consists of a simple stack of convolutional activation blocks. The first convolution uses a simple (1×1) filter with a stride of 1 and projects the input onto 64 output feature maps, where the original input height HX and width WX the semantic segmentation space Y This first convolution is followed by two further convolutional layers. They receive and output a constant {64, 64} feature maps, but reduce the original spatial dimensions of the input variables by a factor of 2, resulting in a latent representation z t The resulting dimensions of the latent space Z are therefore HZ=HX4,WZ=Wx4 and CZ=64. Both layers use a (3×3) filter with a stride of 2. The activation function following each convolutional layer is a leaky ReLU with a slope of 0.2.

[0090] Typically, it is not a single MAU that is responsible for the prediction on the latent space, but a serial stack of MAUs whose output is the predicted latent representation z t+1 in time step t + 1.

[0091] The stack of MAUs is used here as a predictor GMAU(z¯t,z¯t−τ:t−1,Tt−τ:t−1;θGMAU) where θGMAU the parameters of the predictor, e.g. t , e.g. t-τ:t-1 the past spatial states within a fixed receptive field τ and T t-τ:t-1the past temporal states within the same receptive field. The receptive field is effectively a hyperparameter that determines how many past temporal and spatial states are considered for estimating the new temporal state. It differs from the length n of the input sequence z t-n+1:t .

[0092] Comparable to the cell state of an LSTM cell, a temporal state T t uniquely assigned to a single MAU. Since the predictor consists of a stack of MAUs, each MAU has its own temporal state. In general, the temporal state of a k-th MAU is k∈K={1,2,…,K} with Ttk and the spatial initial state of a k-th MAU, which is passed on to the next unit, is denoted by z¯tk.

[0093] Since the output of the encoder EMAU corresponds to the input size of the MAU, we write z¯t=z¯t0. z t The spatial initial state z¯tK the final K-th MAU and thus the final output of the predictor corresponds to the latent representation ẑ t+1 in the next time step t + 1. Accordingly, for the output of this final K-th MAU and the input of the decoder DMAU the layer index can also be omitted, so that z¯tK=z^t+1K=z^t+1.

[0094] A convolutional decoder DMAU(z^t;θDMAU) is for the projection of the latent prediction back into the semantic segmentation space Y Just like the convolutional network, the output size of the network for a single run is defined as ŷ t+1 with the corresponding segmentation mask m̂ t+1 .

[0095] Residual connections are used between each MAU of the predictor. In addition, an information recall scheme is used, which is Fig. 14. The information recall scheme is essentially a U-net-inspired connection between the corresponding encoder and decoder layers. The input feature tensor f ℓ-1 for the ℓ-th decoder layer DlMAU is given by fl−1=dl−1+e−l, where d ℓ-1 the output of the ℓ - 1-th decoder layer and e -ℓ is the output of the ℓ-th encoder layer, counted from the last layer of the encoder.

[0096] The decoder itself consists of two transposed convolutional layers, followed by a leaky ReLU activation function with a slope of 0.2 to maintain the original spatial resolution. Both layers use a (3×3) kernel with a stride of 2 and a constant {64, 64} output feature maps. The third and final layer is a normal convolution with a (1×1) kernel and a stride of 1 to map the 64 latent feature maps back into the semantic segmentation space. Y to transform.

[0097] A MAU has the task of estimating motion information and combining this motion information with the current spatial information z¯tk−1 to make a prediction. To do so, it uses an updatable cell-internal temporal state Ttk, which is unique to each MAU. A single unit consists of two sub-modules: an attention module and a fusion module. The attention module is responsible for estimating motion information using attention mechanisms, while the fusion module is responsible for correctly combining the motion information with the MAU inputs. The following explanation of a single MAU is specific to a kth MAU; therefore, all spatial and temporal states are given the index k or k - 1 for the spatial input of the kth cell.

[0098] Input variables for the attention module are past temporal states Tt−τ:t−1k in the receptive field τ, together with past spatial input variables z¯t−τ:t−1k−1 and the current spatial input variable z¯tk−1 at time step t. To estimate motion information from past temporal states, all temporal states in the receptive field should be weighted according to their importance. Attention is a suitable concept for this task.

[0099] To estimate the correlations between the temporal states v, the correlation between the current spatial state z¯tk−1 and the past spatial conditions z¯t−τ:t−1k−1 considered. An attention map Ak={a1k,…,aτk}, which consists of attention weights a j for each element in the respective array with j ∈ {1, 2, ... , τ}, is calculated as the normalized scalar product between each of the past spatial states and a projected current spatial state, followed by a softmax function. The attention weight aτ(z¯tk−1,z¯t−τk−1) for the last stored temporal state in time step t - τ is given, for example, by aτ(z¯tk−1,z¯t−τk−1)=softmax(FZAtt(z¯tk−1)⋅z¯t−τk−1CZ⋅HZ⋅WZ), where CZ the channel dimension and HZ and WZ the height or width of the latent space Z The function FZAtt is a convolution with a (5×5) filter, a stride of 1, and a constant of 64 input and output feature maps. The attention weights for other time steps are calculated accordingly with the corresponding element of z. t-τ:t-1 calculated.

[0100] The attention weights for each time step in the receptive field are used to assign each element of the past temporal states Tt−τ:t−1k to be weighted individually. After the states have been weighted, they are summed to yield the long-term movement information TLong−termk TLong−termk=∑j=1τajk⋅Tt−jk

[0101] To emphasize the influence of short-term movement information even more, the last time state Tt−1k via a fusion gate Ufk included once more. The fusion gate is calculated as Ufk=sigmoid(FTAtt(Tt−1k)), where FTAtt is also a convolution with a (5×5) filter, a stride of 1, and 64 output dimensions. The fusion gate U f and long-term movement information TLong−termk are combined according to the following formula TAMIk=Ufk⊙Tt−1k+(1−Ufk)⊙TLong−termk, where ⊙ denotes element-wise multiplication and TAMIk is the resulting aggregated motion information of the k-th MAU.

[0102] After the aggregated motion information has been created by the attention module, it is forwarded to the second sub-module of the MAU, the fusion module. In addition to the aggregated motion information TAMIk it only receives the current spatial state z¯tk−1 as input. Both the current spatial state and the aggregated motion information are passed through a convolutional layer with a (5×5) kernel, a stride of 1, and {192, 192} output feature maps. The channel dimension of the projected aggregated motion information and the spatial features is therefore expanded by a factor of 3. The resulting feature tensor for the aggregated motion information is split into three smaller feature tensors along the channel dimension, resulting in the features TAMIk,g, TAMIk,t and TAMIk,s , each of which has the original channel dimension of 64. The same procedure is applied to the projected spatial state, so that the three feature tensors S k,g , S k,t and S k,s from the projected current spatial state z¯tk−1 can be extracted.

[0103] As indicated by the upper index g, the feature tensors TAMIk,g and S k,g for calculating the fusion gates Uftk=sigmoid(TAMIk,g) Ufsk=sigmoid(Sk,g) used so that Uftk and UFSK are the resulting update gates for the new temporal and spatial states, respectively. With these update gates and the other feature tensors TAMIk,t,TAMIk,s S k,t and S k,s the updated spatial state Ttk and the updated temporal state Ttk=Uftk⊙TAMIk,t+(1−Uftk)⊙Sk,t be calculated as follows: z¯tk=Ufsk⊙Sk,s+(1+Ufsk)⊙TAMIk,s. Ttk

[0104] The upper indices t and s indicate whether the tensor is used to calculate the updated temporal state Ttk or the updated spatial state z¯tk is used. While the updated temporal state Ttk remains specific to a single MAU, the new spatial state z¯tk passed to the next MAU. The output of the last MAU in a stack of MAUs can then be used as the latent prediction ẑ t+1 be interpreted. List of reference symbols 10 Receiving a sequence of images 11 Processing the sequence of images using a neural network 12 Providing ground truth data 13 Determining a loss 14 Training the neural network using the loss 15 Outputting predicted future semantic segmentation masks 20 Device 21 Entrance 22 Processing module 23 Loss modulus 24 training modules 25 Control module 26 storage 27 Exit 28 User interface 30 Device 31 storage 32 processor 33 Entrance 34 Exit 40 Device 41 Entrance 42 Processing module 43 Output module 44 Control module 45 storage 46 Exit 47 User interface 50 device 51 storage 52 processor 53 Entrance 54 Exit 60 means of transport 61 Imaging Sensors 62 Assistance system 63 computers 64 Display device 65 Data transmission unit 66 storage 67 Network AV General Processing Module G Ground truth data I Image L loss NN Neural Network S Segmentation mask Student decoder decoder Student encoder Encoder Student network Teacher network Predictor QUOTES CONTAINED IN THE DESCRIPTION

[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited patent literature

[0000] US 2021 / 0146963 A1

[0006] Zitierte Nicht-Patentliteratur

[0000] X. Shi et al.: „Convolutional LSTM network: A machine learning approach for precipitation nowcasting“, Advances in Neural Information Processing Systems Vol. 28 (2015

[0023] X. Ye et al.: „VPTR: Efficient Transformers for Video Prediction“, arXiv:2203.15836 (2022

[0023] Z. Chang et al.: „MAU: A Motion-Aware Unit for Video Prediction and Beyond“, Advances in Neural Information Processing Systems Vol. 34 (2021

[0023]

Claims

[1] Method for training a neural network (NN) for the semantic segmentation of a sequence of images (I), comprising the steps: - receiving (10) a sequence of images (I); - processing (11) the sequence of images (I) by means of the neural network (NN) to predict future semantic segmentation masks (S); - determining (13) a loss (L) between the predicted future semantic segmentation masks (S) and ground truth data (G) of the semantic segmentation of the sequence of images (I); and - Training (14) the neural network (NN) using the determined loss (L). [2] Method according to claim 1, wherein the ground truth data (G) of the semantic segmentation of the sequence of images (I) from a teacher network (FT) provided (12). [3] Method for the semantic segmentation of a sequence of images (I) by means of a neural network (NN) trained according to a method according to claim 1 or 2, comprising the steps: - receiving (10) a sequence of images (I); - processing (11) the sequence of images (I) by means of the neural network (NN) to predict future semantic segmentation masks (S); and - Outputting (15) the predicted future semantic segmentation masks (S). [4] Method according to one of the preceding claims, wherein the neural network (NN) is based on a convolutional network or a transformer or uses motion estimates. [5] Method according to one of the preceding claims, wherein the neural network (NN) uses an attention mechanism. [6] A computer program comprising instructions which, when executed by a computer, cause the computer to carry out the steps of a method according to any one of claims 1 to 5. [7] Neural network (NN) for the semantic segmentation of a sequence of images (I), wherein the neural network (NN) was trained by means of a method according to claim 1 or 2. [8] Device (20) for training a neural network (NN) for the semantic segmentation of a sequence of images (I), comprising: - an input (21) for receiving (10) a sequence of images (I); - a processing module (22) for processing (11) the sequence of images (I) by means of the neural network (NN) to predict future semantic segmentation masks (S); - a loss module (23) for determining (13) a loss (L) between the predicted future semantic segmentation masks (S) and ground truth data (G) of the semantic segmentation of the sequence of images (I); and - a training module (24) for training (14) the neural network (NN) using the determined loss (L). [9] Device (40) for the semantic segmentation of a sequence of images (I) by means of a neural network (NN) trained according to a method according to claim 1 or 2, comprising: - an input (41) for receiving (10) a sequence of images (I); - a processing module (42) for processing (11) the sequence of images (I) by means of the neural network (NN) to predict future semantic segmentation masks (S); and - an output module (43) for outputting (15) the predicted future semantic segmentation masks (S). [10] Means of transport (60) with an imaging sensor system (61), wherein the means of transport (60) comprises a device (40) according to claim 9 or is configured to carry out a method according to one of claims 3 to 5 for the semantic segmentation of a sequence of images (I) of the imaging sensor system (61).

Citation Information

Patent Citations

  • Systems and Methods for Generating Motion Forecast Data for a Plurality of Actors with Respect to an Autonomous Vehicle

    US20210146963A1