Method for training a machine learning model for semantic scene understanding
By leveraging self-supervised learning with data from multiple cameras capturing the same scene, the method addresses the challenge of requiring extensive human annotations for semantic scene understanding, achieving effective semantic scene information determination and improved context understanding.
Patent Information
- Application Number
- DE102023212400
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-08
- Publication Date
- 2025-06-12
AI Technical Summary
Current methods for training machine learning models for semantic scene understanding are hindered by the need for extensive human annotations, which are time and cost intensive, and often rely on limited labeled data.
The method employs a self-supervised learning approach using data from multiple cameras capturing the same scene from different views, allowing the model to learn semantic information about scenes and environments without explicit human annotation.
This approach enables the training of machine learning models that can effectively determine semantic scene information, improving context understanding and reducing the need for manual annotation, while utilizing unlabeled data for pretraining backbone models.
Smart Images

Figure 00000000_0001_ABST 
Figure 00000000_0000_ABST
Abstract
Description
[0001] The invention relates to a method for training a machine learning model for semantic scene understanding. Furthermore, the invention relates to a machine learning model, a computer program, a device, and a storage medium for this purpose. State of the art
[0002] Deep neural networks (DNNs) have a strong ability for data abstraction and outperform classical methods in many areas of machine learning, such as computer vision, object detection, and speech recognition. In computer vision, the best results have been achieved by training large models with many parameters using large amounts of training data. However, providing the human annotations required for supervised learning is time-consuming and costly. This means that typically only a small portion of the available data is labeled. Therefore, recent work has focused on using unlabeled data for pretraining backbone models.
[0003] One way to harness unlabeled data is self-supervised learning, which aims to build background knowledge and a sense of common sense in AI models before fine-tuning the model for a specific downstream task. In self-supervised learning, the model obtains supervisory signals from the data itself, often to predict the same feature space for different views (or augmentations) of the same image. To generate identical embeddings of an object regardless of the view, the model must learn semantic information about that specific object (object understanding). Well-known, conventional methods are described in SimSiam [1] and Dino [2] (references are provided at the end of the description for easy reference). Disclosure of the invention
[0004] The subject matter of the invention is a method having the features of claim 1, a machine learning model having the features of claim 8, a computer program having the features of claim 9, a device having the features of claim 10, and a computer-readable storage medium having the features of claim 11. Further features and details of the invention emerge from the respective subclaims, the description, and the drawings. Features and details described in connection with the method according to the invention naturally also apply in connection with the machine learning model according to the invention, the computer program according to the invention, the device according to the invention, and the computer-readable storage medium according to the invention, and vice versa, so that with regard to the disclosure of the individual aspects of the invention, reference is or can always be made to each other.
[0005] The subject matter of the invention is, in particular, a method for training a machine learning model for semantic scene understanding. Semantic scene understanding can refer to the determination of semantic scene information, i.e., preferably a classification of image information with respect to information about the scene represented therein. In other words, in particular, an understanding of the scene is trained based on the image information.
[0006] The method may initially comprise providing training data. The training data may comprise image information representing a particular scene in an environment. The scene is, for example, a traffic scene in the environment of a vehicle or another scene in the environment of a device. It is conceivable that the scene changes over time and / or that the environment also changes, e.g. while the vehicle / device is moving. Therefore, it is possible for the training data, in addition to the image information for representing one scene, to also comprise image information for further scenes that represent the further scenes of the one or more environments. In this case, it may be crucial that several pieces of image information, in particular several images, are provided for each scene, which possibly show the (same) environment from different perspectives and / or with different views.The images can result from different image sensor sources, which in particular capture the (same) environment with different views / perspectives.
[0007] It is further possible with the method for the image information to result from different image sensor sources in order to represent the respective scene with different views of the environment in the image information. This means in particular that the image information originates from different image sensor sources, such as real or simulated cameras, which represent the (same) environment, but with a different view and / or camera configuration and / or perspective and / or orientation and / or different field of view. The image sensor sources can include, for example, cameras, vehicle cameras, CMOS sensors, CCD sensors, infrared sensors, ultrasonic sensors, radar sensors, LiDAR sensors, thermography sensors, and other types of optical sensors.
[0008] The method can further comprise training the machine learning model based on the provided training data to determine semantic scene information. For this purpose, the environment can be evaluated based on the image information according to the different views in order to determine and, in particular, identify the semantic scene information. In other words, the representation can be evaluated according to the different views, i.e., in particular, multiple representations for representing the scene with different views.
[0009] Furthermore, the method may comprise providing the trained machine learning model in order to use it, for example, for determining the semantic scene information in a vehicle or other device.
[0010] One idea underlying the invention is the way, and in particular the modality, in which the data from a vehicle with multiple cameras or a device with multiple cameras in general is fed into a machine learning model such as a Siamese network for self-supervised learning. In contrast to conventional solutions that use self-supervised learning methods for computer vision (such as SimSiam [1] or Dino [2]), in which different views of an object are created from the same image (through cropping and augmentation), the invention takes advantage of the fact that different images of the same scene are used, captured by different cameras, e.g., of a test vehicle. This not only promotes the learning of object features, but also the learning of semantic information about the specific scenes and environments.
[0011] A scene can be understood as information from an environmental recording that serves to provide a description of the current environment. In other words, a scene can be described by an environmental recording that provides important information about the current environment.
[0012] Scene information can include, for example, information about the location, a context such as a highway, as well as information about the weather and weather conditions. The time of day can also be important and thus be included as scene information.
[0013] The machine learning model, in particular at least one network, can be trained to map the different views to the same feature space to determine the semantic scene information. This can improve context understanding.
[0014] Furthermore, within the scope of the invention, it can be provided that the image information is specific for capturing the different views of the environment by different image sensors, in particular cameras, and / or results from the capture by the different image sensor sources in the form of the different image sensors. Thus, different images with different views of the same environment can be provided as the image information to represent the same scene and used for training. This enables, in particular, an improved understanding of the scene and reduces the effort required, for example, for manual annotation.
[0015] Optionally, it can be provided that the image information comprises various images in which the views of the same surroundings differ in that an image angle and / or a viewing angle differs. A further advantage within the scope of the invention can be achieved if the image sensor sources comprise at least one wide-angle front camera and / or at least one telephoto front camera and / or at least one side camera and / or at least one rear camera, in particular of a vehicle, in order to provide at least one and / or various zoom sections and / or image sections and / or overlaps in the image information to represent the scene. The cameras can be attached to a vehicle, such as a motor vehicle, in order to monitor the surroundings during a journey and to provide the monitoring, for example, for a driving function for automated and / or autonomous driving.
[0016] It is also advantageous if the image sensor sources are embodied as cameras of a vehicle, so that the image information for representing the respective scene is provided in the form of a traffic scene. This makes it possible to detect hazards in road traffic and / or to provide a driving function for automated and / or autonomous driving. The vehicle can be embodied, for example, as a motor vehicle and / or passenger vehicle and / or autonomous vehicle. The vehicle can have a vehicle device, for example for providing an autonomous driving function and / or a driver assistance system. The vehicle device can be designed to control the vehicle at least partially automatically and / or to accelerate and / or decelerate and / or steer.
[0017] Furthermore, within the scope of the invention, it can be provided that the machine learning model is trained for the determination and in particular classification of the semantic scene information based on image points, in particular pixels, of the image information. This makes it possible to obtain a description of the environment from the image information, preferably via a context and / or weather conditions and / or a time of day and / or a traffic situation. The machine learning model is trained in particular for classification and / or object detection in order to be able to reliably determine the semantic scene information based on the image information. Accordingly, the training can result in a trained machine learning model that can be used for this purpose. The use and thus the inference can be provided, for example, in a vehicle or another device.The data points of the training and / or (in the case of inference) input data of the machine learning model can, for example, be pixels of image information, in particular image data, or be based on these in order to carry out the classification and / or object detection of the data points on the basis of the pixels. The image information can comprise sensor and / or image data which at least partially result from capture with a sensor, preferably a camera sensor, and / or which have been at least partially synthesized, i.e. in particular simulate the real data of a sensor. Specifically, it can be provided that the surroundings of a sensor and / or a vehicle and / or a traffic scene are represented by the values of image points, preferably pixels, of the image data. Classification, preferably image classification and / or object detection, can be provided on the basis of these values. This makes it possible, for example, tothat scene information and / or objects in the traffic scene are detected. The classification can also be provided in the form of semantic segmentation (i.e., a pixel- or area-wise classification). The image data can, for example, be images from a radar sensor and / or an ultrasonic sensor and / or a LiDAR sensor and / or a thermal imaging camera. Accordingly, the images can also be implemented as radar images and / or ultrasonic images and / or thermal images and / or LiDAR images.
[0018] Furthermore, it is possible to evaluate the representation according to the different views by aligning the representations at the feature level, in particular by minimizing a distance calculation (e.g., cosine similarity) to determine the semantic scene information. The minimization of a distance calculation can be achieved, for example, by transforming the feature vectors of the representations into a common feature space. Distance calculation is understood in particular to mean that the similarity between the feature vectors of the representations is calculated. A cosine similarity can be used in this context to calculate the similarity between the feature vectors. The cosine similarity is a measure of the similarity between two feature vectors. For example, the cosine of the angle between the two vectors is determined.
[0019] Furthermore, it is conceivable that the machine learning model has at least or exactly two submodels – preferably in parallel paths – and / or is designed with a teacher-student architecture and / or as a Siamese network. Furthermore, training the machine learning model can include: - feeding image information resulting from the capture of a first camera type, preferably a wide-angle front camera, and / or augmentations of this image information into a first model of the machine learning model, preferably into a teacher model, - Feeding image information resulting from the capture of a second camera type, preferably a telephoto, side-view, or even rear-view camera, and / or augmentations of this image information into a second model of the machine learning model, preferably a student model. The invention also relates to a machine learning model trained by a method according to the invention. Thus, the machine learning model according to the invention offers the same advantages as those described in detail with reference to a method according to the invention.
[0020] The invention also relates to a computer program, in particular a computer program product, comprising instructions that, when executed by a computer, cause the computer to execute the method according to the invention. Thus, the computer program according to the invention provides the same advantages as those described in detail with reference to a method according to the invention.
[0021] The invention also relates to a data processing device configured to carry out the method according to the invention. The device can be, for example, a computer that executes the computer program according to the invention. The computer can have at least one processor for executing the computer program. A non-volatile data memory can also be provided, in which the computer program is stored and from which the computer program can be read by the processor for execution. Furthermore, the device can be designed as a control unit for a vehicle.
[0022] The invention may also provide a computer-readable storage medium that has the computer program according to the invention and / or includes instructions that, when executed by a computer, cause the computer to carry out the method according to the invention. The storage medium is designed, for example, as a data storage device such as a hard disk and / or a non-volatile memory and / or a memory card. The storage medium can, for example, be integrated into the computer.
[0023] Furthermore, the method according to the invention can also be implemented as a computer-implemented method.
[0024] Further advantages, features, and details of the invention will become apparent from the following description, which describes exemplary embodiments of the invention in detail with reference to the drawings. The features mentioned in the claims and in the description may be essential to the invention individually or in any combination. They show: Fig. 1 a schematic visualization of a method, a device, a storage medium and a computer program according to embodiments of the invention. Fig. 2 a schematic visualization of a machine learning model according to embodiments of the invention. Fig. 3 a schematic representation of a vehicle with cameras.
[0025] In Fig. 1, a method 100, a device 10, a storage medium 15 and a computer program 20 according to embodiments of the invention are schematically shown. The method 100 is used for training a Fig. 2 visualized machine learning model 50 for a semantic scene understanding. According to a first method step 101, training data is provided, wherein the training data comprises image information 60 representing a respective scene of an environment 2 (see also Fig. 2 and Fig. 3).
[0026] As in Fig. 3, the image information 60 can result from different image sensor sources 5 in order to represent the environment 2 with different views in the image information 60 for the representation of the respective scene.
[0027] According to Fig. 1 and Fig. 2, in a second method step 102, the machine learning model 50 is trained on the basis of the provided training data to determine semantic scene information 70. Then, according to a third method step 103, provision 103 of the trained machine learning model 50 can be provided.
[0028] As in Fig. 3, the image information 60 can be specific for capturing the different views of the environment 2 by different image sensors 5, in particular cameras 5, and / or result from the capture by the different image sensor sources 5 in the form of the different image sensors 5. In this way, different images 60 with different views of the same environment 2 can be provided as the image information 60 to represent the same scene and used for the training 102. Specifically, the image information 60 can comprise different images 60 in which the views of the same environment 2 differ in that an image angle and / or a viewing angle differs. This results, for example, from Fig.3 using the two laterally arranged cameras 5'. The cameras 5, 5' shown are designed, for example, as a wide-angle front camera and / or a telephoto front camera and / or a side camera and / or a rear camera of a vehicle 1.
[0029] According to exemplary embodiments, the invention further comprises a self-supervised learning paradigm for semantic scene understanding. In other words, the method 100 according to embodiments of the invention can enable semantic scene understanding, in contrast to the semantic object understanding provided in conventional methods. For this purpose, in particular, multiple images 60 recorded by different cameras 5 of the same test vehicle 1 can be fed into a Siamese network. The machine learning model 50 can thus be trained to learn semantic information about the current scene in order to map the different camera views to the same embedding.
[0030] A Siamese network is understood in particular to be a machine learning model 50 in which two or more identical neural networks ("twins") are trained in parallel to perform a common task. In other words, the twins are submodels in parallel paths. The Siamese network can have the following structure: An input layer can be provided. This input layer of the network can receive the input data and forward it to the next layer. A processing layer can have several layers of neurons that process the input data and extract features from it. An exemplary number of neurons for the specified purpose is 128 and depends, among other things, on the complexity of the input data and the number of features to be extracted.The linking layer can connect the identical networks by comparing the features from the processing layers and outputting a similarity score. The output layer can output the network's result based on the common task that the identical networks trained in parallel.
[0031] Furthermore, a loss function can be considered a crucial component of a Siamese network. It measures how well the network detects the similarity or dissimilarity of the input data. One possible loss function for the invention is the cross-entropy loss function.
[0032] According to embodiments of the invention, it may be particularly advantageous to use cameras 5 with different image angles and / or viewing angles. For example, a combination of a wide-angle front camera, a telephoto front camera, and additional side cameras and / or even rear cameras can be used. This creates zoom sections as well as overlaps and unique image sections.
[0033] Embodiments of the invention allow self-supervised learning to be used to enable both semantic understanding of scenes and pre-training on domain-relevant data.
[0034] In a method 100 according to embodiments of the invention, it can be provided that a test vehicle 1 is equipped with multiple image sensor sources 5 in the form of video cameras 5, which record a diverse and comprehensive image of the environment 2. The test vehicle 1 can then be used to collect image information 60 in the form of video data 60 from important and complex scenes. It is particularly valuable to collect data from scenes and areas that match the application area of the model (e.g., to enable area adaptation).
[0035] After data acquisition, a Siamese network can be trained using a self-supervised learning technique. It may be advantageous to use one of the proven methods such as SimSiam [1] or Dino [2] (see references at the end of the description).
[0036] When using Dino [2], the teacher model can, for example, be fed with images from the wide-angle front camera (or extensions of these images), while the student model is fed with images from the telephoto, side or even rear camera (or extensions of these images).
[0037] Both the teacher and student embeddings can be fed into a cross-entropy loss to adjust the latent space of the different camera views. In other words, a function such as a cross-entropy loss can be used to determine the difference between the actual and predicted probability distributions. Specifically, the cross-entropy loss function calculates the difference between the actual probability distribution and the probability distribution predicted by the model.
[0038] Subsequently, only the student can be updated, while the teacher's parameters are calculated by an exponential moving average of the student parameters.
[0039] It is possible that the distribution of the various cameras may vary between student and teacher. For example, it is necessary to evaluate which camera images are best directed to the student and which to the teacher. However, it may be advantageous if the teacher's images contain a comparatively large amount of information about the current scene. Furthermore, it may be beneficial to supplement the images, i.e., to augment them, e.g., through random cropping and color distortion.
[0040] The above explanation of the embodiments describes the present invention exclusively within the scope of examples. Of course, individual features of the embodiments can be freely combined with one another, provided they are technically feasible, without departing from the scope of the present invention. [1] Xinlei Chen and Kaiming He. Exploring Simple Siamese Representation Learning. 2020. [2] Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging Properties in Self-Supervised Vision Transformers. 2021. QUOTES CONTAINED IN THE DESCRIPTION
[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited non-patent literature
[0000] Xinlei Chen and Kaiming He. Exploring Simple Siamese Representation Learning. 2020
[0040] Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging Properties in Self-Supervised Vision Transformers. 2021
[0040]
Claims
[1] Method (100) for training a machine learning model (50) for semantic scene understanding, comprising the following steps: - Providing (101) training data, wherein the training data comprises image information (60) representing a respective scene of an environment (2), wherein the image information (60) results from different image sensor sources (5) in order to represent the environment (2) with different views in the image information (60) for the representation of the respective scene, - training (102) the machine learning model (50) on the basis of the provided training data to determine semantic scene information (70), wherein for this purpose the representation is evaluated according to the different views, - Providing (103) the trained machine learning model (50). [2] Method (100) according to claim 1, characterized bythat the image information (60) is specific for capturing the different views of the environment (2) by different image sensors (5), in particular cameras (5), and / or results from the capture by the different image sensor sources (5) in the form of the different image sensors (5), so that different images (60) with different views of the same environment (2) are provided as the image information (60) to represent the same scene and are used for the training (102). [3] Method (100) according to one of the preceding claims, characterized by that the image information (60) comprises different images (60) in which the views of the same environment (2) differ in that an image angle and / or a viewing angle differs. [4] Method (100) according to one of the preceding claims, characterized bythat the image sensor sources (5) comprise at least one wide-angle front camera and / or at least one telephoto front camera and / or at least one side camera and / or at least one rear camera, in particular of a vehicle (1), in order to provide at least one and / or different zoom sections and / or image sections and / or overlaps in the image information (60) for representing the scene. [5] Method (100) according to one of the preceding claims, characterized by that the image sensor sources (5) are designed as cameras of a vehicle, so that the image information (60) for representing the respective scene is provided in the form of a traffic scene. [6] Method (100) according to one of the preceding claims, characterized bythat the machine learning model (50) is trained for determining and in particular classifying the semantic scene information (70) on the basis of image points, in particular pixels, of the image information (60) in order to obtain a description of the environment (2) from the image information (60), preferably via a context and / or weather conditions and / or a time of day and / or a traffic situation, wherein the representation according to the different views is evaluated by adjusting the representations at the feature level, in particular by minimizing a distance calculation in order to determine the semantic scene information (70). [7] Method (100) according to one of the preceding claims, characterized bythat the machine learning model (50) has at least or exactly two submodels in parallel paths and is preferably designed with a teacher-student architecture and / or as a Siamese network, wherein preferably the training (102) of the machine learning model (50) comprises: - feeding image information (60) resulting from the capture of a first camera type, preferably a wide-angle front camera, and / or augmentations of this image information into a first model of the machine learning model (50), preferably into a teacher model, - feeding image information (60) resulting from the capture of a second camera type, preferably a telephoto, side or even rear camera, and / or augmentations of this image information into a second model of the machine learning model (50), preferably into a student model. [8] Machine learning model (50) which has been trained by a method (100) according to one of the preceding claims. [9] Computer program (20) comprising instructions which, when the computer program (20) is executed by a computer (10), cause the computer (10) to carry out the method (100) according to one of the preceding claims. [10] Device (10) for data processing which is arranged to carry out the method (100) according to one of claims 1 to 7. [11] A computer-readable storage medium (15) comprising instructions which, when executed by a computer (10), cause the computer (10) to perform the steps of the method (100) according to any one of claims 1 to 7.