Method and system for training a base model
By training a basic model using a knowledge graph with domain-specific information, the method addresses the challenge of limited training data in vehicle object recognition and trajectory prediction, achieving enhanced performance and adaptability in autonomous driving tasks.
Patent Information
- Application Number
- DE102023211845
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-28
- Publication Date
- 2025-05-28
AI Technical Summary
Existing machine learning models face challenges in object recognition, trajectory prediction, and motion planning of vehicles when limited or no training data is available, particularly in scenarios where new categories or domain-specific knowledge is required.
A method for training a basic model that incorporates a knowledge graph with domain-specific information, allowing the model to understand spatio-temporal relationships and context within driving scenes, even with limited data. This involves generating information matrices from image data and knowledge graphs, and training a deep learning architecture on these matrices.
The approach enables improved object recognition, semantic segmentation, trajectory prediction, and motion planning by leveraging domain-specific knowledge, allowing the model to generalize better and adapt to new scenarios with limited data.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The invention relates to a method for training a base model for object recognition and / or trajectory prediction and / or motion planning of a vehicle. The invention relates to a system for training a base model for object recognition and / or trajectory prediction and / or motion planning of a vehicle. The invention relates to a method for object recognition and / or trajectory prediction and / or motion planning using a base model trained in this way. The invention further relates to a computer program with program code and a computer-readable data carrier. State of the art
[0002] Today, deep neural networks (DNNs) have proven to be extremely powerful tools for machine learning, especially in the area of visual embedding. Using textual descriptions, base models are trained to generate a visual representation of objects. However, these models reach their limits when dealing with scenarios where no data and / or samples of the objects are available during training.
[0003] To overcome this problem of sparse data, approaches based on object attribute classification have been developed. These methods attempt to identify attributes such as shape, color, text, etc., to compensate for the missing amount of information. However, these approaches often rely on manually created and ill-defined structures, making it difficult to integrate new knowledge and adapt the model to other domains or use cases without starting the entire modeling process from scratch.
[0004] From the scientific publication [1] "Learning Visual Models using a Knowledge Graph as a Trainer."; https: / / arxiv.org / pdf / 2103.00020.pdf" a method for training deep neural networks with semantic knowledge graphs / ontologies is known.
[0005] From the scientific publication [2] “Radford, A., Kim, JW, Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (nd). Learning Transferable Visual Models From Natural Language Supervision” a general method is known that learns a foundation model from text and images.
[0006] From the scientific publication [3] “Santurkar, S., Dubois, Y., Taori, R., Liang, P., & Hashimoto, T. (2022), “Is one caption worth a thousand images? A controlled study on representation learning”, http: / / arxiv.org / abs / 2207.07635 it is known how important the descriptive value of image captions is, especially when the amount of data is small.
[0007] From the scientific publication [4] “LINGO-1: Exploring Natural Language for Autonomous Driving. LINGO-1” is known an open-loop driving annotator that combines vision, speech and action to improve the way a basic driving model is perceived and / or interpreted and / or explained and / or trained by a user.
[0008] Approaches, such as those described in [2], train base models using descriptive image captions. These captions are converted into latent vectors, i.e., word embeddings, via language models. The embeddings are then used to guide the learning process of the base models. Both the text corpus and the images are crawled from the web. However, in learning scenarios where no training data is available, a test and / or inference dataset differs from the training dataset.
[0009] For example, the test and / or inference dataset contains objects such as street signs, etc., that were not present in the training data during training of the base model. However, the authors of the paper [2] clearly pointed out that the approach cannot be adapted to new categories. Thus, new categories cannot be included. Furthermore, the approach presented in the paper [2] is limited to specified training samples and does not integrate human knowledge about a specific domain.
[0010] In summary, in the rapidly advancing world of machine learning, learning scenarios where no training data, or only a small and / or insufficient amount of training data, is available pose a particular challenge. In such learning scenarios, machine learning models, for example, in classification and / or segmentation tasks, face the challenge of achieving prediction accuracies even though the objects to be classified and / or segmented have either never been seen before or only very few training examples exist. The limited availability of data in these novel situations makes it extremely challenging to make accurate predictions by a (trained) classification and / or segmentation model.
[0011] To counteract this problem, as described above, the use of high-level and unique attributes has been shown to be crucial for achieving good performance in such data-sparse scenarios. By focusing on these meaningful attributes / features, machine learning models can improve their predictive ability and thus achieve improved results even in situations where data availability is limited.
[0012] An object of the invention is to provide an improved method and / or system for training a base model for object recognition and / or trajectory prediction and / or motion planning of a vehicle.
[0013] The problem is solved by a method for training a basic model for object recognition and / or trajectory prediction and / or motion planning of a vehicle according to the features of patent claim 1. The problem is solved by a system for training a basic model for object recognition and / or trajectory prediction and / or motion planning of a vehicle according to the features of patent claim 8. Disclosure of the invention
[0014] According to a first aspect, a method for training a base model, in particular for scene understanding, which serves as additional input for object recognition and / or semantic segmentation and / or trajectory prediction and / or movement planning of a vehicle is provided, the method comprising the steps: - Providing a training data set of image data, each of which contains information about at least one driving scene from a viewpoint of the vehicle; - Providing a knowledge graph that contains domain-specific knowledge about the at least one driving scene; - optional division of the image data into a variety of image sections; - generating information matrices corresponding to the image sections by assigning domain-specific knowledge about the at least one driving scene, which is extracted from the knowledge graph and / or directly from the image data, to the plurality of image sections of the image data; - Training the base model based on the information matrices; and - Providing the trained base model, in particular for scene understanding for object recognition and / or trajectory prediction and / or motion planning of the vehicle.
[0015] Instead of the step of dividing the image data into a multitude of image sections, the information matrix can also be derived directly from the image data and / or the knowledge graph.
[0016] The object detection and / or semantic segmentation and / or trajectory prediction and / or motion planning of a vehicle is preferably carried out using a respective separate method that is improved by the present method. The object detection and / or semantic segmentation and / or trajectory prediction and / or motion planning of a vehicle itself is therefore not necessarily part of the present method, but is decoupled from it. In other words, detected objects in the image data, in particular based on existing methods, are either related to the knowledge graph (see step 2) or extracted directly from the image data in step 4 and used to create the information matrix.
[0017] It is understood that the steps according to the invention, as well as other optional steps, do not necessarily have to be performed in the order shown, but can also be performed in a different order. Furthermore, additional intermediate steps can be provided. The individual steps can also comprise one or more substeps without thereby departing from the scope of the method according to the invention.
[0018] A knowledge graph is preferably a structured database containing knowledge and / or information across a broad range of domains or topics. The knowledge graph preferably contains domain-specific information. It is preferably a semantic network that represents information in the form of graphs by linking entities (e.g., people, places, objects, events) with their attributes and relationships. Organizing knowledge in the form of a knowledge graph preferably makes it possible to model and understand complex relationships and / or dependencies between different entities. The graph enables questions to be asked, relationships to be explored, connections to be analyzed, and / or new knowledge to be derived. The knowledge graph is preferably used by AI systems and search engines to provide a more comprehensive and context-related answer to user queries.By linking and evaluating information from the knowledge graph, AI systems can develop a deeper understanding of texts, questions and queries and generate more precise and relevant answers.
[0019] “Providing a training dataset” preferably means that training image and / or video data are made available from a database and / or from an optical sensor to be processed by the base model. Each training image and / or video datum has predetermined image attributes that relate to the object and / or subject contained in the respective training image and / or video datum and / or to the at least one domain. Furthermore, the training image and / or video data can each have at least one class label, which is assigned, for example, by an expert or determined in some other way. Extracting image attributes in this case preferably means that the base model extracts image attributes, for example at the pixel level, from the provided training data.This step is preferred to identify the relevant features from the images and / or videos that will be used for subsequent information processing.
[0020] This method can be used to optimize object recognition and / or semantic segmentation and / or trajectory prediction and / or motion planning of a vehicle based on image and / or video data. Semantic attributes are extracted from the image and / or video data, particularly at the visual level, and are converted into textual image information by integrating the knowledge graph, which is provided in the form of information matrices. The method can also be applied to one-dimensional data (e.g., in production), requiring only embeddings generated by a neural network to provide high-level semantic concepts.
[0021] In this case, a base model or foundation model is trained or created that is trained to gain an understanding of (driving) scenes in the context of autonomous driving. The base model specifically aims to understand the spatial-temporal relationships of entities within the domain and their context in driving scenes. The base model also learns to understand how individual driving scenes develop or could potentially develop over time. The base model also preferably learns how individual sections of a driving scene can be predicted and / or supplemented if, for example, only insufficient image information is available for visual scene evaluation. The goal of training such a base model is to use an understanding of driving scenes to support training and improve the performance of various downstream tasks related to autonomous driving, e.g.Object detection, semantic segmentation, motion prediction, etc. The base model is preferably trained on structured information extracted from autonomous driving datasets (e.g., nuScenes) and represented in the form of knowledge graphs. This allows the base model to learn and / or understand relationships, hierarchies, and / or contextual information about the objects present in a driving scene. Training with domain-specific knowledge allows the base model to acquire specialized expertise in that domain, enabling it to generate more accurate and relevant responses.
[0022] The present base model, which is used to understand driving scenes in autonomous driving, preferably enables machine learning tasks in the field of autonomous driving to support and / or optimize them, such as object detection, trajectory prediction, motion planning, etc. Similar to how large language models (LLMs) learn to understand and / or interpret language from large amounts of written documents, the present base model learns to understand and / or interpret and / or predict driving scenes from large amounts of test drives. Just as LLMs can be used to learn various natural language tasks, e.g., translation, summarization, text analysis, the base model can preferably provide its understanding of driving scenes to support the training and / or improve the performance of machine learning tasks in the field of autonomous driving.The base model trained here can either be adapted and / or trained for specific tasks or serve as a supplementary source for existing prior knowledge. The base model trained here can, for example, be used as a model supplement to a visually operating machine learning model to support it in tasks such as object recognition and / or semantic segmentation and / or trajectory prediction and / or motion planning. The base model is trained on the basis of the knowledge graph for driving scenes. The use of this domain-specific knowledge graph offers several advantages. Firstly, one can leverage domain expertise. Domain-specific knowledge graphs preferably provide detailed information about entities, relationships and / or concepts that are specific to the domain.This allows the base model to develop a deeper understanding of the domain, thus improving the generation of more accurate and / or context-relevant answers. The integration of knowledge graphs also enables improved contextual understanding. Domain-specific knowledge graphs, in particular, enable the base model to better understand the context in which information is presented. This is particularly important for tasks that require a nuanced understanding of domain-specific concepts and terminology. Likewise, the use of knowledge graphs can provide improved fact-checking and / or information retrieval for driving scenes. The base model trained with domain-specific knowledge graphs is preferably better able to verify facts and / or retrieve accurate information related to the specific domain.This is crucial for applications where the accuracy and reliability of the information are of utmost importance. The base model trained here also enables improved handling of domain-specific terminology. Many fields have specialized terminology that general machine learning models may not understand sufficiently well. Training with domain-specific knowledge graphs helps the base model to more easily grasp domain-specific terms and to use and / or process them appropriately. The base model trained here can also provide tailored answer generation. In particular, the base model can generate answers that are tailored to the domain and provide users with more relevant and / or accurate information. This is particularly important in areas such as autonomous driving, where precision is crucial.This base model also reduces ambiguity in the interpretation of a driving scene. This is achieved in particular by using domain-specific knowledge graphs. These help to clearly define terms and / or concepts within a driving scene that can have multiple meanings in different contexts. This reduces the likelihood of generating ambiguous or incorrect answers. The base model trained here also enables better integration into existing systems. Domain-specific base models trained with knowledge graphs can be seamlessly integrated into existing systems and / or workflows within the respective domain, thus offering improved functionality and efficiency.
[0023] In this case, training the base model describes a method for fine-tuning a visual machine learning model in order to adapt it to novel and / or extended domain knowledge. This improves, for example, the recognition of traffic signs. Especially when there is only limited access to (training) data, it is important to use clear and meaningful inscriptions and / or labels within a driving scene in order to train more efficiently, i.e. with less domain-specific data, and yet develop a more powerful base model. The knowledge graph used represents, in particular, individual driving scenes with agents and / or objects in the respective (driving) scenes, as well as the relationships between the agents and additional information from extended domain knowledge, e.g. knowledge from geographical map data.The base model trained here is capable of evaluating what real driving scenes look like and / or how such driving scenes (might) develop over time. The base model is trained to determine a spatiotemporal semantic representation of driving scenes. In particular, the base model features a suitable deep learning architecture (e.g., transformer architecture) for this purpose.
[0024] Without further training or fine-tuning, the trained base model is capable of making predictions about missing zones or sections within a driving scene or semantic concepts in a future scene in a spatial area of interest (e.g., to predict potentially occluded objects). Furthermore, the base model is capable of determining the information matrix or area matrix of a next and / or previous driving scene if the relevant context information is available, particularly by incorporating the knowledge graph. Furthermore, the trained base model is capable of supplementing missing context information for multiple area matrices (scene representations).
[0025] The information matrix can preferably always be derived from the ego perspective of a vehicle, whereby the ego vehicle can be positioned within the image sections over which the information matrix is superimposed. For example, the information matrix can consist of 11 columns and 20 rows, and the ego vehicle can be positioned in an image section corresponding to column 6 and row 5.
[0026] The knowledge graph preferably contains a large amount of information and can be further enriched by additional internal and / or external sources of metadata. Many applications, in particular, automatically provide graph-based metadata in addition to image data (autonomous driving, production, IoT, etc.). The present method allows this structured metadata to be utilized in the form of attributes and, in particular, as a high-level description of the respective objects. By applying a knowledge graph to a specific domain, the attributes of the objects contained in the image and / or video data, regardless of whether they were previously contained in image and / or video data, can be extracted and used for similarity comparisons in the embedding vector space with the objects modeled in the knowledge graph.This opens up new possibilities for extending and adapting the model without requiring a complete redesign. In particular, the combination of deep neural networks and knowledge graphs is shown to be a promising direction for improving model performance in data-sparse scenarios. In this way, integrating prior knowledge based on information from a knowledge graph not only improves solution generation but also increases a model's ability to successfully generalize, classify, and / or segment in new and / or unexplored environments.
[0027] The present approach refers to the conceptualization and / or integration of domain knowledge via semantic axioms. This human-created knowledge can be integrated and adapted to the changes occurring in the respective domain. Furthermore, encoding the knowledge in the form of a knowledge graph enables a flexible representation and can be transformed into a vector space using embedding methods. This vector space is used to determine similarity with the attributes extracted from the visual space.
[0028] The invention further relates to an adaptation of a polynomial reduction algorithm for exploiting word size shifts that are matched to the word size of the computer hardware, is based on such technical considerations and can contribute to producing the technical effect of an efficient hardware implementation of the algorithm.
[0029] The present method can be used for the analysis of (image and / or video) data acquired by a sensor. In this case, the term "image and / or video data" can also be replaced by "sensor data". The sensor can determine measurements of the environment in the form of sensor signals, which can be provided, for example, by the following elements: digital images, e.g., video, radar, lidar, ultrasound, motion, thermal images, audio signals and / or specific data, such as 1D data (e.g., in production). In principle, it is also possible to obtain information about elements encoded by a sensor signal based on the sensor signal. In other words, an indirect measurement can be performed based on a sensor signal used as a direct measurement. This is also referred to as virtual sensing.Furthermore, the present method can be used to classify and / or categorize and / or segment the sensor data, in particular to detect the presence or absence of objects in the sensor data and / or to perform a semantic segmentation of the sensor data, e.g. with regard to traffic signs and / or road surfaces and / or pedestrians and / or vehicles and / or other. The present method can also be used to determine one or more continuous values, i.e. to perform a regression analysis, e.g. with regard to a distance and / or a speed and / or an acceleration and / or a tracking of an element, e.g. an object, in the data. The present method can be used to detect anomalies in a technical system.For example, Gaussian deviations and / or other uncertainty values can be used to detect anomalies. The present method can be used to control and / or support a technical system, such as a computer-controlled machine, such as a robotic system, a vehicle, a household appliance, a power tool, a manufacturing machine, a personal assistant, or an access control system. The present method can be used in a system for transmitting information, such as a monitoring system or a medical (imaging) system. The present method can be used for measuring and / or controlling in such a system. The present method can be used to analyze data (e.g., scalar time series), in particular from a sensor, i.e., a perception system.The present method can be used for subsequent operation and / or support of the operation of the technical system.
[0030] For example, it must be ensured that an automated vehicle does not collide with pedestrians. Based on semantic segmentation, a computer calculates depth information of all pedestrians present in an image space, further calculates a trajectory around these pedestrians, and steers the autonomously driving vehicle so that it follows this trajectory so closely that it does not hit any pedestrians. This also applies in principle to any mobile robot, in order to avoid people who might be in its path and / or outside its movement path. The method according to the invention can be effectively used for this purpose.
[0031] Furthermore, the method according to the invention can be used in combination with a regression algorithm to determine an exact spatial orientation of the vehicle, in particular using data from yaw rate and / or linear acceleration sensors of a vehicle.
[0032] In the present case, a control device is also particularly preferably claimed which is included in an autonomous vehicle and / or a robotic system and / or an industrial machine and on which the present method can be carried out at least partially.
[0033] The present method, in particular an actively learning method, can provide a trained base model that can learn to determine at which operating point an engine's exhaust emissions should be tested. For this purpose, the engine is preferably operated at this operating point, the exhaust emissions are measured, and fed into the actively learning attribute learning model as input data until the model is deemed to be sufficiently good.
[0034] In an automated vehicle, the base model described here, in particular an actively learning one, preferably defines predetermined scenarios for which image and / or video data and / or data from alternative sensors are to be collected.
[0035] In a networked physical system, such as a connected automated vehicle, an anomaly detector can also be used to detect whether a selected frame of a predefined length (e.g., 5 s) from an accelerometer time series contains an anomaly. If so, this frame is transmitted to a back-end computer, where it can be used, for example, to define corner cases for testing the ML system, based on the results of which the connected physical system operates.
[0036] In a preferred embodiment, the base model is trained on the basis of the information matrices to determine spatial-temporal relationships of entities within a driving scene and / or a context of the entities within the driving scene and / or a temporal development of the driving scene or across the driving scene. This also allows for improved prediction of future driving scenes with the support of the base model. During training or inference, the last n driving scenes are preferably considered in order to learn or predict the temporal relationship.
[0037] In a preferred embodiment, the domain-specific knowledge about the at least one driving scene contained in the knowledge graph comprises structured information about the at least one driving scene obtained from autonomous driving datasets, in particular from nuScenes datasets, wherein the structured information comprises relationships and / or hierarchies and / or context-related information about the objects occurring in a respective driving scene. The "nuScenes" datasets are a collection of extensive multimodal datasets specifically designed for research and development in the field of autonomous driving. nuScenes contains data recorded with a variety of sensors such as cameras, lidar, radar, and others. These sensors capture a 360-degree view of the environment.The data was collected in various urban environments, including complex traffic scenarios, varying weather conditions, and times of day. A key feature of the nuScenes datasets is their detailed annotation. Objects in the scenes, such as vehicles, pedestrians, and obstacles, are labeled and categorized, facilitating the development and testing of algorithms for the perception and behavior of autonomous vehicles. Researchers and developers in the field of computer vision and autonomous driving use these datasets to train and test algorithms for object detection, trajectory prediction, scene analysis, and other machine learning-based tasks. nuScenes is widely accessible to the research community, facilitating collaboration and comparison of different approaches in the field of autonomous driving.
[0038] In a preferred embodiment, the training data set of image data is generated from test drives with the vehicle and / or from historical driving data with the vehicle. The training data can also be taken from existing databases. The training data can also be supplemented with additional traffic data, for example, drone images and / or high-resolution satellite images of traffic scenes, in particular of traffic junctions and / or intersections and / or traffic lights.
[0039] In a preferred embodiment, the base model comprises a machine learning model, in particular an autoregression-based transformer model or a masking-based transformer model. An autoregression-based transformer model is a type of model used in natural language processing (NLP). It is based on self-attention and cross-attention mechanisms that make it possible to understand relationships between words in a sentence or between sentences, regardless of their position. Autoregressive means that the transformer model works sequentially, with each prediction based on the predictions made so far. For example, during text generation, an autoregressive transformer model predicts each next word based on the words already generated. A masking model describes a model in which parts of the input data are intentionally "masked" or hidden.The masking model is then trained to reconstruct or predict the masked parts. In natural language processing, for example, certain words in a sentence are masked. The model must then use the context of the surrounding words to correctly guess or generate the masked words. A well-known example of a masking model in NLP is BERT (Bidirectional Encoder Representations from Transformers). BERT trains its prediction abilities by randomly masking words in a text and attempting to guess them based only on their context.
[0040] In other words, training the base model can be done either with a dedicated (transformer-based) deep learning architecture, which can start with randomly initialized weights, or by fine-tuning an existing base model (e.g., T5, Roberta, Llama). For very large base models (e.g., Llama), fine-tuning can also be limited to a few layers of the network, for example, using an adapter or LoRA approach.
[0041] This baseline model is trained using a deep neural network (DNN) to create an embedding space based on a large corpus of information, e.g., text, images, or sound. This baseline model is tailored to specific domains, such as autonomous driving, where spatial and temporal dimensions, as well as interactions between road users, are crucial for prediction tasks. By incorporating the knowledge graph, which follows an underlying ontology, high-level information between agents and their positions in terms of space and time can be encoded. This information is used to train the more efficient and effective baseline model for scene understanding tasks in the field of autonomous driving.
[0042] In a preferred embodiment, if the base model comprises a masking-based transformer model, one or more information entries of the information matrices are masked and / or hidden randomly or in a predetermined manner, thus training the base model to predict and / or determine the masked and / or hidden information entries. To optimize the base model, the masking can also be targeted at entities of particular interest or with poor performance. Furthermore, the masking can also be restricted to spatial regions within the image data that are of particular interest from the vehicle's perspective (e.g., directly in front of the ego vehicle).
[0043] It's worth noting that training can be performed for an entire "next driving scene." This means that, preferably, not only individual matrix fields are masked. Rather, the model learns to predict an entire next driving scene based on the previous driving scenes.
[0044] In a preferred embodiment, the base model comprises a pre-trained large language model (LLM). In other words, the base model is fine-tuned from a pre-trained large language model (LLM). The base model can, for example, comprise a pre-trained language model that is trained solely on domain-specific knowledge or in the domain context. This allows for more efficient training.
[0045] In a preferred embodiment, the number of rows and columns of the information matrices corresponds to a number of image sections, wherein each cell of the information matrices comprises domain-specific knowledge in the form of semantic concepts of the entities or events present in the spatial dimensions of the image sections, wherein the domain-specific knowledge comprises information about road infrastructure facilities and / or pedestrians, and / or traffic signs and / or stop areas and / or construction site markings and / or and / or pedestrian crossings and / or potential vehicle trajectories / paths and / or vehicles annotated with actions, and / or context-relevant information,such as, in particular, a distance travelled since a previous driving scene and / or a difference in the orientation of a road user between the driving scene and the previous driving scene and / or a country and / or an intended route and / or direction.
[0046] In one embodiment, more than two information matrices are used during training to learn temporal behavior across individual driving scenes ("snapshots"). The driving scenes are preferably viewed at 2 Hz. In addition to the more than two information matrices, information about the distance the vehicle traveled between the scenes or the steering angle the vehicle applied can also be specified. Additionally, further metadata, such as a country in which the vehicle is traveling, and / or a route and / or direction of travel of the vehicle, can be specified.
[0047] The determination of the information matrices is based in particular on a driving scene representation in the form of a structured knowledge graph, which in particular forms a generic scene knowledge graph (SKG). Based on this, the relevant information of individual driving scenes, in particular from the perspective of a specific road user, is extracted in the form of subgraphs. A "knowledge graph embedding model" can be used in particular. A knowledge graph embedding model is a model that transfers information from a knowledge graph into a numerical vector space. A knowledge graph is preferably a structure that represents knowledge data in the form of entities (objects) and their relationships. Each entity is preferably represented by a unique node in the graph, and relationships between the entities are represented by edges.The goal of a knowledge graph embedding model is to encode the semantic meaning of entities and relationships in a knowledge graph into vectors, specifically latent vectors, so that mathematical operations such as similarity comparisons and / or clustering can be performed on the embeddings. By embedding in a vector space, the abstract relationships between the entities are represented in a compact and computable form that is more easily processed by other machine learning models. There are various approaches to knowledge graph embedding models, such as TransE, TransR, DistMult, ComplEx, etc.
[0048] These subgraphs are preferably transformed into a two-dimensional spatial area matrix, namely the respective driving scene-specific information matrix. The representation is preferably from the perspective of the vehicle that uses the base model. The information matrix preferably has predefined dimensions (number of rows and columns) and is composed of several predefined zones (size of the sub-areas). The information matrix is preferably constructed from the perspective of the road user, i.e., it includes their spatial-temporal orientation in the respective spatial context. Each zone of the information matrix is preferably filled with the semantic concepts, particularly in the form of natural language or textual information, of the scene elements and / or events present in the spatial dimensions of the zone, such as:Road infrastructure facilities, pedestrians, traffic signs, stop areas, pedestrian crossings, potential vehicle trajectories / paths (e.g., right-turn vehicles), vehicles annotated with their actions (car stops, car parks, car accelerates, ...), etc. A header of the information matrix preferably includes additional context-relevant information, including but not limited to the distance traveled since the previous scene, the participant's orientation difference between the current and previous scene, the country, the intended route / direction, etc.
[0049] Next, the base model, based on a deep learning architecture, is trained in a self-supervised manner using the latent vector representation of the respective information matrix. This is preferably done by concatenating the information matrix (including header information) of a current and a predetermined number of previous driving scenes from the vehicle's perspective and then passing these concatenated information matrices to the base model. The deep learning architecture preferably first generates tokens for each of the data elements (the header) and semantic concepts (in the respective information matrix) and concatenates them with special tokens to specify the context information (e.g., distance, orientation difference, ...) and the position coding (e.g., columns, rows, etc.).The base model, if it is a masking model, preferably begins by learning the information matrices and evolving them in light of the context information by randomly masking one or more zones (or one or more semantic concepts) and / or header information (context). Since the base model knows a correct answer for each one based on the training data, it can calculate the loss and use it to adjust the base model's weights. In this way, the base model learns to focus on the relevant aspects of the training data through attention mechanisms.
[0050] It is understood that individual previous driving scenes can be skipped (e.g., only every 2nd, 3rd, etc. scene is shown). Alternatively, the same driving scene can be provided multiple times (with a driving distance of 0) for data enrichment purposes.
[0051] In a preferred embodiment, the at least one image and / or video datum is captured by at least one optical sensor, in particular a camera and / or a lidar sensor and / or a radar sensor and / or an ultrasonic sensor. In principle, other sensors and / or sensor data are also conceivable, as long as they can be processed by the attribute learning model and / or the knowledge graph embedding model and / or converted into a semantic embedding.
[0052] In a preferred embodiment, the at least one image and / or video datum is generated from existing image and / or video data by data augmentation. Data augmentation refers to a technique in machine learning in which new data points are artificially created by transforming and / or modifying existing data. The goal of data augmentation is to increase the volume and diversity of available training data in order to improve the performance and robustness of machine learning models. Various transformations can be applied during data augmentation, depending on the type of data and the requirements of the model. In the field of image processing, in particular, operations such as cropping, scaling, rotating, mirroring, or adding noise can be applied to generate new images that differ slightly from the original data.
[0053] According to a second aspect, a system for training a base model, in particular for scene understanding, is proposed, which serves as additional input for object recognition and / or semantic segmentation and / or trajectory prediction and / or motion planning of a vehicle. The system comprises an evaluation and / or computing device configured to perform at least the following steps: - Providing a training data set of image data, each of which contains information about at least one driving scene from a viewpoint of the vehicle; - Providing a knowledge graph that contains domain-specific knowledge about the at least one driving scene; - optional division of the image data into a variety of image sections; - generating information matrices corresponding to the image sections by assigning domain-specific knowledge about the at least one driving scene, which is extracted from the knowledge graph and / or directly from the image data, to the plurality of image sections of the image data; - Training the base model based on the information matrices; and - Providing the trained base model, in particular for scene understanding, for object recognition and / or trajectory prediction and / or motion planning of the vehicle.
[0054] Instead of the step of dividing the image data into a multitude of image sections, the information matrix can also be derived directly from the image data and / or the knowledge graph.
[0055] The statements made for the method according to the first aspect apply accordingly to the present system, taking into account linguistic variations. All statements and / or features and / or feature descriptions described in connection with the method apply equally to the system and / or the evaluation and / or control device.
[0056] According to a further aspect, a method for object recognition and / or semantic segmentation and / or trajectory prediction and / or motion planning of a vehicle is provided, which method uses a base model trained in this way.
[0057] According to a further aspect, an evaluation and / or control device of an imaging sensor is proposed, which is designed to carry out the method for object recognition and / or semantic segmentation and / or trajectory prediction and / or movement planning of a vehicle. This means that the evaluation and / or control device represents a device or apparatus that is preferably used in conjunction with an imaging sensor. The evaluation and / or control device can, for example, comprise a camera control unit (CCU) or be contained in such a CCU. An imaging sensor is preferably a sensor that serves to capture visual information and convert it into electronic data, such as in digital image processing or medical imaging.Alternatively, an audio sensor can also be used if the data to be classified and / or segmented includes, for example, audio data.
[0058] The present invention also claims a computer program with program code for executing at least parts of the method according to the first aspect or the third aspect, in one of its embodiments, when the computer program is executed on a computer. In other words, the invention provides a computer program (product) comprising instructions which, when the program is executed by a computer, cause the computer to execute the method / steps of the method according to the invention in one of its embodiments.
[0059] The present invention also proposes a computer-readable data carrier with program code of a computer program for implementing at least parts of the method according to the first aspect or the third aspect, in one of its embodiments, when the computer program is executed on a computer. In other words, the invention relates to a computer-readable (storage) medium comprising instructions which, when executed by a computer, cause the computer to execute the method / steps of the method according to the invention in one of its embodiments.
[0060] The described designs and further training courses can be combined as desired.
[0061] Further possible embodiments, developments and implementations of the invention also include combinations of features of the invention described previously or below with regard to the exemplary embodiments that are not explicitly mentioned. Short description of the drawings
[0062] The accompanying drawings are intended to provide a further understanding of embodiments of the invention. They illustrate embodiments and, in the context of the description, serve to explain principles and concepts of the invention.
[0063] Other embodiments and many of the aforementioned advantages will become apparent upon review of the drawings. The elements illustrated in the drawings are not necessarily drawn to scale.
[0064] They show: Fig. 1 is a schematic flowchart of an embodiment of the present method for classifying an image and / or video datum; Fig. 2(A) is a schematic block diagram of an embodiment of the present method; and Fig. 2(B, C) a schematic block diagram of an embodiment of the present method, showing a division into training phases and an inference phase.
[0065] In the figures of the drawings, the same reference symbols designate the same or functionally identical elements, parts or components, unless otherwise stated.
[0066] Fig. 1 shows a schematic flow diagram of a method for classifying at least one image and / or video datum.
[0067] In any embodiment, the method can be carried out at least partially by a system 1, which for this purpose can comprise several components not shown in detail, for example one or more provision devices and / or at least one evaluation and computing device. It is understood that the provision device can be designed jointly with the evaluation and computing device or can be different from it. Furthermore, the system can comprise a storage device and / or an output device and / or a display device and / or an input device.
[0068] The computer-implemented method comprises, with reference to section (A) of the Fig. 2 at least the following steps: In a step S1, a training data set of image data 200 is provided, each of which contains information about at least one driving scene 202, 204 from a viewpoint of a vehicle 206.
[0069] In a step S2, a knowledge graph 208 is provided which has domain-specific knowledge about the at least one driving scene 202, 204.
[0070] In a step S3, the image data 200 is optionally divided into a plurality of image sections 210. Step S3 is to be considered purely optional. Instead of the step S3 of dividing the image data into a plurality of image sections 210, the information matrix 212 can also be derived directly from the image data 200 and / or the knowledge graph 208.
[0071] In a step S4, information matrices 212 corresponding to the image sections 210 are generated by assigning domain-specific knowledge about the at least one driving scene 202, 204, which is extracted from the knowledge graph 208, to the plurality of image sections 210 of the image data 200.
[0072] In a step S5, a base model 214 is trained on the basis of the information matrices 212.
[0073] In a step S6, the trained basic model 216 is provided for object recognition and / or trajectory prediction and / or movement planning of the vehicle.
[0074] Fig. Figure 2, Section (B) shows the trained base model 216, which was trained to understand the driving scenes 202, 204. The trained base model 216 is a masking model. The trained base model 216 receives as input an image datum 218 of a driving scene 220, in particular acquired by an imaging sensor. A subgraph tailored to the driving scenes 220 is preferably extracted from the knowledge graph 208, from which the information matrix 222 for the driving scene 220 is generated. A portion of the information matrix 222 is, for example, empty because, for example, a lack of information prevailed in these sections 224 of the image datum 218. The masking model can predict the information in these empty sections 224. The information matrix 222 can thus be completed, which is indicated by the reference numeral 226.This improves the accuracy in prediction and / or object detection and / or semantic segmentation of the image scene 220.
[0075] Fig. Figure 2, Section (C) shows the trained base model 216 during inference, i.e., during application for object detection. A visually operating object detection model 228 is augmented by the language-based trained base model 216 in order to improve the model's performance in object detection within a driving scene 230. Table 232 shows that the objects (Car, Pedestrian, PedCrossing) included in the driving scene 232 can each be identified with a high probability. Accuracy is improved by the presently trained base model. QUOTES CONTAINED IN THE DESCRIPTION
[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited non-patent literature
[0000] Learning Visual Models using a Knowledge Graph as a Trainer.”; https: / / arxiv.org / pdf / 2103.00020.pdf
[0004] Radford, A., Kim, JW, Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (nd). Learning Transferable Visual Models From Natural Language Supervision
[0005] Santurkar, S., Dubois, Y., Taori, R., Liang, P., & Hashimoto, T. (2022), “Is one label worth more than a thousand images? A controlled study of representation learning,” http: / / arxiv.org / abs / 2207.07635
[0006] LINGO-1: Natural Language Research for Autonomous Driving. LINGO-1
[0007]
Claims
[1] Method for training a base model (214) for object recognition and / or semantic segmentation and / or trajectory prediction and / or motion planning of a vehicle (206), the method comprising the steps: - Providing (S1) a training data set of image data (200) each containing information about at least one driving scene (202, 204) from a viewpoint of the vehicle (206); - providing (S2) a knowledge graph (208) comprising domain-specific knowledge about the at least one driving scene (202, 204); - optionally dividing (S3) the image data (200) into a plurality of image sections (210); - generating (S4) information matrices (212) corresponding to the image sections (210) by assigning domain-specific knowledge about the at least one driving scene (202, 204), which is extracted from the knowledge graph (208) and / or directly from the image data (200), to the plurality of image sections (210) of the image data (200); - training (S5) the base model (214) on the basis of the information matrices (212); and - Providing (S6) the trained basic model (216), in particular for scene understanding, for object recognition and / or trajectory prediction and / or movement planning of the vehicle (206). [2] Method according to claim 1, wherein the base model (214) is trained on the basis of the information matrices (212) to determine spatial-temporal relationships of entities within a driving scene (202, 204) and / or a context of the entities within the driving scene (202, 204) and / or a temporal development of the driving scene (202, 204). [3] Method according to claim 1 or 2, wherein the domain-specific knowledge about the at least one driving scene (202, 204) comprised in the knowledge graph (208) comprises structured information about the at least one driving scene (202, 204) which is obtained from data sets for autonomous driving, in particular from nuScenes data sets, wherein the structured information comprises relationships and / or hierarchies and / or context-related information about the objects occurring in a respective driving scene (202, 204). [4] Method according to one of the preceding claims, wherein the training data set of image data (200) is generated by test drives with a vehicle and / or by historical driving data with a vehicle. [5] Method according to one of the preceding claims, wherein the base model (214) comprises a machine learning model, in particular an autoregression-based transformer model or a masking-based transformer model. [6] The method of claim 5, wherein, if the base model (214) comprises a masking-based transformer model, one or more information entries of the information matrices (212) are masked and / or hidden randomly or in a predetermined manner so as to train the base model (214) to predict and / or determine the masked and / or hidden information entries. [7] Method according to one of the preceding claims, wherein the base model (214) comprises a pre-trained large language model (LLM). [8] Method according to one of the preceding claims, wherein the number of rows and columns of the information matrices (212) corresponds to a number of image sections, wherein each cell of the information matrices (212) has domain-specific knowledge in the form of semantic concepts of the entities or events present in the spatial dimensions of the image sections (210), wherein the domain-specific knowledge comprises information about road infrastructure facilities and / or pedestrians, and / or traffic signs and / or stop areas and / or construction site markings and / or pedestrian crossings and / or potential vehicle trajectories / paths and / or vehicles annotated with actions, and / or further context-relevant information,in particular how a distance travelled since a previous driving scene and / or a difference in the orientation of a road user between the driving scene and the previous driving scene and / or a country and / or an intended route and / or direction. [9] Method according to one of the preceding claims, wherein the image data (200) are acquired by at least one optical sensor or generated by data augmentation from existing image and / or video data. [10] System (1) for training a base model (214) for object recognition and / or semantic segmentation and / or trajectory prediction and / or movement planning of a vehicle (206), the system (1) comprising an evaluation and / or computing device which is designed to carry out at least the following steps: - providing a training data set of image data (200) each containing information about at least one driving scene (202, 204) from a viewpoint of the vehicle (206); - providing a knowledge graph (208) comprising domain-specific knowledge about the at least one driving scene (202, 204); - optionally dividing the image data (200) into a plurality of image sections (210); - generating information matrices (212) corresponding to the image sections (210) by assigning domain-specific knowledge about the at least one driving scene (202, 204), which is extracted from the knowledge graph (208) and / or directly from the image data (200), to the plurality of image sections (210) of the image data (200); - training the base model (214) based on the information matrices (208); and - Providing the trained base model (216), in particular for scene understanding, for object recognition and / or trajectory prediction and / or movement planning of the vehicle (206). [11] Method for object recognition and / or semantic segmentation and / or trajectory prediction and / or motion planning of a vehicle, which uses a base model (216) trained according to one of claims 1 to 9. [12] Evaluation and / or control device of an imaging sensor, which is designed to carry out a method according to claim 11. [13] Computer program with program code to carry out at least parts of a method according to one of claims 1 to 9 and / or 11 when the computer program is executed on a computer. [14] Computer-readable data carrier with program code of a computer program for carrying out at least parts of a method according to one of claims 1 to 9 and / or 11 when the computer program is executed on a computer.
Citation Information
Cited By
Intelligent driving system data enhancement method based on spatio-temporal joint hierarchical priori knowledge
CN122414271A