Method for controlling at least one system component of a system according to environmental information generated by the system

The system addresses limitations in environmental information classification by using a camera and visual question answering to generate text descriptions for controlling system components, enhancing environmental awareness and coordination through V2X communication.

DE102024115995A1Pending Publication Date: 2025-12-11AUDI AG
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
DE102024115995
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-07
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Existing systems are limited in their ability to classify and share environmental information, leading to a loss of information and restrictive object classification, which can hinder effective environmental awareness and control.

Method used

A system that captures image data using a camera, annotates position data with a GNSS module, and generates a text description using a visual question answering system to control system components based on detected triggers, enabling the sharing of enriched environmental information via V2X communication.

Benefits of technology

Enhances environmental awareness by accurately identifying relevant triggers, allowing for proactive control measures and improved system coordination through real-time, context-aware information sharing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a method for controlling at least one system component (22) of a system (24) according to environmental information (12) generated by the system (24). First, image data is acquired by means of a detection device (8) of the system (24) and fed into it. Upon detection of a trigger (6), position data (14) of the system (24) is annotated to the image data. Subsequently, a text description message (26) is created for the corresponding image data using a visual question answering system (10) of the system (24), thereby generating environmental information (12). Based on the environmental information (12), at least one system component (22) is then controlled.For this purpose, it is provided that a warning message (16) and / or a navigation instruction and / or the provision of environmental information (12) is initiated as a message (26) to one or at least one other system (28) when it is at a specified distance from the system (24) and / or a trajectory that is at least partially identical to that of the system (24) is being followed.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method for controlling at least one system component of a system according to environmental information generated by the system.

[0002] Sharing environmental information is a goal of V2X, Vehicle-to-Everything communication. For example, to transmit environmental information, a message can be sent from one system to at least one other system, particularly using a Collective Perception Message. Obstacles or objects can be classified by the first system and transmitted in the message as detected obstacles or objects to one or more other systems. Upon detection of such an obstacle, a system component, such as an automatic braking system, can also be activated to reduce speed or bring the system to a complete stop.

[0003] German patent application DE 11 2018 000 899 T5 discloses a method and systems for identifying objects from a 3D point cloud and a 2D image. The method can include determining a first set of 3D suggestions using Euclidean clustering on the 3D point cloud and determining a second set of 3D suggestions from the 3D point cloud based on a 3D convolutional neural network. The method can include combining the first and second sets of 3D suggestions to determine a set of 3D candidates. The method can also include projecting the first set of 3D suggestions onto the 2D image and determining a first set of 2D suggestions using a 2D convolutional neural network.The procedure may involve merging the projected first set of 3D proposals and the first set of 2D proposals to determine a set of 2D candidates, and then merging the set of 3D candidates and the set of 2D candidates.

[0004] EP 3 675 121 A2 discloses one or more embodiments comprising a virtual personal assistant module running on a virtual personal assistant system. The virtual personal assistant module receives initial sensor data from a first sensor, which is part of a plurality of sensors. The virtual personal assistant module analyzes the initial sensor data to generate an initial result. The virtual personal assistant module receives second sensor data from a second sensor, also part of the plurality of sensors. The virtual personal assistant module analyzes the second sensor data and the initial result to generate a second result. The virtual personal assistant module outputs natural language audio to the user based on the second result.

[0005] In the current state of the art, it is disadvantageous that object classification can often only be implemented for a limited number of objects, leading to a loss of information when generating or sharing environmental information. Furthermore, a dependency on the objects defined in messages within the standards can be restrictive. Limiting the number of objects to be classified can also restrict the generated environmental information.

[0006] The invention is based on the objective of providing a system that is able to forward environmental information to other systems depending on the detected context or trigger.

[0007] The problem is solved by the subject matter of the independent patent claims. Advantageous embodiments of the invention are described by the dependent patent claims, the following description, and the figure.

[0008] The invention relates to a method for controlling at least one system component of a system according to environmental information generated by the system. The following steps are performed for this purpose: - Capturing image data using a capture device of the system, such as a camera, and feeding or introducing the image data into the capture device, - Upon detection of a predetermined trigger: Annotating the system's position data to the image data, e.g., by using a GNSS module (GNSS - Global Navigation Satellite System) in the camera and Generating a text description message for the corresponding image data using the system's visual question answering system, thereby generating environment information. - According to the environmental information: Controlling at least one system component, comprising: generating a warning message and / or a navigation instruction and / or providing the environmental information as a message to one or at least one other system if it is within a specified distance of the system and / or is following a trajectory that is at least partially identical to the system.

[0009] The system detects a "trigger," such as a vehicle or an obstacle in close proximity to the system, for example, 5 meters to 500 meters away, or within a predefined radius, such as 5 meters to 2 kilometers. The detection device can therefore be configured or trained to recognize the corresponding triggers, for example, using a specially trained CNN (Convolutional Neural Network).

[0010] The system can also be equipped with various sensors, such as cameras and / or radar sensors and / or lidar sensors.

[0011] "Predetermined" can mean that a (recognizable) trigger is predetermined.

[0012] The (first) system can, using appropriate state-of-the-art sensors, initiate object tracking towards a detected obstacle and thereby estimate a trajectory of the obstacle that is at least partially identical, by making a prediction about its future trajectory based on environmental information and / or a movement pattern of the detected obstacle.

[0013] The visual question answering (VQA) system can be a state-of-the-art VQA system. One possibility is to use a combination of a CNN and a recurrent neural network (RNN). The CNN can be used to extract visual features from the image, while the RNN, for example, designed as a long short-term memory (LSTM), serves to understand the question and generate an appropriate answer.

[0014] Furthermore, the image data can be enriched with GNSS data, so that a detected trigger can be located, at least for the time of detection.

[0015] The VQA system or module can then generate a text description message or text related to the image data. A description of an implementation example can be found below.

[0016] The image data, together with the annotated GNSS data and the text description message, can be considered an example of environmental information. This environmental information can be augmented with further data. For example, data on traffic density, road closures, accidents, and / or construction sites can be used. Another example is the use of data concerning (current) news reports, which provide information about events that could affect the system's environment, e.g., within a radius of 1 to 100 kilometers. As a result, a system component, such as a display device, can be triggered to generate a warning message, for example, that a traffic jam is approaching along the system's current trajectory.In this context, navigation instructions or directions can then be generated, describing a trajectory to bypass the event, such as a traffic jam or obstacle. Additionally or alternatively, at least partial or complete autonomous control can be provided to execute longitudinal and / or lateral guidance of the system according to the navigation instructions. Furthermore, it can be provided that a system component is activated based on this, such as an intelligent heating system, when an outside temperature below a predefined threshold is measured, e.g., between 0 and 10 degrees Celsius.

[0017] The invention enables (contextual) detection of the environment, whereby triggers (e.g., events) that may be relevant to the user are recognized, and controls, e.g., to bypass the trigger, for example by changing the trajectory, in particular by issuing a corresponding navigation instruction, are executed.

[0018] This also has the advantage that conflict-free driving in traffic is likely to be increased by, for example, the system avoiding heavily trafficked routes (at a current or later time).

[0019] The invention also includes further developments that result in additional advantages.

[0020] Further training stipulates that the system can be a vehicle, bicycle, mobile robot, or camera. It is also possible for the system to be a combination of these devices (e.g., a vehicle and a camera). Depending on the configuration, different user scenarios arise, as, for example, a vehicle and a bicycle might travel at least partially different routes. This flexibility allows the system to achieve versatile and / or precise environmental mapping.

[0021] A training course stipulates that the trigger includes: - a query from a user of the system and / or - an obstacle that lies at a predetermined distance or distance from the system and / or is located on a trajectory to be traversed by the system and / or - predefined sensor measurements above a predefined threshold and / or - reaching a predefined zone that is stored in a navigation map of the system and / or - a time-controlled request.

[0022] The query can include a specific question from the user, who wants to know, for example, whether there is an accident and / or roadwork within the next 2 to 20 kilometers along a given trajectory. This enables the user to actively initiate the query and quickly obtain environmental information without having to wait for a potential trigger.

[0023] To detect an obstacle located at a predetermined distance from the system and / or on the system's intended trajectory, image data, for example from a 360-degree camera, is continuously captured. The user can configure what constitutes an obstacle, so that, for instance, other vehicles are not recognized as obstacles, but a bicycle and / or a pedestrian on the road are. This prevents the user from receiving warnings and / or triggering control actions for objects and / or other vehicles that they can already perceive, as the driver must remain attentive. Even if the user has already noticed, for example, a cyclist ahead, this approach still offers the advantage of increased road safety.

[0024] An example of a predefined sensor measurement above a threshold is a temperature measurement below a predefined threshold, which activates a system component, e.g., an intelligent heating system, to bring the interior, e.g., of a vehicle, to a desired or predefined temperature. The system can be equipped with a rain sensor that, when a predefined threshold of the sensor data or the measured rainfall intensity is exceeded, triggers a warning message, such as "Caution, risk of wet conditions," and / or automatically adjusts the system's speed.

[0025] The term "zone" can refer to various areas, such as a 30 km / h zone, a zone around a kindergarten, a school, or a highway. In such zones, the driving style can be adjusted, at least partially autonomously, by controlling at least one system component as described above. Furthermore, a warning message can be issued to advise the driver that a slower speed is advisable, especially if a kindergarten or school is located nearby, i.e., within the detected zone.

[0026] Regarding the time-controlled query, this means that, for example, a query is initiated every t time intervals, where t specifically refers to 5 seconds to 5 minutes. The query can include one or more questions along with predefined answers or answers to be provided in real time. A specific embodiment of this can be found below.

[0027] A further training course stipulates that the time-controlled request is initiated depending on... - a speed of the system and / or - a number and / or type of detected objects in the system's environment and / or - a recognized environment.

[0028] A "time-triggered request" is a request that is automatically initiated after a specific period of time has elapsed. This period of time can be predefined.

[0029] The system's speed can refer to its current speed of movement. In this context, the time-based request can be triggered more or less frequently, depending on how fast the system is moving. At higher speeds, e.g., above 80 to 100 kilometers per hour, the system can be configured to send requests more frequently (e.g., every five seconds) to ensure it always has up-to-date information and can react quickly to changes.

[0030] The term "number and / or type of detected objects in the system's environment" can refer to the fact that the time-based query can be made dependent on how many objects (e.g., obstacles and / or other vehicles and / or people) are detected in the system's vicinity or within a predefined radius, e.g., 500 meters to 10 kilometers, and what type of objects these are (e.g., stationary obstacles or moving objects). If many objects are detected, e.g., more than 4 or 8, or if certain objects (e.g., pedestrians) are present in the environment, the system can increase the query frequency (e.g., a query every five seconds) to obtain accurate and / or up-to-date information or environmental information.

[0031] The system's environment can encompass various factors, such as the type of terrain (urban and / or rural and / or rough terrain) and / or weather conditions and / or lighting conditions. The timed request can be adjusted accordingly, depending on the nature and characteristics of the environment. In complex and / or unsafe environments, it may be necessary to make requests more frequently (e.g., every five seconds) to ensure conflict-free navigation. This can be configured as needed.

[0032] A further training program provides that the text description message is generated using the Visual Question Answering System by linking at least one of several predefined questions with predefined, at least partially similar features from the captured image data, whereby the Visual Question Answering System is trained with a dataset containing image data or features from image data, associated questions and corresponding answers.

[0033] Ideally, a trained VQA model with an accuracy of at least 70% to 90% should be selected on a validated test dataset or dataset.

[0034] It may be provided that, if a text description message is required, the image data, along with at least one of five predefined questions and corresponding information, is passed to the VQA model so that pre-defined or real-time feedback or answers are provided by the VQA model.

[0035] The system can, for example, capture image data of a road on which a cyclist is riding. The VQA module can automatically generate and answer a question such as, "Is there a cyclist on the road?" Based on this answer, the VQA module can then generate a text description message that states, for example: "A cyclist is riding east on XY road." or "A cyclist is 400 meters away from you."

[0036] By asking questions in real time, the VQA module can react immediately to new information and update the text description message accordingly. This ensures that the generated message always reflects the current conditions. The intention is to generate an answer that most closely matches the question based on the training data—that is, the question-answer pairs on which the VQA module was trained—and thus has the highest probability.

[0037] Further training stipulates that the environmental information and / or position data additionally include an orientation angle and / or speed and / or a planned trajectory or route of the system and are automatically and / or manually transmitted to one or more other systems. "Planned trajectory" refers to the predefined path or route that the system is to take. The orientation angle refers to the direction or position in which the system is oriented relative to the detected obstacle. By including the orientation angle, e.g., measurable by an electronic compass and / or MU (Inertial Measurement Unit), the system can determine its position relative to the environment more accurately. This enables precise localization and orientation, which is particularly advantageous in applications such as navigation or robotics.The orientation angle can help to accurately detect and track objects and / or obstacles in the environment, especially when they are located in different viewing directions. Furthermore, the user can selectively forward or send environmental information generated by their system to at least one other system, e.g., in the immediate vicinity, particularly within a radius of 10 meters to 100 kilometers. Additionally or alternatively, it can be configured so that at least one other system in the immediate vicinity automatically receives the environmental information.

[0038] The user has the option to selectively forward, or have forwarded, the environmental information generated by their system to other systems. This allows the user to control data sharing and ensure that only relevant information is sent to at least one other system.

[0039] By strategically sharing environmental information, multiple systems can collaborate and exchange information. This can improve the performance and efficiency of the respective systems that rely on shared data, for example, in system coordination, such as in at least partially autonomous vehicles.

[0040] The ability to send environmental information to other systems in real time allows recipients to immediately access updated information and react accordingly. This is particularly important in dynamic environments where conditions can change rapidly. The range of environmental information transmission can be scaled as needed, from at least one local system in close proximity to at least one remote system, e.g., more than 100 kilometers away, thus covering a larger geographical area. This makes the system flexible and adaptable to various deployment scenarios.

[0041] In summary, the system is designed to exchange or share environmental information (in the form of a message) to obtain a comprehensive and / or accurate perception of its surroundings. The message can, for example, be a textual description, allowing all types of environmental information to be described textually by the VQA module. Suitable devices for creating and receiving such messages can be found in the prior art. Message transmission can occur, for example, via mobile networks, Bluetooth, and / or WLAN (Wireless Local Area Network).

[0042] Forwarding the message, i.e., the environmental information, can help to identify potential dangers early and / or to take preventive measures.

[0043] Further training stipulates that the provisioning of information occurs either by making the message accessible via an edge cloud and / or cloud infrastructure to one or more additional systems, or by direct transmission to them. The sharing of environmental information can therefore be achieved by sharing the message with other systems via an edge cloud and / or cloud infrastructure. Additionally or alternatively, the message can be transmitted directly from one system to at least one other system via V2V (vehicle-to-vehicle) communication.

[0044] Further training involves expanding or updating the system through - a feedback system that receives feedback from the user after the procedure has been carried out and / or - Addition of further datasets containing image data or features from image data, associated questions and corresponding answers and / or - Versioning in the system's backend and / or by receiving update data via mobile network and / or WLAN.

[0045] Once the system has completed the process, it can activate a feedback system to gather feedback from the user. For example, the system can display a rating request to the user to determine their satisfaction.

[0046] For example, after calculating a navigation route or issuing a navigation instruction, the system can display a notification to the user and ask them to rate the accuracy and effectiveness of the route. The feedback can then be collected and used, for example, to trigger improvements, perhaps by having it manually evaluated by a designated team of experts.

[0047] The system can be expanded or updated by adding further datasets to improve its performance. This can include image data or features from image data, as well as related questions and answers. For example, image data from various traffic situations can be linked with questions and corresponding answers to cover a wider range of scenarios. This may involve further processing of the image data collected during the process, for example, manually by the aforementioned team of experts.

[0048] Version control in the system's backend allows for efficient management of updates and improvements. This enables the management of different system versions and the ability to revert to older versions when needed. This makes it easier for the system to revert to previous states if, for example, problems arise with the execution of a control function that did not occur in an earlier version.

[0049] The system can receive update data via cellular network and / or Wi-Fi to keep its functionality and performance up to date. This may include the integration of updates and / or patches and / or new training data.

[0050] For use cases or application situations that may arise during the procedure and are not explicitly described here, it may be provided that, according to the procedure, an error message and / or a request for user feedback is issued and / or a default setting and / or a predetermined initial state is set.

[0051] The invention also includes the control device for the system. The control device can comprise a data processing device or a processor circuit configured to perform an embodiment of the method according to the invention. For this purpose, the processor circuit can comprise at least one microprocessor and / or at least one microcontroller and / or at least one FPGA (Field Programmable Gate Array) and / or at least one DSP (Digital Signal Processor). In particular, a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or an NPU (Neural Processing Unit) can be used as the microprocessor. Furthermore, the processor circuit can comprise program code configured to perform the embodiment of the method according to the invention when executed by the processor circuit. The program code can be stored in a data memory of the processor circuit.The processor setup can be based on at least one circuit board and / or at least one SoC (System on Chip).

[0052] The system according to the invention can be designed, for example, as a motor vehicle or preferably as a motor vehicle, in particular as a passenger car or truck, or as a passenger bus or motorcycle.

[0053] As a further solution, the invention also includes a computer-readable storage medium comprising program code which, when executed by the control device or a computer or computer network, causes it to execute an embodiment of the method according to the invention. The storage medium can be provided at least partially as a non-volatile data storage medium (e.g., as flash memory and / or as an SSD - solid state drive) and / or at least partially as a volatile data storage medium (e.g., as RAM - random access memory). The storage medium can be located within the computer or computer network. However, the storage medium can also be operated, for example, as an app store server and / or cloud server on the internet. The computer or computer network can provide a processor circuit with, for example, at least one microprocessor.The program code can be provided as binary code and / or as assembly code and / or as source code of a programming language (e.g. C) and / or as a program script (e.g. Python).

[0054] The invention also includes combinations of the features of the described embodiments. The invention therefore also includes realizations that each exhibit a combination of the features of several of the described embodiments, provided that the embodiments have not been described as mutually exclusive.

[0055] The following are exemplary embodiments of the invention described. This is illustrated by: Fig. 1 a schematic representation of an embodiment according to the method according to the invention.

[0056] The exemplary embodiments described below are preferred embodiments of the invention. In these exemplary embodiments, the described components each represent individual features of the invention, which can be considered independently of one another and each further develops the invention independently. Therefore, the disclosure is intended to include combinations of features of the embodiments other than those shown. Furthermore, the described embodiments can also be supplemented by further features of the invention already described.

[0057] In the figure, identical reference symbols denote functionally equivalent elements.

[0058] Fig. Figure 1 illustrates a schematic representation of an embodiment according to the idea. Shown are a data set 2, a question selection 4, a trigger 6, a detection device 8, a VQA module 10, environmental information 12, position data 14, a warning message 16, a navigation note 18, the provision of a message 20, a system component 22, a system 24, a message 26, further systems 28, and message processing 30.

[0059] First, as in Fig.As shown in Figure 1, image data is acquired by means of a capture device 8 of the system 24 and fed into it. Upon detection of a trigger 6, position data 14 of the system 24 can be annotated to the image data. Subsequently, a text description message 26 can be created for the corresponding image data using a visual question answering system or VQA module 10 of the system 24, generating environmental information 12. Based on the environmental information 12, at least one system component 22, such as an intelligent heating system, can then be controlled.This may include the provision that a warning message 16 and / or a navigation instruction and / or the provision of environmental information 12 as a message 26, symbolically represented here as V2X, Vehicle-to-Everything communication, is sent to (one or more) further systems 28 when these are at a specified distance from the system 24 and / or a trajectory that at least partially coincides with that of the system 24 is being followed.

[0060] For example, based on a detected trigger 6, a selection of N questions (where N is a number greater than 1) from dataset 2 can be extracted. The corresponding answers can then be retrieved from these questions, generating a text description for the captured image data. Alternatively, the extracted questions can be applied to the image data, causing VQA module 10 to generate a potentially new or previously non-existent text description.

[0061] Processing message 30 from another system 28 may involve enriching it with environmental information 12. This means that the environmental information 12 can be enriched with further position data 14 and / or sensor data in general.

[0062] Overall, the examples show how a method and system 24 for querying, capturing and sharing environmental information can be provided using a VQA module 10 and distribution via a message 26, in particular using V2X communication. Reference symbol list 2 data sets 4 Question Selection 6 triggers 8 Detection device 10 VQA module 12. Environmental information 14 Position data 16 Warning message 18 Navigation note 20. Providing the message 22 System component 24 System 26th message 28 more systems 30. Processing the message QUOTES INCLUDED IN THE DESCRIPTION

[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature

[0000] DE 11 2018 000 899 T5

[0003] EP 3 675 121 A2

[0004]

Claims

[1] Method for controlling at least one system component (22) of a system (24) according to environment information (12) generated by the system (24), comprising the steps: - Acquisition of image data by means of an acquisition device (8) of the system (24) and feeding the image data into the acquisition device (8), - Upon detection of a trigger (6): Annotating position data (14) of the system (24) to the image data and generating a text description message (26) for the corresponding image data using a visual question answering system (10) of the system (24), thereby generating environment information (12), - According to the environmental information (12): Controlling the at least one system component (22), comprising: generating a warning message (16) and / or a navigation instruction and / or providing the environmental information (12) as a message (26) to one or at least one other system (24) when it is at a specified distance from the system (24) and / or is following a trajectory that is at least partially identical to that of the system (24). [2] Method according to claim 1, wherein the system (24) is a vehicle or bicycle or mobile robot or a camera. [3] Method according to any of the preceding claims, wherein the trigger (6) comprises: - a query from a user of the system (24) and / or - an obstacle that lies at a predetermined distance from the system (24) and / or is located on a trajectory to be traversed by the system (24) and / or - predefined sensor measurements above a predefined threshold and / or - reaching a predefined zone stored in a navigation map of the system (24) and / or - a time-controlled request. [4] Method according to claim 3, wherein the time-controlled request is initiated depending on - a speed of the system (24) and / or - a number and / or type of detected objects in the environment of the system (24) and / or - a recognized environment. [5] Method according to one of the preceding claims, wherein the text description message (26) is generated by means of the visual question answering system (10) by linking at least one of several predefined questions with predefined at least partially similar features from the captured image data, wherein the visual question answering system (10) is trained with a data set (2) containing image data or features from image data, associated questions and corresponding answers. [6] Method according to any of the preceding claims, wherein the environmental information (12) and / or the position data (14) additionally comprise an orientation angle and / or a speed and / or a planned trajectory of the system (24) and are automatically and / or provided to one or more further systems (28). [7] Method according to any of the preceding claims, wherein the provision is carried out by making the message (26) accessible via an edge cloud and / or cloud infrastructure to one or more further systems (28) or by direct transmission to them. [8] A method according to any of the preceding claims, wherein the system (24) is extended by - a feedback system (24) that receives feedback from the user after the procedure has been carried out and / or - Addition of further data sets (2) containing image data or features from image data, associated questions and corresponding answers and / or - Versioning in the backend of the system (24) and / or by receiving update data via mobile network and / or WLAN, Wireless Local Area Network. [9] System (24) comprising a control device comprising a processor unit which has program instructions which, when executed by the processor unit, cause it to perform a method according to one of the preceding claims. [10] Motor vehicle comprising a system (24) according to claim 9.

Citation Information

Patent Citations

  • Method for environment recognition for navigation system in car, involves storing data of object or feature in storage, and classifying object or feature by comparison of data after visual inspection of object or feature

    DE102008041679A1

  • procedures to improve road safety

    DE102017008492A1

  • Method and device for selecting and transmitting sensor data from a first to a second motor vehicle

    DE102017223585A1

  • Method and warning device for warning a vehicle user of a potential hazardous situation

    DE102021206634A1

  • Joint 3D object acquisition and orientation estimation via multimodal fusion

    DE112018000899T5