Method for answering environment-related questions from a vehicle occupant and driver assistance system

An artificial neural network-based system processes verbal questions and environmental data to provide accurate and efficient answers, improving driver comfort and safety in vehicles.

DE102023005196A1Pending Publication Date: 2025-06-18MERCEDES BENZ GROUP AG
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
DE102023005196
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-16
Publication Date
2025-06-18

AI Technical Summary

Technical Problem

Existing vehicle systems lack the ability to efficiently and accurately answer environment-related questions from occupants using human-like interactions, limiting driver comfort and safety.

Method used

A method utilizing an artificial neural network, particularly a large language model, to process verbal questions and environmental data from various sources, generating answers by combining textual and image information, and outputting responses through a vehicle's output units.

Benefits of technology

Enhances driver comfort and safety by providing quick and accurate answers to environment-related queries, leveraging multiple information sources for enhanced information retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a method for answering environment-related questions (Q) from a vehicle occupant, comprising the following steps (101 to 106): - Recording a question (Q) asked orally by the vehicle occupant about a vehicle environment (5), - Feeding the recorded question (Q) into an artificial neural network (ANN), - Feeding environmental information (10.1 to 10.n) into the artificial neural network (ANN), - Generation of an answer (R) based on processing of the environment-related information (10.1 to 10.n) depending on the required information from the input question (Q) by means of the artificial neural network (ANN) and - Output of the response (R) to the vehicle occupant. Furthermore, the invention relates to a driver assistance system (2).
Need to check novelty before this filing date? Find Prior Art

Description

The invention relates to a method for answering environmental questions from a vehicle occupant. The invention furthermore relates to a driver assistance system having at least one arithmetic unit for carrying out the method.GB 2600093 A discloses a method for capturing an image of the surroundings of a motor vehicle by means of an assistance system, wherein a voice command for controlling a camera unit that is input by a person within the motor vehicle is captured.From WO 2023 / 072400 A1 a method for obtaining a response to a person-placed query relating to a physical object is known.The invention is based on the object of specifying a novel method for answering environmental questions from a vehicle occupant and a driver assistance system which is improved compared with the prior art.The object with respect to the method is achieved according to the invention by the specified features of claim 1. The object with respect to the driver assistance system is achieved according to the invention by the specified features of claim 10.Advantageous embodiments of the invention are the subject matter of the dependent claims.The method according to the invention for answering environmental questions from a vehicle occupant comprises the following steps:detecting a question as to a vehicle environment which is announced by the vehicle occupant,feeding the detected question into an artificial neural network,feeding environmental information into the artificial neural network,generating a response based on processing the environment-related information as a function of requested information from the input question by means of the artificial neural network; andoutputting the response to the vehicle occupant.The invention relates to a method for answering environmental questions from a vehicle occupant, wherein data can be drawn from various information sources and processed in order to generate a response.Human-like interactions between occupants and vehicles are an essential component of luxury vehicles.The advantages achieved with the invention consist in particular in increasing driver comfort and personal safety and driver safety, wherein questions about the immediate environment or about the immediate environment of the vehicle can be answered quickly and very largely precisely. In this case, it is possible to make and answer questions of interest regarding environmental data, for example environmental objects, as well as questions relating to a journey and / or an operation.Examples of questions may be:What for a car marker is the red car driving in front of me?How long the bridge is in front of mir?How many vehicles are in front of mir?How is my visibility?Drive I to a end of jamming?Is a child behind my vehicle?where there is a free parking space around mich?The method according to the invention enables simple and convenient information retrieval with little effort.The detected verbally posed question, in particular a question spoken in speech, can be converted into a text form before being fed into the artificial neural network. A conversion into text form can be carried out by means of at least one transcription unit. The transcription unit can be a so-called speech-to-text program, for example a software or a tool. The occupant's speech can be recorded. The speech is converted to text. For this purpose, algorithms for speech-to-text conversion already known from the prior art can be used. The question, which is in text form, and environment-related information, for example in the form of images, can be supplied as input to the artificial neural network.Alternatively or optionally additionally, the verbal question can be detected by means of a speech recognition, in particular by means of a speech recognition program, a software or a tool. In speech recognition, the words or instructions spoken by the occupant may be analyzed by artificial intelligence.The artificial neural network can comprise at least one large language model, wherein information required in the input question can be determined by means of the large language model and compared with the input environmental information in order to generate a response.Given a textual input and a further information input, for example in image form, the large language model (also referred to as LLM for short) may be capable of generating a suitable textual output. This textual output combines the information required from the textual input with the information from the image. Such multimodal models thus have both the capability of understanding and generating texts and the capability of extracting information from further information sources, for example from images. In this case, a so-called embedding module can be used to code input text and input images into vectors. These vectors can be used by a so-called transformer decoder to autoregressively generate an output text.Training of the artificial neural network, in particular of the large language model, can proceed as follows:Initially, only the text understanding of the artificial neural network, in particular of the multimodal LLM, can be trained. This means that a huge dataset with text, e.g. from the Internet, can be used to train the model to understand texts and generate meaningful texts or text building blocks. A specific training task can be, for example, the completion of a sentence ("next token prediction").Subsequently, the network, in particular the model, can be trained by means of text-image combinations. Such text-image combinations can likewise be obtained from the Internet, for example from blog articles, for example from so-called "interleaved text images" and / or from subtitles of movies. A concrete training task may again be the completion of the text input if only a part thereof is specified ("next token prediction"). This may result in a multimodal LLM.The response can be output by means of a vehicle-side output unit. The output unit can be, for example, a loudspeaker and / or an optical display unit. The response can be output acoustically via a loudspeaker. The output of the response generated by the artificial neural network may be converted into a voice output and played back to the occupant via speakers of the vehicle. The response can be output as text via an optical display unit.The response can be output as speech. For example, the response can be read to the vehicle occupant.The oral question can be detected by means of a vehicle-side input unit. For example, the input unit may comprise at least one microphone. In a development, the vehicle occupant can also question in text form.As environmental information, at least image data and / or fused image data and / or map data and / or raster map data and / or object data and / or static environmental data and / or raw sensor data can be fed into the artificial neural network.A plurality of vehicle-side capturing units, for example cameras, can capture a plurality of images which can be merged into an image. For this purpose, already existing algorithms can be used which merge a plurality of images into one another.The environmental information can be ascertained and provided by at least one vehicle-side detection unit and / or a vehicle-side navigation unit. The Internet can also provide environment-related and object-related information.For example, a question posed by the vehicle occupant, for example after analysis and / or filtering of required information from the question, can be combined with information of one camera or more cameras which observe the vehicle environment, in order to generate and output an environment-related or environment-related response to the question. For example, a question posed by the vehicle occupant, for example after analysis and / or filtering of requested information from the question, can be combined with arbitrary sensor information in order to generate and output an environmental or environment-related response to the question. The response can include any information, i.e. any number of information and / or information from different information sources, such as navigation maps and / or high-definition maps and / or vehicle-side detection units, camera units and / or sensor units. The Internet can also be used as the information source.A plurality of environment-related information items from various sensors, cameras and / or other and different information sources can be fed into the artificial neural network. The environmental information may be combined to generate and output an even more mature response.The artificial neural network, in particular the multimodal LLM, can be integrated into the vehicle, for example into a driver assistance system of the vehicle, for answering questions. The text of the vehicle occupant extracted from the speech can serve as the question input. As environmental information, an image input may be input to the artificial neural network. An image and / or a fused image from a plurality of images of at least one vehicle-side detection unit, for example a camera, can / can be used as information input, for example image input, in particular information feed. Alternatively or optionally additionally, a raster map of the vehicle environment and / or map data, for example navigation map data, can / can be used as information input. Alternatively or optionally additionally, object data, which can be obtained, for example, from a navigation map and / or from the Internet and / or from images, and / or static environment data, can / can serve as information input.The raster map can represent the environment of the vehicle in a bird's eye view. A raster map can be seen as a generalization of an image. Different information can be displayed in different channels of the raster map. A typical image has, for example, 3 channels (RGB), a raster map can also contain significantly more channels depending on the given information. A representation of the raster map may be flexible, and any information may be mapped into more than 3 channels. In this case, existing multimodal LLMs can only be expanded in order to be able to process raster maps with potentially multiple channels instead of images with only 3 channels. This can be an adaptation in the neural network, which originally generates the embedding from the images.The training can furthermore proceed as follows:training the text understanding on existing text data sets, e.g. parts of the Internet. Additional data sets are required.training the model by means of text-raster map combinations. For this purpose, a data record can be created. This consists, for example, of raster maps and appropriate textual descriptions of these raster maps.Typical information which can be displayed in such a raster map is, for example:raw sensor information that can be transformed into a bird's eye view. Sensor information may include, for example, camera images, lidar point clouds, radar point clouds, and / or ultrasound data.detected objects that can be transformed into bird's eye view, wherein object detectors and tracking algorithms are applied to the sensor information in order to detect the objects. The objects can be determined and transformed via so-called bounding boxes, speed vectors (comprising two channels, one channel in x- and one channel in y-direction) and / or object classes (one channel per object class).static information of a map that can be transformed into bird's eye view. Static information may include tracks and their directions (one direction may be encoded in multiple channels) and / or zebra stripes.A driver assistance system comprises at least one arithmetic unit for carrying out the method described above for answering environmental questions from a vehicle occupant. To generate a response, all information from a vehicle sensor system, for example raw sensor information, and other static information, such as map data and / or raster map data, and / or object data, can be combined. A multimodal large language model can be trained to interpret this specific information and, in combination with the question posed, for example converted into text form, process it in an environment-specific manner.Any information sources can be transferred to the raster map representation and thus be used for finding the response. An already installed or present vehicle-side detection unit may be sufficient to provide information detected in the vehicle environment for the purpose of finding a response. For example, a lidar sensor may return sensor measurements in unexposed areas at night. For example, the vehicle occupant may ask for the next parking spot at night.Complex facts that go beyond human capabilities can also be recognized by the artificial neural network, in particular by the large language model, and influence it in the response. For example, the vehicle occupant may ask for a sharp curve whether he must expect a dangerous field path exit after this curve.Exemplary embodiments of the invention are explained in more detail below with reference to drawings.The following are shown: FIG. 1 schematically shows a vehicle having a driver assistance system, comprising at least one computing unit for carrying out a method for answering environmental questions from a vehicle occupant, FIG. 2 schematically shows a block diagram of a method for answering environmental questions from a vehicle occupant, and FIG. 3 schematically shows a further block diagram of the method for answering environmental questions from a vehicle occupant.Corresponding parts are provided with the same reference numerals in all figures.FIG. 1 schematically shows a vehicle 1 having a driver assistance system 2, comprising at least one computing unit 3 for carrying out a method for answering environmental questions Q from a vehicle occupant.The vehicle 1 may include a plurality of sensing units 4. The detection unit 4 can be at least one sensor and / or a camera for detecting a vehicle environment 5.The driver assistance system 2 can comprise a navigation system. The driver assistance system 2, in particular its computing unit 3, can be designed to acquire and process information from different information sources, for example from the navigation system, from the Internet and / or the detection unit 4. The computing unit 3 can be, for example, a processor, a central computing unit or a data processing unit.For the detection of a question Q that is put orally, i.e. acoustically, by a vehicle occupant, the vehicle 1 can comprise at least one input unit 6, for example a microphone.In order to output a response R to the question Q, the vehicle 1 can comprise at least one output unit 7. The output unit 7 can comprise at least one loudspeaker.The driver assistance system 2 can be coupled to the detection unit 4, the input unit 6 and the output unit 7, respectively.A question Q of the vehicle occupant detected by means of the input unit 6 can be transmitted to the computing unit 3.For finding and generating a response R, a plurality of items of environmental information 10.1 to 10.n can be supplied to the arithmetic unit 3.The computing unit 3 can be designed to process the environment-related information 10.1 to 10.n as a function of requested information from the fed-in question Q and to generate at least one response R. Subsequently, the response R in the vehicle 1 can be output to the vehicle occupant.The arithmetic unit 3 can have an artificial neural network ANN for analyzing the fed-in question Q and for processing the environmental information 10.1 to 10.n.FIG. 2 schematically shows a block diagram of a method for answering environmental questions Q from a vehicle occupant.The method may comprise at least the following steps:Step 101: Input or acquisition of a question Q by verbally speaking the question Q from the vehicle occupant.Step 102: Convert Question Q to text form.Step 103: Feeding the question Q in text form to the artificial neural network KNN.Step 104: Feeding environment-related information 10.1 to 10.n to the artificial neural network ANN.Step 105: Generation of a response R by processing the environment-related information 10.1 to 10.n as a function of required information from the fed-in question Q by means of the artificial neural network KNN.Step 106: Output the response R to the vehicle occupant.FIG. 3 schematically shows a further block diagram of the method for answering environmental questions Q from a vehicle occupant.First, the vehicle occupant can activate the driver assistance system 2 by a voice command, for example. Subsequently, the vehicle occupant can place a question Q orally, i.e. talk loud.The question Q announced by the vehicle occupant can be detected by the input unit 6. The detected question Q can be converted into a text form QT by means of a transcription unit not shown in detail. The transcription unit can be part of the arithmetic unit 3. Subsequently, the question Q is fed in text form QT into the artificial neural network ANN, in particular a so-called large language model. The artificial neural network KNN can analyze the question Q and filter information from the question Q required by means of algorithms, for example. Subsequently, depending on the required information, the environment-related information 10.1 to 10.5 required for answering can be determined from the question Q.For example, raw sensor information or sensor data, which can be ascertained using vehicle-side detection units 4, can be used as the environmental information 10.1. As further environmental information 10.2, object data can be used which can be detected, for example, with the vehicle-side detection unit 4. As further environmental information 10.3, static data, for example static environmental data, which have been stored in maps and / or in the navigation system, can be used. For example, as further environmental information 10.4, 10.5, images or image data of at least one first vehicle-side detection unit 4 and images or image data of at least one second vehicle-side detection unit 4 can be used.The environmental information 10.1 to 10.5 can be directly supplied to the artificial neural network KNN. Optionally in addition, this information can be preprocessed before the environment-related information 10.1 to 10.3 is fed into the artificial neural network ANN in order to create a raster map 11. The environment-related information 10.4, 10.5 can be combined to form a fused image 12, in particular to form fused image data. Subsequently, the raster map 11 and / or the fused image 12 can be fed into the artificial neural network ANN.The artificial neural network KNN can then generate a response R from the input question Q on the basis of processing the environmental information 10.1 to 10.n and / or the raster map 11 and / or the fused image 12 as a function of required information. The generated response R can be present as text. The generated response R may be converted to speech using known algorithms. The response R can be output via the output unit 7, for example output as speech to the vehicle occupant.List of reference characters1 Vehicle 2 Driver assistance system 3 Computing unit 4 Detection unit 5 Vehicle environment 6 Input unit 7 Output unit 10.1 to 10.n Environment-related information 11 Raster map 12 Fused image 101 to 106 Step ANN Artificial neural network Q Question QT Text form R ResponseReferences included in the specificationThis list of documents cited by the applicant has been produced in an automated manner and is only included for the better information of the reader. The list is not part of the German patent application or utility model application. The DPMA does not take any adhesion for any faults or omissions.Patent Literature citedGB 2600093 A

[0002] WO 2023 / 072400 A1

[0003]

Claims

Method for answering environmental questions (Q) from a vehicle occupant, characterized bythe following steps (101 to 106): - detection of a question (Q) to a vehicle environment (5) which is announced by the vehicle occupant, - feeding the detected question (Q) to an artificial neural network (ANN), - feeding environmental information (10.1 to 10.n) to the artificial neural network (ANN), - generation of a response (R) on the basis of processing of the environmental information (10.1 to 10.n) as a function of required information from the fed question (Q) by means of the artificial neural network (ANN), and - output of the response (R) to the vehicle occupant.Method according to claim 1, characterised in that the acquired verbal question (Q) is converted into a text form (QT) before being fed into the artificial neural network (KNN).Method according to claim 2, characterised in that the conversion in text form (QT) is carried out by means of at least one transcription unit.Method according to one of the preceding claims, characterized in that the artificial neural network (ANN) comprises at least one large language model, wherein information required in the input question (Q) is determined by means of the large language model and is compared with the input environmental information (10.1 to 10.n) in order to generate a response (R).Method according to one of the preceding claims, characterized in that the response (R) is output by means of a vehicle-side output unit (7).Method according to one of the preceding claims, characterized in that the response (R) is output as speech.Method according to one of the preceding claims, characterized in that the question (Q) which is put into the public is detected by means of a vehicle-side input unit (6).Method according to one of the preceding claims, characterized in that at least - image data and / or - fused image data and / or - map data and / or - raster map data and / or - object data and / or - static environment data and / or - raw sensor data are fed into the artificial neural network (ANN) as environment-related information (10.1 to 10.n).Method according to one of the preceding claims, characterized in that the environmental information (10.1 to 10.n) is determined and provided by at least one vehicle-side detection unit (4) and / or a vehicle-side navigation unit.Driver assistance system (2) having at least one arithmetic unit (3) for carrying out a method for answering environmental questions (Q) from a vehicle occupant according to one of the preceding Claims 1 to 9.

Citation Information

Patent Citations

  • Procedures for operating a speech dialogue system and speech dialogue system

    DE102019217751A1

  • CONTEXT-SENSITIVE SPEECH DIALOGUE SYSTEM

    DE102019219406A1

  • Method and system for providing information requested in a motor vehicle about an object in the vicinity of the motor vehicle

    DE102021130155A1

  • A method for capturing an image of the surroundings of a motor vehicle by an assistance system of the motor vehicle

    GB2600093A

  • Method for conducting dialog between human and computer

    WO2019011356A1