Assistance system and assistance procedure for object localization
The assistance system with a user interface, imaging sensors, and LLM efficiently locates lost items in vehicles by analyzing images, addressing the challenge of inefficient item searches and reducing user frustration.
Patent Information
- Application Number
- DE102024119360
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-08
- Publication Date
- 2026-01-08
AI Technical Summary
Drivers face frustration and delays in locating lost or hidden items within or outside their vehicles, such as keys or an ice scraper, due to inefficient search methods.
An assistance system utilizing a user interface, imaging sensors, and a Large Language Model (LLM) to analyze vehicle images and provide localization results to users, enabling quick and efficient object retrieval.
Facilitates rapid and accurate identification of lost items by processing image data with an LLM, providing visual or auditory guidance to users, enhancing search efficiency and reducing time spent on item location.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The present disclosure relates to an assistance system for object localization, a vehicle with such an assistance system, an assistance method for object localization, and a storage medium for executing the assistance method. The present disclosure relates in particular to the localization of lost objects inside and / or outside a vehicle. State of the art
[0002] A common problem for drivers is finding items they suspect are missing but can't immediately locate. This might involve searching for house or car keys that have accidentally fallen out of a pocket, or for an ice scraper in the trunk. Often, this search leads to frustration and delays, especially when time is of the essence. Disclosure of the invention
[0003] The purpose of this disclosure is to specify an assistance system for object localization, a vehicle with such an assistance system, an assistance method for object localization and a storage medium for executing the assistance method, which can help vehicle users to quickly and efficiently locate lost or hidden objects.
[0004] This problem is solved by the subject matter of the independent claims. Advantageous embodiments are specified in the dependent claims.
[0005] According to an independent aspect of the present disclosure, an assistance system for object localization is specified. The assistance system comprises - a user interface module that is set up to receive user input from a user to locate at least one object in an area of the vehicle; - an imaging sensor system designed to capture at least part of the vehicle's area and provide corresponding image data; and - an assistance module that is set up to: - to transmit the image data together with a localization request for at least one object to a Large Language Model, and - to receive a localization result with respect to at least one object from the Large Language Model, - where the user interface module is set up to communicate the localization result to the user visually and / or audibly.
[0006] According to the invention, image data is transmitted to a Large Language Model (LLM) in response to user input, along with a localization request. The LLM analyzes the image data to locate objects specified by the user in the input. For example, the user input might be "Find my house key," and the image data might include images of a vehicle's footwell. The LLM can then analyze the footwell images to locate the house key. The localization result can then be communicated to the user, for example, in the form of a voice message ("Your house keys are under the passenger seat"). As a result, the user can be assisted in quickly and efficiently finding lost or hidden objects.
[0007] The user interface module and the assistance module may include software components / algorithms that are set up to run on at least one processor and thereby perform the functionalities of the respective module.
[0008] The minimum item can be any object of practical use or personal value that can be lost inside and / or outside the vehicle and that the user wishes to find. Typical examples include keys (e.g., house keys, car keys), an ice scraper, a wallet, a mobile phone, sunglasses, a USB stick, a parking ticket, or similar personal items. The user often searches for such items to be able to continue using them immediately or to avoid loss and inconvenience.
[0009] Preferably, the localization request includes a name and / or a description of the at least one object. The name of the at least one object could be, for example, "house key". The description of the at least one object could include properties and / or an appearance of the at least one object, such as "keychain with several keys and a gold key fob".
[0010] Preferably, the user input includes, or is, a verbal user input.
[0011] The user interface module can include at least one input device for receiving the user's speech input and at least one output device for outputting the localization result.
[0012] The at least one input device may include a speech input device for receiving the user's speech input and optionally a touch-sensitive input device, such as a touch field or touch pad, and / or a tactile input device, such as a switch (e.g. push button and / or rotary switch) or other mechanically actuated key elements.
[0013] The at least one output device can comprise at least one loudspeaker and / or at least one display device for outputting the localization result. The at least one display device can comprise a display, in particular an LCD display, a plasma display, or an OLED display. Additionally or alternatively, the at least one display device can comprise a projection device configured to display information directly in the driver's field of vision, in particular to project it onto a windshield.
[0014] Preferably, the user interface module is permanently installed in the vehicle. In some embodiments, the user interface module can be provided by the vehicle's infotainment system, such as a head unit or a pillar-to-pillar system.
[0015] However, the present disclosure is not limited to this, and the user interface module can, in other embodiments, for example, be provided by the user's mobile device or be the user's mobile device itself. The mobile device can be in direct or indirect communication with the vehicle (e.g., Bluetooth, mobile network, etc.) to implement the functionalities according to the invention.
[0016] The term mobile device includes in particular smartphones, but also other mobile phones, personal digital assistants (PDAs), tablets, notebooks, smartwatches, smart glasses and all current and future electronic devices equipped with communication technology.
[0017] To locate at least one object, a Large Language Model (LLM) is used. A Large Language Model is a machine learning model from the field of generative artificial intelligence and is generally used to generate content. Large Language Models are based on neural networks with a transformer architecture and often use deep learning algorithms. ChatGPT is one example of such a Large Language Model.
[0018] The Large Language Model is configured to process at least the image data and the localization request. Specifically, the Large Language Model can be configured to understand and link text and images together.
[0019] Preferably, the assistance module is configured to transmit the localization request to the Large Language Model in text form.
[0020] Preferably, the assistance module is configured to receive the localization result provided by the Large Language Model in text form.
[0021] Preferably, the Large Language Model is implemented in a central unit (e.g., server or backend) and / or in the cloud.
[0022] Preferably, the vehicle, and in particular the assistance system, comprises a communication module. The vehicle's communication module can be configured for communication via a mobile network. The mobile network can be, for example, an LTE network or a 5G network. This allows the vehicle to communicate with the Large Language Model (or the central unit and / or the cloud) via the mobile network in order to implement the functionalities according to the invention.
[0023] Preferably, the imaging sensor system comprises interior sensors, wherein the interior sensors are configured to detect at least a portion of a first area within the vehicle and to provide corresponding initial image data. In some embodiments, the interior sensors may include one or more interior cameras and optionally one or more light sources (e.g., infrared light sources).
[0024] Preferably, the interior sensors are arranged such that they detect at least partially obscured areas within the vehicle. Partially obscured areas within a vehicle are areas that are not immediately visible or accessible to the user, often due to structural elements, trim, or other fixtures. These areas can be difficult to reach and tend to accumulate items that can easily be lost. Examples include areas between the seats and the center console, areas between the seats and doors, areas under the seats, areas in the trunk (e.g., under the carpet or cover), areas in the door pockets, and areas behind the rear seats.
[0025] Preferably, the imaging sensor system comprises an external sensor system, wherein the external sensor system is configured to detect at least a portion of a second area outside the vehicle and to provide corresponding second image data. In some embodiments, the external sensor system may include one or more external cameras and optionally one or more light sources (e.g., infrared light sources).
[0026] Preferably, the exterior sensors are arranged such that they detect at least partially obscured areas outside the vehicle. Partially obscured areas outside a vehicle are areas that are not immediately visible or accessible to the user, often due to structural elements, trim, or other obstacles. Examples include areas under the bumper and the underbody.
[0027] Preferably, the assistance module is further configured to transmit reference image data to the Large Language Model, wherein the reference image data relates to or shows the area of the vehicle without the at least one object. This allows the Large Language Model to reliably identify the at least one object by comparing the reference image data, in which the at least one object is not visible, with the current image data, in which the at least one object is visible, particularly taking into account a label and / or description of the at least one object contained in the localization request.
[0028] Preferably, the assistance module is configured to transmit the reference image data together with the (current) image data and the localization request to the Large Language Model.
[0029] Preferably, the assistance module is further configured to transmit a description of the reference image data to the Large Language Model along with the reference image data. For example, the reference image data can comprise several images of different areas inside and / or outside the vehicle, with a description of what is visible in each image (e.g., footwell behind the driver's seat, trunk, area between the door and seat, etc.). The description of the reference image data can be used to specify the localization result or the location of the at least one object in more detail for the user.
[0030] Preferably, the assistance module is further configured to detect people in the image data; and to transmit the image data to the Large Language Model only if no people are detected in the image data. This ensures the privacy of the individuals.
[0031] Preferably, the assistance module is further configured to recognize people in the image data; and to anonymize the people in the image data and transmit the image data with the anonymized people to the Large Language Model. For example, the people can be displayed in a single color and / or pixelated. This ensures the privacy of the individuals.
[0032] Preferably, the assistance system is configured to capture image data only from areas inside and / or outside the vehicle where no people are present.
[0033] Preferably, the assistance system is configured to process image data containing people locally within the vehicle using software (e.g., artificial intelligence, LLM, etc.) and optionally transmit image data without people to the Large Language Model outside the vehicle. For example, a local artificial intelligence or a local Large Language Model can detect whether people are present in the image data and transmit the image data to the Large Language Model outside the vehicle if it is determined that no people are present in the image data.
[0034] Preferably, the assistance module is further configured to transmit historical image data to the Large Language Model, with the historical image data pertaining to the area of the vehicle over a specific period. In other words, the past can be taken into account: The assistance system can capture multiple images over time, which allows the Large Language Model, for example, to indicate that at least one object (e.g., a key) is in the vehicle but hidden under another object (e.g., a jacket). Additionally or alternatively, the Large Language Model could determine that the at least one object was in a specific position for a certain period of time but has since been removed (e.g., the key was on the seat two days ago but has since been removed). Instead of storing the image, an abstract description such as "empty seat" or "seat with a car key on it" could also be stored.This would allow for faster searching of old data.
[0035] According to another independent aspect of the present disclosure, a vehicle, in particular a motor vehicle, is specified. The vehicle comprises the object localization assistance system according to the embodiments of the present disclosure.
[0036] The term "vehicle" includes cars, trucks, vans, buses, motorhomes, motorcycles, etc., used for the transport of people, goods, etc. In particular, the term includes motor vehicles for passenger transport.
[0037] According to another independent aspect of the present disclosure, an assistive method for object localization is specified. The assistive method comprises receiving, by a user interface module, a user input for localizing at least one object in an area of the vehicle; acquiring, by an imaging sensor, at least a part of the area of the vehicle and providing corresponding image data; transmitting the image data together with a localization request for the at least one object to a Large Language Model; receiving a localization result with respect to the at least one object from the Large Language Model; and optically and / or acoustically outputting, by the user interface module, the localization result to the user.
[0038] The object localization assistance procedure can implement the aspects of the object localization assistance system described in this document.
[0039] According to another independent aspect of the present disclosure, a software (SW) program is specified. The SW program can be configured to run on one or more processors and thereby execute the object localization assistance procedure described in this document.
[0040] According to another independent aspect of the present disclosure, a storage medium is specified. The storage medium may include a software program configured to run on one or more processors and thereby execute the object localization assistance procedure described in this document.
[0041] According to another independent aspect of the present disclosure, software with program code is specified. The software is configured to perform the object localization assistance procedure when the software runs on one or more software-controlled devices.
[0042] According to another independent aspect of the present disclosure, an object localization assistance system is specified. The assistance system comprises one or more processors; and at least one memory connected to the one or more processors and containing instructions that can be executed by the one or more processors to perform the object localization assistance procedure described in this document.
[0043] A processor or processor module is a programmable computing unit, i.e., a machine or an electronic circuit that controls other elements according to given instructions and thereby advances an algorithm (process). Brief description of the drawings
[0044] Examples of the manifestation of the revelation are shown in the figures and are described in more detail below. They show: Fig. 1 schematically a vehicle with an assistance system for object localization according to embodiments of the present disclosure, Fig. 2 schematic objects in an area of the vehicle, and Fig. 3 a flowchart of an assistance method for object localization according to embodiments of the present disclosure. Implementations of the revelation
[0045] Unless otherwise noted, the same reference symbols are used for identical and equivalent elements in the following.
[0046] Fig. Figure 1 schematically shows a vehicle 10 with an assistance system 100 for object localization according to embodiments of the present disclosure. Fig. Figure 2 schematically shows objects OB1, OB2 in an area of vehicle 10, which the user has lost and wants to find again.
[0047] In the example of the Fig. Figure 2 shows an object OB1 inside vehicle 10 and an object OB2 outside vehicle 10. The object can be any item of practical use or personal value that can be lost inside and / or outside the vehicle and that the user wants to find. Typical examples include keys (e.g., house keys, car keys), an ice scraper, a wallet, a mobile phone, sunglasses, a USB stick, a parking ticket, or similar personal items.
[0048] The assistance system 100 comprises a user interface module 110, which is configured to receive user input (e.g., voice input) for locating at least one object OB1, OB2 in an area of the vehicle 10. In some embodiments, the localization request may include a name and / or a description of the at least one object OB1, OB2. The name of the at least one object OB1, OB2 may, for example, be "house key". The description of the at least one object OB1, OB2 may include properties and / or an appearance of the at least one object OB1, OB2, such as "a keyring with several keys and a gold key fob".
[0049] The assistance system 100 further includes an imaging sensor system 120, which is configured to capture at least part of the area of the vehicle 10 and provide corresponding image data. The imaging sensor system 120 can include interior and / or exterior sensors, such as multiple interior and / or exterior cameras.
[0050] The imaging sensor 120 can be arranged to detect at least partially obscured areas inside and / or outside the vehicle 10. Partially obscured areas are areas that are not immediately visible or accessible to the user, often due to structural elements, trim, or other fixtures. Examples of obscured areas inside the vehicle 10 include areas between the seats and the center console, areas between the seats and doors, areas under the seats, areas in the trunk (e.g., under the carpet or cover), areas in the door pockets, and areas behind the rear seat. Examples of obscured areas outside the vehicle 10 include areas under the bumper and an underbody area.
[0051] The assistance system 100 further comprises an assistance module 130, which is configured to transmit the image data of the imaging sensor 120 together with a localization request based on the received user input for the at least one object OB1, OB2 to a Large Language Model LLM; and to receive a localization result with respect to the at least one object OB1, OB2 from the Large Language Model LLM.
[0052] For this purpose, the vehicle 10, in particular the assistance system 100, can include a communication module configured to communicate with the Large Language Model (LLM). The Large Language Model (LLM) can be implemented in a central unit 20 (e.g., server or backend) and / or in the cloud; however, the present disclosure is not limited to this.
[0053] Communication between the communication module and the Large Language Model (LLM) can take place via a communication link 1, which, for example, uses a mobile network. The mobile network can be, for example, an LTE or 5G network. In some embodiments, the communication module can include a SIM unit. The communication module can be configured to communicate with the central unit 20 via the mobile network. This allows the vehicle 10 to communicate with the Large Language Model (or the central unit 20) via the mobile network in order to implement the functionalities according to the invention.
[0054] The user interface module 110 is further configured to communicate the localization result obtained from the Large Language Model LLM to the user visually and / or audibly (e.g., "Your key is located under the driver's seat").
[0055] In some embodiments, the assistance module 130 can be further configured to transmit reference image data to the Large Language Model (LLM), wherein the reference image data relates to or shows the area of the vehicle 10 without the at least one object OB1, OB2. This allows the Large Language Model (LLM) to reliably identify the at least one object OB1, OB2 by comparing the reference image data, in which the at least one object OB1, OB2 is not visible, with the current image data, in which the at least one object OB1, OB2 is visible, particularly taking into account a label and / or description of the at least one object OB1, OB2 contained in the localization request.
[0056] In some embodiments, the assistance module 130 can be further configured to transmit a description of the reference image data to the Large Language Model (LLM) along with the reference image data. For example, the reference image data can comprise several images of different areas inside and / or outside the vehicle 10, with a description of what is visible in each image (e.g., footwell behind the driver's seat, trunk, area between the door and seat, etc.). The Large Language Model (LLM) can use the description of the reference image data to specify the localization result, or the location of the at least one object OB1, OB2, in more detail for the user.
[0057] In some embodiments, the assistance module 130 can be further configured to detect people in the image data and to transmit the image data to the Large Language Model (LLM) only if no people are detected in the image data. This ensures the privacy of the individuals.
[0058] In some embodiments, the assistance module 130 can be further configured to process image data containing people locally within the vehicle 10 using software (e.g., artificial intelligence, LLM, etc.) and to transmit image data without people to the Large Language Model (LLM) outside the vehicle 10. For example, a local artificial intelligence or a local Large Language Model can detect whether people are present in the image data and transmit the image data (or only a portion of the image data) to the Large Language Model (LLM) outside the vehicle 10 if it is determined that no people are present in the image data.
[0059] In some embodiments, the assistance module 130 can be further configured to transmit historical image data to the Large Language Model (LLM), wherein the historical image data relates to the area of the vehicle 10 over a specific period. In other words, the past can be taken into account: The assistance system 100 can capture multiple images over time, which allows the Large Language Model (LLM), for example, to indicate that at least one object OB1, OB2 (e.g., a key) is in the vehicle 10 but is hidden under another object (e.g., under a jacket). Additionally or alternatively, the Large Language Model (LLM) could determine that at least one object OB1, OB2 was in a specific position for a certain period of time but has since been removed (e.g., the key was on the seat two days ago but has since been removed).Instead of saving the image itself, an abstract description such as "empty seat" or "seat with a car key on it" could be saved. This would allow for faster searching of older data.
[0060] Fig. Figure 3 schematically shows a flowchart of an assistance method 300 for object localization according to embodiments of the present disclosure. The assistance method 300 can be implemented by suitable software that can be executed by one or more processors (e.g., a CPU).
[0061] The assistance procedure 300 comprises, in block 310, receiving, by a user interface module, a user input for localizing at least one object in an area of the vehicle; in block 320, capturing, by an imaging sensor, at least a part of the area of the vehicle and providing corresponding image data; in block 330, transmitting the image data together with a localization request for the at least one object to a Large Language Model; in block 340, receiving a localization result with respect to the at least one object from the Large Language Model; and in block 350, optically and / or acoustically outputting, by the user interface module, the localization result to the user.
[0062] According to the invention, image data is transmitted to a Large Language Model (LLM) in response to user input, along with a localization request. The LLM analyzes the image data to locate objects specified by the user in the input. For example, the user input might be "Find my house key," and the image data might include images of a vehicle's footwell. The LLM can then analyze the footwell images to locate the house key. The localization result can then be communicated to the user, for example, in the form of a voice message ("Your house keys are under the passenger seat"). As a result, the user can be assisted in quickly and efficiently finding lost or hidden objects.
[0063] Although the invention has been further illustrated and explained in detail by means of preferred embodiments, the invention is not limited by the disclosed examples, and other variations can be derived from them by a person skilled in the art without departing from the scope of protection of the invention. It is therefore clear that a multitude of possible variations exist. It is also clear that the embodiments mentioned as examples are truly only examples and are not to be understood in any way as limiting, for example, the scope of protection, the possible applications, or the configuration of the invention.Rather, the preceding description and the description of the figures enable the person skilled in the art to implement the exemplary embodiments in concrete terms, whereby the person skilled in the art, with knowledge of the disclosed inventive concept, can make various changes, for example with regard to the function or the arrangement of individual elements mentioned in an exemplary embodiment, without leaving the scope of protection defined by the claims and their legal equivalents, such as further explanations in the description.
Claims
[1] Assistance system (100) for object localization, comprising: - a user interface module (110) that is configured to receive user input from a user to locate at least one object (OB1, OB2) in an area of the vehicle (10); - an imaging sensor system (120) configured to capture at least part of the area of the vehicle (10) and to provide corresponding image data; and - an assistance module (130) that is set up to: - to transmit the image data together with a localization request for at least one object (OB1, OB2) to a Large Language Model (LLM); and - to receive a localization result with respect to at least one object (OB1, OB2) from the Large Language Model (LLM), - wherein the user interface module (110) is set up to communicate the localization result to the user visually and / or audibly. [2] Assistance system (100) according to claim 1, wherein the imaging sensor system (120) comprises interior sensors, wherein the interior sensors are configured to detect at least a part of a first area within the vehicle (10) and to provide corresponding first image data. [3] Assistance system (100) according to claim 1 or 2, wherein the imaging sensor system (120) comprises an outside space sensor system, wherein the outside space sensor system is configured to detect at least a part of a second area outside the vehicle (10) and to provide corresponding second image data. [4] Assistance system (100) according to one of claims 1 to 3, wherein the assistance module (130) is further configured to transmit reference image data to the Large Language Model (LLM), wherein the reference image data relates to the area of the vehicle (10) without the at least one object (OB1, OB2). [5] Assistance system (100) according to claim 4, wherein the assistance module (130) is further configured to transmit a description of the reference image data to the Large Language Model (LLM) together with the reference image data. [6] Assistance system (100) according to one of claims 1 to 5, wherein the assistance module (130) is further configured to: - To identify people in the image data; and - to only transmit the image data to the Large Language Model (LLM) if no people are recognized in the image data, or - to anonymize the people in the image data and to transmit the image data with the anonymized people to the Large Language Model (LLM). [7] Assistance system (100) according to one of claims 1 to 6, wherein the assistance module (130) is further configured to transmit historical image data to the Large Language Model (LLM), wherein the historical image data relates to the area of the vehicle (10) over a certain period of time. [8] Vehicle (10), in particular motor vehicle, comprising the assistance system (100) according to any one of claims 1 to 7. [9] Assistance procedures (300) for object localization, comprising: - Receiving (310), through a user interface module (110), a user input from a user to locate at least one object (OB1, OB2) in an area of a vehicle (10); - Detection (320), by means of an imaging sensor (120), of at least part of the area of the vehicle (10) and provision of corresponding image data; - Transmitting (330) the image data together with a localization request for the at least one object (OB1, OB2) to a Large Language Model (LLM); - Receiving (340) a localization result with respect to the at least one object (OB1, OB2) from the Large Language Model (LLM); and - optical and / or acoustic output (350), through the user interface module (110), of the localization result to the user. [10] Storage medium comprising a software program configured to run on one or more processors and thereby to execute the assistance method (300) according to claim 9.
Citation Information
Patent Citations
SYSTEM AND METHOD FOR PROVIDING A MISTAKE NOTIFICATION
DE102019114392A1
Method for detecting at least one object in a vehicle interior
DE102022004324A1
USE OF SCENE-AWARE CONTEXT FOR DIALOGUE-ORIENTED AI SYSTEMS AND APPLICATIONS
DE102023124120A1
Apparatus and method for detecting falling object
US20200074640A1
Electronic apparatus and operation method thereof
US20200134334A1