Assistance system and assistance method for object localization
The assistance system with a user interface, imaging sensors, and LLM aids in quickly locating lost vehicle items by analyzing image data, improving search efficiency and reducing frustration.
Patent Information
- Application Number
- PCT/EP2025/063260
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-08
- Filing Date
- 2025-05-14
- Publication Date
- 2026-01-15
AI Technical Summary
Drivers face frustration and delays in locating lost or hidden items within or outside their vehicles, such as keys or an ice scraper, due to inefficient search methods.
An assistance system utilizing a user interface, imaging sensors, and a Large Language Model (LLM) to analyze image data and provide localization results to users, enabling quick and efficient object retrieval.
Facilitates rapid and accurate identification of lost objects by processing image data with LLM, providing visual or auditory guidance to users, enhancing search efficiency and reducing time spent on object localization.
Smart Images

Figure EP2025063260_15012026_PF_FP_ABST
Abstract
Description
[0001] Assistance system and assistance procedure for object localization
[0002] The present disclosure relates to an assistance system for object localization, a vehicle with such an assistance system, an assistance method for object localization, and a storage medium for executing the assistance method. The present disclosure relates in particular to the localization of lost objects inside and / or outside a vehicle.
[0003] State of the art
[0004] A common problem for drivers is finding items they suspect are missing but can't immediately locate. This might involve searching for house or car keys that have accidentally fallen out of a pocket, or for an ice scraper in the trunk. Often, this search leads to frustration and delays, especially when time is of the essence.
[0005] Disclosure of the invention
[0006] The purpose of this disclosure is to specify an assistance system for object localization, a vehicle with such an assistance system, an assistance method for object localization and a storage medium for executing the assistance method, which can help vehicle users to quickly and efficiently locate lost or hidden objects.
[0007] This problem is solved by the subject matter of the independent claims. Advantageous embodiments are specified in the dependent claims.
[0008] According to an independent aspect of the present disclosure, an assistance system for object localization is specified. The assistance system comprises a user interface module configured to receive user input for localizing at least one object in an area of the vehicle; imaging sensors configured to capture at least a portion of the vehicle's area and provide corresponding image data; and an assistance module configured to: transmit the image data together with a localization request for the at least one object to a Large Language Model; and receive a localization result with respect to the at least one object from the Large Language Model, wherein the user interface module is configured to communicate the localization result to the user optically and / or acoustically.
[0009] According to the invention, image data, along with a localization request, is transmitted to a Large Language Model (LLM) upon user input. The LLM analyzes the image data to locate objects specified by the user in the input. For example, the user input might be "Find my house key," and the image data might include images of a vehicle's footwell. The LLM can then analyze the footwell images to locate the house key. The localization result can then be communicated to the user, for example, in the form of a voice message ("Your house keys are under the passenger seat"). As a result, the user can be assisted in quickly and efficiently finding lost or hidden objects.
[0010] The user interface module and the assistance module may include software components / algorithms that are set up to run on at least one processor and thereby perform the functionalities of the respective module.
[0011] The minimum item can be any object of practical use or personal value that can be lost inside and / or outside the vehicle and that the user wishes to find. Typical examples include keys (e.g., house keys, car keys), an ice scraper, a wallet, a mobile phone, sunglasses, a USB stick, a parking ticket, or similar personal items. The user often searches for such items to be able to continue using them immediately or to avoid loss and inconvenience.
[0012] Preferably, the localization request includes a name and / or a description of the at least one object. The name of the at least one object could be, for example, "house key". The description of the at least one object could include properties and / or an appearance of the at least one object, such as "keychain with several keys and a gold key fob".
[0013] Preferably, the user input includes, or is, a verbal user input.
[0014] The user interface module can include at least one input device for receiving the user's speech input and at least one output device for outputting the localization result.
[0015] The at least one input device may include a speech input device for receiving the user's speech input and optionally a touch-sensitive input device, such as a touch field or touch pad, and / or a tactile input device, such as a switch (e.g. push button and / or rotary switch) or other mechanically actuated key elements.
[0016] The at least one output device can comprise at least one loudspeaker and / or at least one display device for outputting the localization result. The at least one display device can comprise a display, in particular an LCD display, a plasma display, or an OLED display. Additionally or alternatively, the at least one display device can comprise a projection device configured to display information directly in the driver's field of vision, in particular to project it onto a windshield.
[0017] Preferably, the user interface module is permanently installed in the vehicle. In some embodiments, the user interface module can be provided by the vehicle's infotainment system, such as a head unit or a pillar-to-pillar system.
[0018] However, the present disclosure is not limited to this, and the user interface module can, in other embodiments, for example, be provided by the user's mobile device or be the user's mobile device itself. The mobile device can be in direct or indirect communication with the vehicle (e.g., Bluetooth, mobile network, etc.) to implement the functionalities according to the invention.
[0019] The term mobile device includes in particular smartphones, but also other mobile phones, personal digital assistants (PDAs), tablets, notebooks, smartwatches, smart glasses and all current and future electronic devices equipped with communication technology.
[0020] To locate at least one object, a Large Language Model (LLM) is used. A Large Language Model is a machine learning model from the field of generative artificial intelligence and is generally used to generate content. Large Language Models are based on neural networks with a transformer architecture and often use deep learning algorithms. ChatGPT is one example of such a Large Language Model.
[0021] The Large Language Model is configured to process at least the image data and the localization request. Specifically, the Large Language Model can be configured to understand and link text and images together. Preferably, the assistance module is configured to transmit the localization request to the Large Language Model in text format.
[0022] Preferably, the assistance module is configured to receive the localization result provided by the Large Language Model in text form.
[0023] Preferably, the Large Language Model is implemented in a central unit (e.g., server or backend) and / or in the cloud.
[0024] Preferably, the vehicle, and in particular the assistance system, comprises a communication module. The vehicle's communication module can be configured for communication via a mobile network. The mobile network can be, for example, an LTE network or a 5G network. This allows the vehicle to communicate with the Large Language Model (or the central unit and / or the cloud) via the mobile network in order to implement the functionalities according to the invention.
[0025] Preferably, the imaging sensor system comprises interior sensors, wherein the interior sensors are configured to detect at least a portion of a first area within the vehicle and to provide corresponding initial image data. In some embodiments, the interior sensors may include one or more interior cameras and optionally one or more light sources (e.g., infrared light sources).
[0026] Preferably, the interior sensors are arranged such that they detect at least partially obscured areas within the vehicle. Partially obscured areas within a vehicle are areas that are not immediately visible or accessible to the user, often due to structural elements, trim, or other fixtures. These areas can be difficult to reach and tend to accumulate items that can easily be lost. Examples include areas between the seats and the center console, areas between the seats and doors, areas under the seats, areas in the trunk (e.g., under the carpet or cover), areas in the door pockets, and areas behind the rear seats.
[0027] Preferably, the imaging sensor system comprises an external sensor system, wherein the external sensor system is configured to detect at least a portion of a second area outside the vehicle and to provide corresponding second image data. In some embodiments, the external sensor system may include one or more external cameras and optionally one or more light sources (e.g., infrared light sources).
[0028] Preferably, the exterior sensors are arranged such that they detect at least partially obscured areas outside the vehicle. Partially obscured areas outside a vehicle are areas that are not immediately visible or accessible to the user, often due to structural elements, trim, or other obstacles. Examples include areas under the bumper and the underbody.
[0029] Preferably, the assistance module is further configured to transmit reference image data to the Large Language Model, wherein the reference image data relates to or shows the area of the vehicle without the at least one object. This allows the Large Language Model to reliably identify the at least one object by comparing the reference image data, in which the at least one object is not visible, with the current image data, in which the at least one object is visible, particularly taking into account a label and / or description of the at least one object contained in the localization request.
[0030] Preferably, the assistance module is configured to transmit the reference image data, along with the (current) image data and the localization request, to the Large Language Model. Preferably, the assistance module is further configured to transmit a description of the reference image data to the Large Language Model along with the reference image data. For example, the reference image data can comprise several images of different areas inside and / or outside the vehicle, with a description of what is visible in each image (e.g., footwell behind the driver's seat, trunk, area between the door and seat, etc.). The description of the reference image data can be used to specify the localization result or the location of the at least one object in more detail for the user.
[0031] Preferably, the assistance module is further configured to detect people in the image data; and to transmit the image data to the Large Language Model only if no people are detected in the image data. This ensures the privacy of the individuals.
[0032] Preferably, the assistance module is further configured to recognize people in the image data; and to anonymize the people in the image data and transmit the image data with the anonymized people to the Large Language Model. For example, the people can be displayed in a single color and / or pixelated. This ensures the privacy of the individuals.
[0033] Preferably, the assistance system is configured to capture image data only from areas inside and / or outside the vehicle where no people are present.
[0034] Preferably, the assistance system is configured to process image data containing people locally within the vehicle using software (e.g., artificial intelligence, LLM, etc.) and optionally to transmit image data without people to the Large Language Model outside the vehicle. For example, a local artificial intelligence or a local Large Language Model can detect whether people are present in the image data and transmit the image data to the Large Language Model outside the vehicle if it is determined that no people are present in the image data. Preferably, the assistance module is further configured to transmit historical image data to the Large Language Model, wherein the historical image data covers the area of the vehicle over a specific period of time.In other words, the past can be taken into account: The assistance system can capture multiple images over time, allowing the Large Language Model, for example, to indicate that at least one object (e.g., a key) is in the vehicle but hidden under another object (e.g., a jacket). Additionally or alternatively, the Large Language Model could determine that at least one object was in a specific position for a certain period of time but has since been removed (e.g., the key was on the seat two days ago but has since been removed). Instead of storing the image, an abstract description such as "empty seat" or "seat with a car key on it" could also be stored. This would enable faster searches in older data.
[0035] According to another independent aspect of the present disclosure, a vehicle, in particular a motor vehicle, is specified. The vehicle comprises the object localization assistance system according to the embodiments of the present disclosure.
[0036] The term "vehicle" includes cars, trucks, vans, buses, motorhomes, motorcycles, etc., used for the transport of people, goods, etc. In particular, the term includes motor vehicles for passenger transport.
[0037] According to another independent aspect of the present disclosure, an assistive method for object localization is specified. The assistive method comprises receiving, by a user interface module, a user input for localizing at least one object in an area of the vehicle; acquiring, by an imaging sensor, at least a part of the area of the vehicle and providing corresponding image data; transmitting the image data together with a localization request for the at least one object to a Large Language Model; receiving a localization result with respect to the at least one object from the Large Language Model; and optically and / or acoustically outputting, by the user interface module, the localization result to the user.
[0038] The object localization assistance procedure can implement the aspects of the object localization assistance system described in this document.
[0039] According to another independent aspect of the present disclosure, a software (SW) program is specified. The SW program can be configured to run on one or more processors and thereby execute the object localization assistance procedure described in this document.
[0040] According to another independent aspect of the present disclosure, a storage medium is specified. The storage medium may include a software program configured to run on one or more processors and thereby execute the object localization assistance procedure described in this document.
[0041] According to another independent aspect of the present disclosure, software with program code is specified. The software is configured to perform the object localization assistance procedure when the software runs on one or more software-controlled devices.
[0042] According to another independent aspect of the present disclosure, an object localization assistance system is specified. The assistance system comprises one or more processors and at least one memory connected to the one or more processors and containing instructions that can be executed by the one or more processors to perform the object localization assistance procedure described in this document. A processor or processor module is a programmable arithmetic unit, i.e., a machine or electronic circuit that controls other elements according to given instructions and thereby executes an algorithm (process).
[0043] Brief description of the drawings
[0044] Examples of the manifestation of the revelation are shown in the figures and are described in more detail below. They show:
[0045] Figure 1 schematically shows a vehicle with an assistance system for object localization according to embodiments of the present disclosure,
[0046] Figure 2 schematically shows objects in one area of the vehicle, and
[0047] Figure 3 shows a flowchart of an assistive method for object localization according to embodiments of the present disclosure.
[0048] Implementations of the revelation
[0049] Unless otherwise noted, the same reference symbols are used for identical and equivalent elements in the following.
[0050] Figure 1 schematically shows a vehicle 10 with an assistance system 100 for object localization according to embodiments of the present disclosure. Figure 2 schematically shows objects OBI, OB2 in an area of the vehicle 10, which the user has lost and wants to find again.
[0051] In the example shown in Figure 2, an object OBI is shown inside vehicle 10 and an object OB2 is shown outside vehicle 10. The object can be any item of practical use or personal value that can be lost inside and / or outside the vehicle and that the user wants to find. Typical examples include keys (e.g., house keys, car keys), an ice scraper, a wallet, a mobile phone, sunglasses, a USB stick, a parking ticket, or similar personal items.
[0052] The assistance system 100 comprises a user interface module 110, which is configured to receive user input (e.g., voice input) from a user for locating at least one object OBI, OB2 in an area of the vehicle 10. In some embodiments, the localization request may include a name and / or a description of the at least one object OBI, OB2. The name of the at least one object OBI, OB2 may, for example, be "house key". The description of the at least one object OBI, OB2 may include properties and / or an appearance of the at least one object OBI, OB2, such as "keychain with several keys and a gold key fob".
[0053] The assistance system 100 further includes an imaging sensor system 120, which is configured to capture at least part of the area of the vehicle 10 and provide corresponding image data. The imaging sensor system 120 can include interior and / or exterior sensors, such as multiple interior and / or exterior cameras.
[0054] The imaging sensor 120 can be arranged to detect at least partially obscured areas inside and / or outside the vehicle 10. Partially obscured areas are areas that are not immediately visible or accessible to the user, often due to structural elements, trim, or other fixtures. Examples of obscured areas inside the vehicle 10 include areas between the seats and the center console, areas between the seats and doors, areas under the seats, areas in the trunk (e.g., under the carpet or cover), areas in the door pockets, and areas behind the rear seat. Examples of obscured areas outside the vehicle 10 include areas under the bumper and an underbody area.The assistance system 100 further comprises an assistance module 130, which is configured to transmit the image data of the imaging sensor 120 together with a localization request based on the received user input for the at least one object OBI, OB2 to a Large Language Model LLM; and to receive a localization result with respect to the at least one object OBI, OB2 from the Large Language Model LLM.
[0055] For this purpose, the vehicle 10, in particular the assistance system 100, can include a communication module configured to communicate with the Large Language Model (LLM). The Large Language Model (LLM) can be implemented in a central unit 20 (e.g., server or backend) and / or in the cloud; however, the present disclosure is not limited to this.
[0056] Communication between the communication module and the Large Language Model (LLM) can take place via a communication link 1, which, for example, uses a mobile network. The mobile network can be, for example, an LTE or 5G network. In some embodiments, the communication module can include a SIM unit. The communication module can be configured to communicate with the central unit 20 via the mobile network. This allows the vehicle 10 to communicate with the Large Language Model (or the central unit 20) via the mobile network in order to implement the functionalities according to the invention.
[0057] The user interface module 110 is further configured to communicate the localization result obtained from the Large Language Model LLM to the user visually and / or audibly (e.g., "Your key is located under the driver's seat").
[0058] In some embodiments, the assistance module 130 can be further configured to transmit reference image data to the Large Language Model (LLM), wherein the reference image data relates to or shows the area of the vehicle 10 without the at least one object OBI, OB2. This allows the Large Language Model (LLM) to reliably identify the at least one object OBI, OB2 by comparing the reference image data, in which the at least one object OBI, OB2 is not visible, with the current image data, in which the at least one object OBI, OB2 is visible, particularly taking into account a label and / or description of the at least one object OBI, OB2 contained in the localization request.
[0059] In some embodiments, the assistance module 130 can be further configured to transmit a description of the reference image data to the Large Language Model (LLM) along with the reference image data. For example, the reference image data can comprise several images of different areas inside and / or outside the vehicle 10, with a description of what is visible in each image (e.g., footwell behind the driver's seat, trunk, area between the door and seat, etc.). The Large Language Model (LLM) can use the description of the reference image data to specify the localization result, or the location of the at least one object OBI, OB2, in more detail for the user.
[0060] In some embodiments, the assistance module 130 can be further configured to detect people in the image data and to transmit the image data to the Large Language Model (LLM) only if no people are detected in the image data. This ensures the privacy of the individuals.
[0061] In some embodiments, the assistance module 130 can be further configured to process image data containing people locally within the vehicle 10 using software (e.g., artificial intelligence, LLM, etc.) and to transmit image data without people to the Large Language Model (LLM) outside the vehicle 10. For example, a local artificial intelligence or a local Large Language Model can detect whether people are present in the image data and transmit the image data (or only a portion of the image data) to the Large Language Model (LLM) outside the vehicle 10 if it is determined that no people are present in the image data. In some embodiments, the assistance module 130 can also be configured to transmit historical image data to the Large Language Model (LLM), wherein the historical image data pertains to the area of the vehicle 10 over a specific period of time.In other words, the past can be taken into account: The assistance system 100 can capture multiple images over time, which allows the Large Language Model (LLM), for example, to indicate that at least one object OBI, OB2 (e.g., a key) is in vehicle 10 but is hidden under another object (e.g., a jacket). Additionally or alternatively, the Large Language Model (LLM) could determine that at least one object OBI, OB2 was in a specific position for a certain period of time but has since been removed (e.g., the key was on the seat two days ago but has since been removed). Instead of storing the image, an abstract description such as "empty seat" or "seat with a car key on it" could also be stored. This would allow for faster searching of older data.
[0062] Figure 3 schematically shows a flowchart of an assistance method 300 for object localization according to embodiments of the present disclosure. The assistance method 300 can be implemented by suitable software that can be executed by one or more processors (e.g., a CPU).
[0063] The assistance method 300 comprises, in block 310, receiving, by a user interface module, a user input for locating at least one object in an area of the vehicle; in block 320, acquiring, by an imaging sensor, at least a part of the area of the vehicle and providing corresponding image data; in block 330, transmitting the image data together with a localization request for the at least one object to a Large Language Model; in block 340, receiving a localization result with respect to the at least one object from the Large Language Model; and in block 350, optically and / or acoustically outputting the localization result to the user by the user interface module. According to the invention, image data together with a localization request are transmitted to a Large Language Model in response to a user input.The Large Language Model analyzes image data to locate objects specified by the user in the input. For example, the user input might be "Find my house key," and the image data could include pictures of a vehicle's footwell. The Large Language Model can then analyze the footwell images to locate the house key. The localization result can then be communicated to the user, for example, as a voice message ("Your house keys are under the passenger seat"). As a result, the user can be helped to quickly and efficiently find lost or hidden objects.
[0064] Although the invention has been further illustrated and explained in detail by means of preferred embodiments, the invention is not limited by the disclosed examples, and other variations can be derived from them by a person skilled in the art without departing from the scope of protection of the invention. It is therefore clear that a multitude of possible variations exist. It is also clear that the embodiments mentioned as examples are truly only examples and are not to be understood in any way as limiting, for example, the scope of protection, the possible applications, or the configuration of the invention.Rather, the preceding description and the description of the figures enable the person skilled in the art to implement the exemplary embodiments in concrete terms, whereby the person skilled in the art, with knowledge of the disclosed inventive concept, can make various changes, for example with regard to the function or the arrangement of individual elements mentioned in an exemplary embodiment, without leaving the scope of protection defined by the claims and their legal equivalents, such as further explanations in the description.
Claims
1. Patent claims 1. An assistance system (100) for object localization, comprising: a user interface module (HO) configured to receive user input for localizing at least one object (OBI, OB2) in an area of the vehicle (10); an imaging sensor (120) configured to capture at least a part of the area of the vehicle (10) and provide corresponding image data; and an assistance module (130) configured to: transmit the image data together with a localization request for the at least one object (OBI, OB2) to a Large Language Model (LLM); and receive a localization result with respect to the at least one object (OBI, OB2) from the Large Language Model (LLM), wherein the user interface module (110) is configured to communicate the localization result to the user optically and / or acoustically.
2. Assistance system (100) according to claim 1, wherein the imaging sensor system (120) comprises interior sensors, wherein the interior sensors are configured to detect at least a part of a first area within the vehicle (10) and to provide corresponding first image data.
3. Assistance system (100) according to claim 1 or 2, wherein the imaging sensor system (120) comprises an external sensor system, wherein the external sensor system is configured to detect at least a part of a second area outside the vehicle (10) and to provide corresponding second image data.
4. Assistance system (100) according to one of claims 1 to 3, wherein the assistance module (130) is further configured to transmit reference image data to the Large Language Model (LLM), wherein the reference image data relates to the area of the vehicle (10) without the at least one object (OBI, OB2).
5. Assistance system (100) according to claim 4, wherein the assistance module (130) is further configured to transmit a description of the reference image data to the Large Language Model (LLM) together with the reference image data.
6. Assistance system (100) according to one of claims 1 to 5, wherein the assistance module (130) is further configured to: to identify persons in the image data; and to transmit the image data to the Large Language Model (LLM) only if no persons are identified in the image data, or to make the persons in the image data unrecognizable and to transmit the image data with the unrecognized persons to the Large Language Model (LLM).
7. Assistance system (100) according to one of claims 1 to 6, wherein the assistance module (130) is further configured to transmit historical image data to the Large Language Model (LLM), wherein the historical image data relates to the area of the vehicle (10) over a certain period of time.
8. Vehicle (10), in particular motor vehicle, comprising the assistance system (100) according to one of claims 1 to 7.
9. Assistance procedures (300) for object localization, including: Receiving (310), through a user interface module (HO), a user input from a user to locate at least one object (OBI, OB2) in an area of a vehicle (10); Capturing (320) by means of an imaging sensor (120) at least a part of the area of the vehicle (10) and providing corresponding image data; Transmit (330) the image data together with a localization request for the at least one object (OBI, OB2) to a Large Language Model (LLM); Receiving (340) a localization result with respect to the at least one object (OBI, OB2) from the Large Language Model (LLM); and optical and / or acoustic output (350), through the user interface module (110), of the localization result to the user.
10. Storage medium comprising a software program configured to run on one or more processors and thereby to execute the assistance method (300) according to claim 9.