Method for detecting a mobile terminal in a vehicle

By combining multiple image captures to eliminate reflections and adapt to lighting conditions, the method improves the reliability of recognizing mobile terminal content in vehicles, facilitating intuitive vehicle function control.

WO2025149193A1PCT designated stage expired Publication Date: 2025-07-17BAYERISCHE MOTOREN WERKE AG
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/079319
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-08
Filing Date
2024-10-17
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing methods for recognizing content on a mobile terminal in a vehicle are hindered by reflections and varying lighting conditions, particularly in ambient light, making it difficult to reliably capture and interpret displayed information without specialized coatings or films.

Method used

A method involving multiple image captures from different angles or positions, combined to eliminate reflections and improve recognition, using image processing to align and merge views, allowing for reliable content recognition without additional hardware modifications.

Benefits of technology

Enhances the reliability and efficiency of recognizing content on mobile devices within vehicles by effectively removing reflections and adapting to varying lighting conditions, enabling seamless control of vehicle functions without additional device connections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024079319_17072025_PF_FP_ABST
    Figure EP2024079319_17072025_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method for detecting a mobile terminal (2) in a vehicle. According to the method, first image data is captured by means of at least one image capturing device (4) in the interior (1) of a vehicle, said first image data including at least one first view of a terminal (2) of the user (10) at a first point in time. Second image data is captured by means of the image capturing device (4), said second image data including at least one second view of the terminal (2) of the user (10) at a second point in time. The first and second image data are combined in order to obtain a composite view of the terminal (2) of the user (10). Contents (22) displayed on the terminal (2) of the user (10) are detected by analyzing the combined image data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Method for detecting a mobile terminal in a vehicle

[0002] The present invention relates to a method for detecting a mobile terminal in a vehicle. In particular, content displayed on a mobile terminal is to be captured using an image capture device in order to control a function of the vehicle.

[0003] Modern vehicles have a multitude of functions that can be controlled by a user, i.e., accessed, started, stopped, configured, etc. This can include, in particular, infotainment system functions such as entering a navigation destination, playing a piece of music, calling a contact, or the like. These infotainment system functions are typically controlled by the user interacting directly with the vehicle, for example, by pressing buttons or touching a touchscreen, or by voice control or gesture control.

[0004] However, operating the system in the vehicle itself can sometimes be cumbersome, for example when data such as a navigation destination or the title of a song has to be entered using a keyboard on a touchscreen. Such input is often easier on a user's device, such as a smartphone. In some cases, the data is even already available on the smartphone, for example if a user has already listened to a particular song on their smartphone and now wants to hear it in the vehicle. Therefore, methods are known in which data is transferred from the user's device to the vehicle, for example via Bluetooth or a cloud service.

[0005] It is also known to capture content shown on the display of a mobile device using a camera in the vehicle. Since modern vehicles are often equipped with corresponding cameras, for example for interior monitoring, this infrastructure can be used to read in displayed content. The content can be displayed text that is recognized via OCR. This way, navigation destinations, for example, can be easily transferred from a smartphone to the vehicle. Machine-readable codes such as QR codes can also be used and captured by a camera in the vehicle. In order to reliably recognize displayed content, in particular text or codes, sufficient capture of the displayed content is necessary. In particular, reliable recognition requires capturing the entire content or at least parts of the content that are relevant for recognition.However, particularly in ambient light outdoors, especially due to sunlight, glare or reflections often occur, which can depend on the position in which a user holds their smartphone relative to the camera in the vehicle. If it is darker inside the vehicle than outside, such reflections can be particularly strong. It is therefore often difficult to hold the smartphone up to the camera without reflections, even when trying to move it back and forth. To prevent or reduce reflections, special films or coatings are known that can be applied to the display of a smartphone. However, such a film may not always be available and may also be undesirable for some users.

[0006] The present invention is based on the object of providing an improved method for recognizing a mobile device in a vehicle. In particular, the recognition of content displayed on a mobile device is to be improved.

[0007] This object is achieved according to the teaching of the independent claims. Various embodiments and further developments of the invention are the subject of the dependent claims.

[0008] A first aspect of the invention relates to a method, in particular a computer-implemented method, for recognizing a mobile terminal in a vehicle. In the method, first image data are captured by at least one image capture device in the interior of a vehicle, wherein the first image data contain at least a first view of a user's terminal at a first point in time. Furthermore, second image data are captured by the image capture device and / or another image capture device in the interior of the vehicle, wherein the second image data contain at least a second view of the user's terminal at a second point in time. The first and second image data are combined to obtain a combined view of the user's terminal. Content displayed on the user's terminal is recognized by evaluating the combined image data.The aforementioned method according to the first aspect is therefore based in particular on the evaluation of image data captured by an image capture device, such as a camera. In this case, several views of the user's terminal device are recorded using the image capture device (or possibly several image capture devices) so that content on the terminal device, in particular content displayed on a display of the terminal device, can be recognized therefrom. By capturing several views, which are then combined to form an overall view, deficiencies in the individual views that could hinder the recognition of the content based on the individual view can be eliminated. These can in particular be reflections on the display of a mobile terminal device, such as a smartphone. The views should be different in order to obtain an overall view without deficiencies.In practice, this can easily occur if a user rarely holds their device completely still. The views can be captured by the same image capture device or by different image capture devices located at different positions within the vehicle interior. Reflections in the overall view can thus be replaced by corresponding parts in other views, thus eliminating them. This allows for better and more reliable recognition of displayed content.

[0009] The term “vehicle” as used herein refers in particular to a passenger car, including all types of motor vehicles, hybrid and battery-powered electric vehicles, as well as vehicles such as vans, buses, trucks, delivery vans and the like.

[0010] The term "image capture device" used here refers in particular to a camera, especially a digital camera. The camera can capture still images (photos) or moving images (videos), especially in the visible light range. An image capture device can capture or record such an image and output corresponding image data. The image capture device can also serve as a monitoring device, i.e., the image capture device and the monitoring device can be one and the same device. It is understood that multiple image capture devices can be present in the interior of a vehicle, which can be located at different locations. Captured image data can then be processed by a central control device.The term "mobile device" as used here refers in particular to electronic devices that are portable and capable of communicating wirelessly with other devices or networks. These are handy devices that can often be carried in a pocket or held in the hand. Mobile devices include smartphones, tablets, laptops, smartwatches, and other similar devices. These devices generally use cellular networks, Wi-Fi, Bluetooth, or UWB to establish a wireless connection and enable data transmission. The mobile device is particularly configured to display content, in particular on a screen. The user of a device can in particular be the driver of the vehicle. However, it is understood that a passenger and, in principle, any vehicle occupant can also be a user of the device within the meaning of the invention.

[0011] The term "content" of a terminal device, as used herein, refers in particular to information presented on a display device, such as a display, of a terminal device. The content does not necessarily have to be displayed on the terminal device in an app that corresponds to the type of vehicle function. The content can simply consist of, for example, unformatted text or machine-readable code. Nevertheless, the content can also be examined with regard to its presentation, such as font format, colors, shapes, logos, or the like, which can be used to identify the vehicle function. A code can also contain more than just text, such as additionally specifying the type of content or even a corresponding control code for the vehicle function.

[0012] The term "monitoring device" used here refers in particular to a device suitable for monitoring the interior of a vehicle, in particular the vehicle occupants, and in particular their positions and movements within the space. This can be an image capture device. An infrared (IR) camera (e.g., a near-infrared (NIR) camera) can be used. An IR image is well suited for monitoring because it is robust against changing light conditions, for example, even in strong sunlight. Recordings are also possible in the dark. It goes without saying that a suitable IR light source can be present. It is also conceivable to create a three-dimensional model of the driver or the vehicle interior, for example, using a time-of-flight (TOF) camera.Using the three-dimensional model, distances to the camera, and thus distances between objects and their movements can be calculated. The term “function of a vehicle” (synonymous with the term “vehicle function”) as used herein is to be understood in particular as a function of an infotainment system of a vehicle. This can include proprietary applications of the vehicle manufacturer and / or applications implemented by third-party providers. The control of a vehicle function, which is also described here, can in particular include calling up or starting a vehicle function assigned to a recognized content depending on the recognized content. However, the control of a vehicle function can also merely be a display of data relating to a recognized content orthe associated vehicle function (possibly together with the recognized content), for example on a screen of the vehicle, in particular a display of the infotainment system.

[0013] The terms "comprises," "includes," "includes," "has," "has," "with," or any other variation thereof, as used herein, are intended to cover non-exclusive inclusion. For example, a method or apparatus that includes or has a list of elements is not necessarily limited to those elements, but may include other elements not expressly listed or that are inherent in such a method or apparatus.

[0014] Furthermore, unless explicitly stated to the contrary, "or" refers to an inclusive "or" and not an exclusive "or." For example, a condition A or B is satisfied by one of the following conditions: A is true (or present) and B is false (or absent), A is false (or absent) and B is true (or present), and both A and B are true (or present).

[0015] The terms "a" or "an" as used herein are defined as "one or more." The terms "another" and "another," and any other variations thereof, are defined as "at least one other."

[0016] The term “plurality” as used here shall mean “two or more”.

[0017] The term “configured” or “set up” to fulfil a specific function (and respective modifications thereof) is to be understood within the meaning of the invention that the corresponding device is already in a design or setting in which it can carry out the function or is at least adjustable - i.e. configurable - so that it can carry out the function after being set accordingly. The configuration can be carried out, for example, by appropriately setting parameters of a process sequence or of switches or the like for activating or deactivating functionalities or settings. In particular, the device can have a plurality of predetermined configurations or operating modes, so that the configuration can be carried out by selecting one of these configurations or operating modes.

[0018] Preferred embodiments of the method are described below, which can be combined with each other as well as with the other aspects of the invention described, unless this is expressly excluded or is technically impossible.

[0019] In some embodiments, the method further comprises determining a position and / or orientation of the mobile terminal in the vehicle interior by means of at least one monitoring device or by means of the image capture device, wherein the image data are only combined if the position and / or orientation of the mobile terminal at the first point in time differs from that at the second point in time. If the position and / or orientation of the terminal in space changes, the corresponding view captured by the same image capture device also changes. As briefly mentioned above, it is advantageous if the views, i.e. the image data that are combined, are different. In this way, deficiencies or errors in the individual views can be compensated for by corresponding parts in the other view(s).If merging is performed only when the views are actually different, the efficiency of the process can be improved, since merging (essentially) identical views does not provide any benefit.

[0020] In some embodiments, the first and second image data are part of an image sequence that is captured by the image capture device, wherein the first and second image data are selected from the image sequence. However, the second image data are only selected if the position and / or orientation of the mobile terminal device differs at the first point in time from that at the second point in time. As just described, different views are to be combined to obtain an overall view from which the content can be reliably recognized. For this purpose, an image sequence, such as a video stream, can be captured, wherein individual images from the image sequence are selected for combination when a difference from previous views arises. For this purpose, the position and / or orientation of the terminal device in space can advantageously be tracked, for example using a monitoring device such as a time-of-flight camera.Of course, this can also be determined using simple images from a normal camera.

[0021] In some embodiments, the position and / or orientation of the mobile device comprises an angle between the mobile device and the image capture device. In particular, the position and / or orientation of the mobile device can be determined to be different at the first and second points in time if a difference in the angle exceeds a predetermined threshold. In particular for mirroring or reflections, an angle of the device, more precisely its display relative to the camera (image capture device), is more crucial than, for example, the distance to the camera. Movements that change the angle are therefore particularly suitable for changing reflections, so that an overall view can be compiled without deficiencies, from which displayed content can be reliably recognized.

[0022] In some embodiments, the method further comprises determining whether a function of the vehicle can be assigned to the recognized content. If no function can be assigned to the content, further image data is selected from the image sequence at at least one further point in time following the first and second points in time and combined with the other image data. The recognition of the content can in particular serve to control a function of the vehicle, as explained in more detail below. Therefore, content can be recognized as "valid" if a corresponding vehicle function can be assigned to it. An image sequence with different views can now be recorded until enough different image data have been collected from which an overall view can be created, from which valid content can then be recognized. During this "iteration", the user can move their device accordingly.If the natural movement or a movement made by the user on his or her own initiative is not sufficient, it may also be possible to provide a corresponding indication to the user.

[0023] In some embodiments, the method further comprises the following prior to capturing the image data by means of the image capturing device(s): capturing monitoring data for at least part of the interior of the vehicle by means of at least one monitoring device and determining (i.e. checking) whether the user's terminal device is at least partially located in a predetermined area of ​​the monitoring data, wherein the capturing of the image data by means of the image capturing device(s) is only carried out if it is determined that the user's terminal device is at least partially located in the predetermined section.

[0024] This measure, which precedes the actual recognition of content on the end device, allows the process to be made more reliable. Image analysis is not carried out continuously, which would require additional computing power and could also lead to undesirable results, for example if a user does not actually intend to control a vehicle function. The image data for the possible recognition of content on the end device is only recorded if a user actually holds the device within a detection range (the specified range). A monitoring device that may already be present in the vehicle can be used for this purpose, e.g. to monitor the driver for other purposes.However, the monitoring data as such may not be suitable for detecting content displayed on a device, which is why, if it is positively determined that a device is located within the specified area, image data is captured as explained above. The monitoring device may nevertheless also be an image capture device, as explained in more detail below.

[0025] In some embodiments, the method further comprises determining whether the mobile device, when in use, is at least partially located in a predetermined area of ​​the part of the interior that lies within a detection range of an image capture device (see description above). It may thus be provided to not detect every use of a mobile device, for example when a user uses their smartphone solely for themselves, e.g., typing a message or making a call. The predetermined area of ​​the part of the interior may be an area in front of an interior camera, which may be located in the interior mirror. A user must therefore then actively hold their device in front of the camera for use to be detected.

[0026] In some embodiments, before recognizing content on the user's terminal device by evaluating the image data, the method further comprises the following: determining a section of the image data that represents the terminal device from the monitoring data (in particular the predetermined area described above), and determining whether the section contains image data that is suitable for recognizing content on the user's terminal device, wherein the recognition of content on the terminal device is only carried out if it is determined that the section contains image data that is suitable for recognizing content on the terminal device. In other words, the monitoring data can be used to determine the section of the image data that is relevant for recognition. In this case, it can be determined based on coordinates where, for example, a display of the terminal device is located.For example, the image data may not be suitable for recognizing content if the device's display is not switched on. This can be determined, for example, by brightness differences in the image data. The lighting conditions on the device may also be unsuitable for recognizing content, for example, because the display appears too dark. This can be done, for example, by determining pixels and their saturation, for example, if the number of pixels with a certain (minimum saturation) is less than a specified value.

[0027] In some embodiments, when combining the image data, (only those) image portions from the first and second image data whose saturation exceeds a threshold value are taken into account. The saturation can be used, in particular, as an indicator of mirroring or reflections on a display, since these areas appear locally strongly overexposed or white in the image data. These portions can then be removed by replacing them with corresponding image regions from other views that do not contain this mirroring (or contain it elsewhere) when combining the image data. A corresponding histogram can be used to determine the saturation, which can be done pixel by pixel.

[0028] In some embodiments, the image data is captured by the image capture device in the visible light range. This can be done using an RGB camera. This method is most suitable for recognizing content appropriated on a terminal device's display.

[0029] In some embodiments, the monitoring data is acquired by the monitoring device, which is configured to acquire monitoring data for at least part of the vehicle interior, wherein the monitoring data preferably comprises image data in the infrared light range (IR) and / or three-dimensional model data. Driver monitoring systems typically operate with such IR cameras. This method is suitable for monitoring a vehicle interior because image data acquired in the IR range is robust to different lighting conditions. In other words, this method delivers reliable results both in strong sunlight and in the dark (especially compared to image data in the visible light range). Data on a three-dimensional model, recorded, for example, with a time-of-flight camera, allows a detailed three-dimensional view of the vehicle interior.

[0030] In some embodiments, recognizing content comprises recognizing text and / or recognizing a machine-readable code. Such recognition can utilize known methods, such as text recognition from image data using OCR. Classification or parsing can then be performed to assign the recognized text (or code) to a vehicle function. For example, a recognized text can be used to determine whether it is a title of a piece of music or an address. Depending on this, an app from a streaming provider can then be called up accordingly, for example, to play the recognized piece of music, or the navigation system can be started with appropriate route guidance to the recognized address.

[0031] In some embodiments, the method comprises assigning the recognized content to a function of the vehicle and controlling the function of the vehicle depending on the recognized content. The recognized content is assigned to a function of the vehicle, in other words, the type of content is recognized so that the content can be assigned to a vehicle function, so that this function can then be controlled, in particular called up or configured, depending on the recognized content, for example with the recognized content. This procedure significantly simplifies the control of a vehicle function because no coupling, connection, or other electronic connection of the terminal device to the vehicle is necessary. Therefore, a vehicle function can also be controlled, for example, by users who do not have to connect their terminal device to the vehicle beforehand. The control can, for example,This can be done simply by holding a smartphone up to a camera in the interior of the vehicle. From the image captured by the smartphone, or more precisely the display with the content shown thereon, an associated vehicle function with the appropriate content can be called up quickly and easily. Due to the improved image recognition by combining individual images as explained above, the control of vehicle functions can be carried out particularly reliably. A second aspect of the invention relates to a system for data processing, comprising at least one processor which is configured to carry out the method according to the first aspect of the invention. The system has, in particular, at least one lighting device in the interior of a vehicle.

[0032] In some embodiments, the system further comprises at least one image capture device configured to capture image data in the visible range (RGB) of light. As explained above, content of a terminal device, in particular content displayed on a display of a terminal device, can be recognized in this way.

[0033] In some embodiments, the system further comprises at least one monitoring device configured to acquire monitoring data for at least part of the vehicle's interior, particularly in the form of image data in the infrared (IR) range and / or three-dimensional model data. As explained above, it is advantageous to monitor the interior of a vehicle using an infrared camera or a time-of-flight camera.

[0034] A third aspect of the invention relates to a computer program comprising instructions which, when executed on a system according to the second aspect, cause the system to carry out the method according to the first aspect.

[0035] The computer program can, in particular, be stored on a non-volatile data carrier. This is preferably a data carrier in the form of an optical data carrier or a flash memory module. This can be advantageous if the computer program as such is to be handled independently of a processor platform on which the one or more programs are to be executed. In another implementation, the computer program can be present as a file on a data processing unit, in particular on a server, and can be downloadable via a data connection, for example the Internet or a dedicated data connection, such as a proprietary or local network. Furthermore, the computer program can have a plurality of interacting individual program modules.

[0036] The system according to the second aspect can accordingly comprise a program memory in which the computer program is stored. Alternatively, the system can also be configured to access an external computer program, for example, available on one or more servers or other data processing units, via a communication connection, in particular to exchange data with it that is used during the execution of the method or computer program or that represents outputs of the computer program.

[0037] The features and advantages explained with reference to the first aspect of the invention also apply accordingly to the further aspects of the invention. This also applies to the features and advantages explained with reference to the second aspect of the invention.

[0038] Further advantages, features and possible applications of the present invention will become apparent from the following detailed description in conjunction with the drawings.

[0039] It shows:

[0040] Fig. 1 is a flowchart of a method according to an embodiment;

[0041] Fig. 2 is a view of a vehicle interior with a terminal held by a user in a detection area; and

[0042] Fig. 3 a view of a terminal with displayed content.

[0043] Throughout the figures, the same reference numerals are used for the same or corresponding elements of the invention.

[0044] Fig. 1 shows a method 100 for detecting a mobile terminal 2 in a vehicle, which in the example shown leads to the control of a vehicle function. The method 100 can be carried out in particular in a data processing system of the vehicle (not shown). The method 100 is explained below with reference to Fig. 2 and Fig. 3. Fig. 2 shows a section of a vehicle interior 1 with a driver as user 10 of a smartphone as an exemplary mobile terminal 2. Fig. 3 shows a smartphone 2 in the hand of a user with content 22 displayed on the display 21. Reflections 23 can obstruct the view of the displayed content 22. However, as explained below, these can be eliminated using the method 100, in particular by recording an image sequence and combining image data.

[0045] In a step S1, monitoring data for at least part of the vehicle interior 1 can be recorded. The vehicle can be equipped with a driver monitoring system (DMS), which monitors the driver and possibly also other parts of the vehicle interior 1. This monitoring can be active permanently while driving (i.e., in particular, beginning and ending with the ignition being switched on or off). For this purpose, a monitoring device 3 can be arranged in the vehicle interior 2, for example, in the base 5 of the interior mirror 6. From this position, the driver 10 and large parts of the vehicle interior 1 can be clearly seen.

[0046] An infrared (IR) camera (e.g., a near-infrared (NIR) camera) can be used. An IR image is well-suited for surveillance because it is robust against changing light conditions, for example, even in strong sunlight. Recordings in the dark are also possible. It goes without saying that a suitable IR light source may be present. It is also conceivable to create a three-dimensional model of the driver 10 or the vehicle interior 1, for example, using a time-of-flight (ToF) camera. Using the three-dimensional model, distances to the camera, and thus distances between objects and their movements, can be calculated.

[0047] From the recorded monitoring data, it can be determined whether a mobile terminal 2 is detected that is currently in use. In principle, this can include any use of the terminal 2 that a user 10 can make in the vehicle, without the terminal 2 having to be held in a special area (see explanations below), for example. It may be sufficient if the user 10 picks up their terminal 2 or the like. For security or convenience reasons, for the purpose of detecting use, it can also be determined, for example, who is using the terminal 2, in particular to distinguish between driver and passenger, and / or whether the vehicle is stationary or moving.For example, it can be provided that use by the driver is only recognized as such when the vehicle is stationary, or that use while driving is only recognized when the terminal device 2 is held in the area 7 described below, which is indicated in Fig. 2 by dashed lines. This area 7 is an excerpt from the monitoring data; in the case of (IR) image data, this can simply be a two-dimensional area that has been determined to be suitable for recognizing content displayed on a terminal device. In the vehicle interior, this can correspond to an area near camera 3 (monitoring device) (or camera 4 (image capture device), see below). The arrangement in the base 5 of the interior mirror 6 at the top center of the vehicle interior 1 is also easily accessible for a user to position a terminal device 2 for recognition in the area 7 in a suitable manner. It is understood that the area shown in Fig.The area 7 indicated in Figure 2 is merely exemplary. The position, size, and shape of area 7 can be adjusted according to desired requirements.

[0048] The following steps of method 100 are advantageously only carried out when terminal device 2 enters area 7. On the one hand, the method can be carried out more reliably, since no evaluation of all surveillance data takes place, even if terminal device 2 is, for example, not in the field of view of camera 3 (or camera 4, see below).

[0049] Before image data is captured for evaluation in step S2, it can first be checked whether the terminal device 2 is suitably positioned in front of the camera 4. For example, for reliable recognition of displayed content, the terminal device 2 should be held in an orientation that is neither too tilted nor too twisted. Furthermore, the terminal device 2 should not be obscured, for example, by the hand of the user 10, which may be an indication that the display 21 of the terminal device 2 is not in the field of vision or at least that large parts of the display 21 are obscured by fingers or parts of the hand of the user 10. Furthermore, excessively rapid movement of the terminal device 2 may result in the content 22 not being recognized. Firstly, recognition of content 22 works more reliably if the terminal device 2 is not moved too much.On the other hand, a rapid movement of the terminal device 2 in area 7 may indicate that it has unintentionally entered area 7. Any reflections 23 on the display 21 can then be removed as explained below.

[0050] The image capture device 4 can, in particular, be a camera 4 that captures image data in the visible light range and can therefore be referred to as an RGB camera. While an IR camera or ToF camera as a monitoring device 3 delivers a suitable image under very different lighting conditions, the monitoring data 3 then does not contain any (RGB) colors. However, these are necessary for the recognition of content 22 displayed (in the visible light range) on a terminal device 2, such as on a display 21 of a smartphone (see Fig. 3). In an IR image, a terminal device 2, such as a smartphone, can only be recognized as a (usually) essentially rectangular object, on which the display 21 can at best be distinguished as a rectangular monochrome surface.

[0051] While in Fig. 2 the monitoring device 3 and the image capture device 4 are shown as separate units, they can also be combined in one unit, which in particular can be an RGB-IR camera, ie a camera which delivers image data in both the RGB and IR ranges.

[0052] The image capture device 4 can, in particular, also be arranged in the base 5 of the interior mirror 6. Therefore, the viewing areas of the monitoring device 3 and the image capture device 4 can be assumed to be largely congruent, at least in area 7 (or even substantially congruent if the monitoring device 3 and the image capture device 4 are formed as a single unit). Therefore, image data located in area 7 is particularly captured. To expand the coverage of the interior, it can also be provided to combine the image data from several cameras.

[0053] In step S2, image data is captured, in particular, as individual images of an image sequence. For example, a corresponding video stream can be recorded with camera 4, with the individual frames representing the individual images and providing the image data for method 100. Recording of the image sequence can start when user 10 holds their terminal device in area 7, which—in other words—means that terminal device 2 becomes visible in the camera image.

[0054] Recording can then be continued until the displayed content can be recognized (in particular by combining the image data of the individual images). By executing the hand movement (see arrows in Fig. 2) to align the end device 2 relative to the camera 4 in the interior 1, the user 10 moves and rotates the possibly reflective surface of the display 21 to such an extent that the reflections are located at different points in the respective views at different times in the video sequence. For this purpose, the information regarding how the end device is oriented in space can also be used to easily select individual images of the sequence (step S3) in which the reflections occur at different points. By combining the selected recordings (step S4), the reflections 23 can now be removed from the image content, since the displayed content 21 on the display 21 does not change.The reflections 23 can preferably be completely removed or at least reduced to such an extent that a valid content 21 can be recognized, which can subsequently be used to control a vehicle function.

[0055] The image data is now evaluated using image analysis. First, content 22 displayed on the terminal device 2 is recognized (step S5). In particular, known methods for optical character recognition (OCR) can be used for this purpose if the content is text, or for reading machine-readable codes, for example if the content is a code such as a QR code. The recognized content can then be assigned to a vehicle function in a step S6, for example using known parsing or classification methods. For example, a recognized address can be assigned to the vehicle function "navigation system," a recognized piece of music can be assigned to the vehicle function "music player," etc. For this purpose, other optical features such as formatting, color, etc., for example to call up a specific streaming provider, or even data from a code can be used for the assignment.

[0056] In a step S7, the correspondingly assigned vehicle function is then controlled depending on the recognized content. Controlling a vehicle function can in particular comprise calling or starting a vehicle function assigned to a recognized content, in particular depending on the recognized content, i.e. the vehicle function is configured with the recognized content. For example, if the content 22 is an address, the vehicle function 9 "navigation system" can be called directly with the address as the set destination address, thus eliminating the need to manually enter the address in the navigation system in the vehicle. If the recognized piece of music is a song, a streaming provider with the recognized piece of music can be called directly, etc.

[0057] However, controlling a vehicle function can also simply comprise displaying data associated with recognized content, for example on a screen 8 of the vehicle, in particular a display of the infotainment system. Then, if necessary, provision can be made to offer a query for confirmation by the user, for example, whether the recognized content is correct and / or whether a function associated with the recognized content should be executed. For example: "The following address was recognized. Should route guidance be started?", "The following piece of music was recognized. Should it be played?", etc.

[0058] For example, the process might work like this in practice: A user opens an app on their smartphone, for example to display a QR code. The user turns the smartphone and moves it towards the interior camera. If the smartphone is visible in the image from the interior camera, recording starts and saves images. At the same time, the camera software estimates the corner points of the smartphone display and calculates the position of the smartphone in 3D space. Optionally, the coordinates of the hand holding the smartphone with the reflective display surface can be calculated. If the position of the smartphone display changes significantly, especially the planar angles between the interior camera and the smartphone display (i.e. more than a predefined threshold), the relevant image (from the image sequence) is selected.A complete image is finally calculated from the selected images, primarily by blending and stitching the individual images together (sometimes referred to as "stitching"). The combined image is used to recognize the complete QR code, ultimately controlling a corresponding vehicle function.

[0059] While at least one exemplary embodiment has been described above, it should be noted that a large number of variations exist. It should also be noted that the described exemplary embodiments are only non-limiting examples and are not intended to limit the scope, applicability, or configuration of the devices and methods described herein. Rather, the foregoing description will provide a guide to implementing at least one exemplary embodiment, with the understanding that various changes in the operation and arrangement of the elements described in an exemplary embodiment may be made without departing from the subject matter defined in the appended claims, as well as their legal equivalents.

[0060] 100 Methods for detecting a mobile device in a vehicle

[0061] 1 Vehicle interior 2 End device (smartphone)

[0062] 3 Monitoring device (IR camera)

[0063] 4 Image capture device (RGB camera)

[0064] 5 Interior mirror base

[0065] 6 Interior mirror 7 Area

[0066] 8-display infotainment system

[0067] 9 Vehicle function

[0068] 10 users (drivers)

[0069] 21 Display device (smartphone) 22 Displayed content

[0070] 23 Reflections

Claims

CLAIMS 1. A method (100) for detecting a mobile terminal (2) in a vehicle, the method comprising: - capturing first image data by means of at least one image capturing device (4) in the interior (1) of a vehicle, wherein the first image data contain at least a first view of a terminal (2) of the user (10) at a first point in time; - capturing second image data by means of the image capturing device (4) and / or another image capturing device in the interior (1) of the vehicle, wherein the second image data contain at least a second view of the terminal (2) of the user (10) at a second point in time; - combining the first and second image data to obtain a composite view of the terminal (2) of the user (10); and - Recognizing a content (22) displayed on the terminal (2) of the user (10) by evaluating the combined image data.

2. The method of claim 1, further comprising: - Determining a position and / or orientation of the mobile terminal (2) in the vehicle interior (1) by means of at least one monitoring device (3) or by means of the image capture device (4), wherein the image data are only combined if the position and / or orientation of the mobile terminal (2) at the first time differs from that at the second time.

3. The method according to claim 2, wherein the first and second image data are part of an image sequence which is captured by the image capture device (4), the method further comprising: - selecting the first and second image data from the image sequence, wherein the second image data are only selected if the position and / or orientation of the mobile terminal (2) at the first time differs from that at the second time.

4. The method according to claim 2 or 3, wherein the position and / or orientation of the mobile terminal comprises an angle between the mobile terminal (2) and the image capture device (4).

5. The method according to claim 4, wherein the position and / or orientation of the mobile terminal (2) are determined to be different at the first and second times if a difference in angle exceeds a predetermined threshold value.

6. The method according to any one of claims 2 to 5, further comprising: - determining whether a function (9) of the vehicle can be assigned to the recognized content (22), wherein, if no function (9) can be assigned to the content (22), further image data are selected from the image sequence at at least one further time point following the first and second time points and are combined with the other image data.

7. A method according to any one of the preceding claims, wherein the method further comprises, prior to capturing the image data: - Acquiring monitoring data for at least part of the interior (1) of the vehicle by means of at least one monitoring device (3) or by means of the image acquisition device (4); - determining whether the terminal (2) of the user (10) is at least partially located in a predetermined area (7) of the monitoring data; wherein the acquisition of the image data is only carried out if it is determined that the terminal (2) of the user (10) is at least partially located in the predetermined area (7).

8. The method according to any one of the preceding claims, further comprising: - Assigning the recognized content (22) to a function (9) of the vehicle; and - Controlling the function (9) of the vehicle depending on the detected content (22).

9. Method according to one of the preceding claims, wherein, when combining the image data, image components from the first and second image data are taken into account whose saturation exceeds a threshold value.

10. Method according to one of the preceding claims, wherein the recognition of a content (22) comprises the recognition of a text and / or the recognition of a machine-readable code.

11. A data processing system comprising at least one processor configured to carry out the method according to any one of the preceding claims.

12. A computer program comprising instructions which, when executed on a system according to claim 11, cause the system to carry out the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Information entry via in-vehicle camera

    US20170213098A1

  • Adaptive camera control for reducing motion blur during real-time image capture

    US20200374455A1