Method for detecting a mobile terminal in a vehicle

The method uses an image capture device and machine learning to analyze the layout of mobile terminal content, improving recognition accuracy and enabling direct control of vehicle functions by identifying relevant parameters without user interaction.

WO2025149195A1PCT designated stage expired Publication Date: 2025-07-17BAYERISCHE MOTOREN WERKE AG
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/079321
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-08
Filing Date
2024-10-17
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing methods for controlling vehicle functions using a mobile terminal are cumbersome and prone to errors due to difficulties in accurately recognizing and assigning relevant content from the terminal's display to the appropriate vehicle function, especially when text recognition produces ambiguous results or includes irrelevant information.

Method used

A method utilizing an image capture device and a trained machine learning model to analyze the design of content displayed on a mobile terminal, determining the associated vehicle function based on the layout, and controlling the function without the need for direct user interaction or device connection.

Benefits of technology

Enhances the reliability of content recognition by accurately assigning displayed content to vehicle functions, allowing seamless control of vehicle systems such as navigation or media playback directly from the mobile terminal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024079321_17072025_PF_FP_ABST
    Figure EP2024079321_17072025_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method for detecting a mobile terminal (2) in a vehicle. According to the method, image data is captured by means of at least one image capturing device (4) in the interior (1) of a vehicle, said image data including at least one view of a terminal (2), and contents displayed on the terminal (2) are detected. In order to detect the contents, the presentation (24) of the displayed contents is detected by analyzing the image data, and a vehicle function (9) associated with the displayed contents (22) is determined by means of a trained machine learning model using the detected presentation (24). Additionally, a function parameter in the displayed contents (22) is determined. The invention also relates to a method for training the machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Method for detecting a mobile terminal in a vehicle

[0002] The present invention relates to a method for detecting a mobile terminal in a vehicle. In particular, content displayed on a mobile terminal is to be captured using an image capture device in order to control a function of the vehicle.

[0003] Modern vehicles have a multitude of functions that can be controlled by a user, i.e., accessed, started, stopped, configured, etc. This can include, in particular, infotainment system functions such as entering a navigation destination, playing a piece of music, calling a contact, or the like. These infotainment system functions are typically controlled by the user interacting directly with the vehicle, for example, by pressing buttons or touching a touchscreen, or by voice control or gesture control.

[0004] However, operating the system in the vehicle itself can sometimes be cumbersome, for example when data such as a navigation destination or the title of a song has to be entered using a keyboard on a touchscreen. Such input is often easier on a user's device, such as a smartphone. In some cases, the data is even already available on the smartphone, for example if a user has already listened to a particular song on their smartphone and now wants to hear it in the vehicle. Therefore, methods are known in which data is transferred from the user's device to the vehicle, for example via Bluetooth or a cloud service.

[0005] It is also known to capture content shown on a mobile device display using a camera in the vehicle. Since modern vehicles are often equipped with corresponding cameras, for example for interior monitoring, this infrastructure can be used to read displayed content. The content can be displayed text that is recognized using OCR. This way, for example, navigation destinations can be easily transferred from a smartphone to the vehicle. Machine-readable codes such as QR codes can also be used and captured by a camera in the vehicle. Even if displayed text is correctly recognized using text recognition, it can still be difficult to correctly identify or assign the actually relevant text.For example, displayed content may contain various texts or words, only some of which are relevant for correct control of an intended function. For example, a song title must be distinguished from other text, such as descriptions or labels in an app, in order to actually control a music function. It may also happen that, for example, several points of interest (POI) are recognized on a map using text recognition, but only one of them is of interest to the user. In this case, all of the recognized text could be processed and the user then presented with a list. However, this is not desirable, particularly from the user's perspective. In other cases, text recognition may lead to confusion or ambiguous results, for example if a song title is recognized that is also a name or designates a location.Then the result of the text recognition cannot be clearly assigned to a function, e.g. "media" or "navigation", i.e. a vehicle function cannot be clearly controlled, or at least not without user intervention.

[0006] The present invention is based on the object of providing an improved method for recognizing a mobile device in a vehicle. In particular, the recognition of content displayed on a mobile device is to be improved.

[0007] This object is achieved according to the teaching of the independent claims. Various embodiments and further developments of the invention are the subject of the dependent claims.

[0008] A first aspect of the invention relates to a method, in particular a computer-implemented method, for recognizing a mobile terminal in a vehicle. In the method, image data is captured by means of at least one image capture device in the interior of a vehicle, wherein the image data contain at least one view of a terminal. Content displayed on the terminal is recognized. To recognize the content, a design of the displayed content is captured by evaluating the image data, and a function of the vehicle associated with the displayed content is determined on the basis of the captured design using a trained machine learning model. In addition, a functional parameter in the displayed content is determined. The aforementioned method according to the first aspect is therefore based in particular on evaluating image data captured by means of an image capture device, such as a camera.A trained machine learning model is used to recognize the displayed content. In particular, the model can be used to assign a design, i.e. a layout of the displayed content, to a function of the vehicle. The model is trained with corresponding training data before use, as will be explained in more detail later. The method is thus able to recognize, based on the design of the displayed content, which vehicle function this content can be assigned to. For example, it can be recognized whether the displayed content is a view of a music app on the end device. In this case, a song title can be recognized as a function parameter, and a corresponding media or entertainment application in the vehicle is determined as associated.However, if the machine learning model detects that the device is a map display or a navigation app based on its design, this can be assigned to a navigation function in the vehicle, and a location, in particular, can be determined as a navigation destination as a functional parameter. This can increase the reliability of content recognition, as the system can assign or classify the recognized content to a vehicle function (also referred to as a "domain") based on its layout using the trained machine learning model.

[0009] The term “vehicle” as used herein refers in particular to a passenger car, including all types of motor vehicles, hybrid and battery-powered electric vehicles, as well as vehicles such as vans, buses, trucks, delivery vans and the like.

[0010] The term "image capture device" used here refers in particular to a camera, especially a digital camera. The camera can capture still images (photos) or moving images (videos), particularly in the visible light range. An image capture device can capture or record such an image and output corresponding image data. The image capture device can also serve as a monitoring device, i.e., the image capture device and the monitoring device can be one and the same device.

[0011] The term "mobile device" used here refers in particular to electronic devices that are portable and capable of communicating wirelessly with other devices or networks. These are handy devices that can often be carried in a pocket or held in the hand. Mobile devices include smartphones, tablets, laptops, smartwatches, and other similar devices. The mobile device is particularly configured to display content, in particular on a screen. The user of a device can in particular be the driver of the vehicle. However, it is understood that a passenger and, in principle, any vehicle occupant can also be a user of the device within the meaning of the invention.

[0012] The term “content” of a terminal device, as used herein, is to be understood in particular as information presented on a display device, such as a display, of a terminal device. The content does not necessarily have to be displayed on the terminal device in an app that corresponds in type to the vehicle function. The content can simply consist of, for example, unformatted text or machine-readable code. Nevertheless, the content is examined in particular with regard to its design (i.e. layout, arrangement, presentation), which can, for example, also include font format, colors, shapes, logos or the like in order to identify the associated vehicle function. In particular, displayed content also includes a function parameter for the associated function of the vehicle, i.e. a parameter with which the function is called or executed. This can, for example,a song title played in a media application or a location as a navigation destination.

[0013] The term "monitoring device" used here refers in particular to a device suitable for monitoring the interior of a vehicle, in particular the vehicle occupants, and in particular their positions and movements within the space. This can be an image capture device. An infrared (IR) camera (e.g., a near-infrared (NIR) camera) can be used. An IR image is well suited for monitoring because it is robust against changing light conditions, for example, even in strong sunlight. Recordings are also possible in the dark. It goes without saying that a suitable IR light source can be present. It is also conceivable to create a three-dimensional model of the driver or the vehicle interior, for example, using a time-of-flight (TOF) camera.Using the three-dimensional model, distances to the camera, and thus distances between objects and their movements, can be calculated. The term "function of a vehicle" (synonymous with the term "vehicle function"), as used herein, refers in particular to a function of a vehicle's infotainment system, such as "navigation" or "entertainment." This can include proprietary applications of the vehicle manufacturer and / or applications implemented by third-party providers. The control of a vehicle function, also described here, can in particular include calling or starting a vehicle function associated with a recognized content depending on the recognized content, i.e., with a specific function parameter, e.g., starting navigation to a specific destination, playing a specific song with the entertainment application, etc.However, controlling a vehicle function may also merely comprise displaying data relating to a recognized content or the associated vehicle function (possibly together with the recognized content), for example on a screen of the vehicle, in particular a display of the infotainment system.

[0014] The terms "comprises," "includes," "includes," "has," "has," "with," or any other variation thereof, as used herein, are intended to cover non-exclusive inclusion. For example, a method or apparatus that includes or has a list of elements is not necessarily limited to those elements, but may include other elements not expressly listed or that are inherent in such a method or apparatus.

[0015] Furthermore, unless explicitly stated to the contrary, "or" refers to an inclusive "or" and not an exclusive "or." For example, a condition A or B is satisfied by one of the following conditions: A is true (or present) and B is false (or absent), A is false (or absent) and B is true (or present), and both A and B are true (or present).

[0016] The terms "a" or "an" as used herein are defined as "one or more." The terms "another" and "another," and any other variations thereof, are defined as "at least one other."

[0017] The term "plurality" as used here is to be understood in the sense of "two or more". The term "configured" or "set up" to perform a specific function (and respective variations thereof) is to be understood, within the meaning of the invention, that the corresponding device is already in a design or setting in which it can perform the function or is at least adjustable - i.e. configurable - so that it can perform the function after being set accordingly. The configuration can be carried out, for example, by appropriately setting parameters of a process sequence or of switches or the like for activating or deactivating functionalities or settings. In particular, the device can have a plurality of predetermined configurations or operating modes, so that configuration can be carried out by selecting one of these configurations or operating modes.

[0018] Preferred embodiments of the method are described below, which can be combined with each other as well as with the other aspects of the invention described, unless this is expressly excluded or is technically impossible.

[0019] In some embodiments, the function parameter is recognized in a sub-area of ​​the displayed content, which is determined based on the detected design and depending on the specific function. In this way, the function parameter, such as text or a code, can be recognized in a targeted manner, since the design and the associated function indicate the location on the displayed content at which the function parameter is contained in this layout. For example, a song title can be easily distinguished from other descriptions that may also be visible. Using navigation as an example, a navigation destination can be easily filtered out from other locations, since the sub-area in which the navigation destination is actually located in the specific view is known and not, for example, another location.By limiting recognition to the sub-area, relevant text can be easily distinguished from irrelevant text (when text is recognized as a function parameter). Furthermore, recognition can be performed more efficiently and quickly, as the entire displayed content does not have to be searched for the function parameter and, if necessary, identified.

[0020] In some embodiments, an anchor point is further determined in the captured design of the displayed content, wherein the anchor point defines the sub-area of ​​the displayed content in which the functional parameter can be recognized. Such an anchor point can be easily determined using the trained machine learning model. Starting from the anchor point, a bounding box of the sub-area in which the functional parameter can be recognized can then be defined. The system can use the machine learning model to learn where in which design and for which function such an anchor point can be found.

[0021] In some related embodiments, the anchor point is thus a position in the design of the displayed content, which is determined using the machine learning model. For example, the model may have acquired knowledge about where the song title can be found in a music app, for example, under a cover, whereby other text, such as labels, can be ignored.

[0022] In some embodiments, determining the anchor point comprises detecting a marker in the captured design, wherein the detected marker is determined as the anchor point. For example, in a particular design of displayed content, a specific symbol or icon may indicate the location where the functional parameter can be found. The machine learning model may then be particularly capable of finding this marker in a design, even if its position is flexible, such as a POI on a map that is marked with a specific symbol. In this way, the correct functional parameter can be detected with increased reliability, even with flexible designs.

[0023] In some embodiments, detecting the design includes detecting a pattern of the displayed content and comparing the pattern with known patterns. The aforementioned marking can also be detected by pattern matching. Pattern recognition can simplify the assignment of the detected design to a specific vehicle function, particularly using the machine learning model. For example, a design can have certain features of an arrangement of content (i.e., a pattern), which can then be easily assigned to a function, i.e., a domain.

[0024] In some embodiments, the method comprises controlling the function of the vehicle depending on the recognized functional parameter. The recognized content is assigned to a function of the vehicle by recognizing the design using a machine learning model, as described, so that this function can then be called up or configured depending on the recognized content, more precisely the recognized functional parameter. This procedure significantly simplifies the control of a vehicle function, since no coupling, connection, or other electronic connection of the end device to the vehicle is necessary. Therefore, a vehicle function can also be controlled, for example, by users who do not have to connect their end device to the vehicle beforehand. The control can be carried out, for example, simply by holding a smartphone up to a camera in the interior of the vehicle. From the image recorded by the smartphone, orMore precisely, the display with the content shown on it allows a corresponding vehicle function with the appropriate content to be called up quickly and easily.

[0025] In some embodiments, the function parameter comprises text that is recognized using optical character recognition. Alternatively, the function parameter can also comprise a machine-readable code. Such recognition can rely on known methods, for example, optical character recognition from the image data.

[0026] In some embodiments, the method further comprises the following prior to capturing the image data by means of the image capturing device: capturing monitoring data for at least part of the interior of the vehicle by means of at least one monitoring device and determining (i.e. checking) whether the user's terminal is at least partially located in a predetermined area of ​​the monitoring data, wherein the capturing of the image data by means of the image capturing device is only carried out if it is determined that the user's terminal is at least partially located in the predetermined section.

[0027] This measure, which can precede the actual recognition of content on the end device, allows the process to be designed with increased reliability. Image analysis is not carried out continuously, which would require additional computing power and could also lead to undesirable results, for example if a user does not actually intend to control a vehicle function. The image data for the possible recognition of content on the end device is only recorded if a user actually holds the device within a detection range (the specified range). A monitoring device that may already be present in the vehicle can be used for this purpose, e.g. to monitor the driver for other purposes.However, the monitoring data as such may not be suitable for detecting content displayed on a device, which is why, if it is positively determined that a device is located within the specified area, image data is captured as explained above. The monitoring device may nevertheless also be an image capture device, as explained in more detail below.

[0028] In some embodiments, the method further comprises determining whether the mobile device, when in use, is at least partially located in a predetermined area of ​​the part of the interior that lies within a detection range of an image capture device (see description above). It may thus be provided to not detect every use of a mobile device, for example when a user uses their smartphone solely for themselves, e.g., typing a message or making a call. The predetermined area of ​​the part of the interior may be an area in front of an interior camera, which may be located in the interior mirror. A user must therefore then actively hold their device in front of the camera for use to be detected.

[0029] In some embodiments, before recognizing content on the user's terminal device by evaluating the image data, the method further comprises the following: determining a section of the image data representing the terminal device from the monitoring data (in particular the predetermined area described above), and determining whether the section contains image data suitable for recognizing content on the user's terminal device, wherein the recognition of content on the terminal device is only carried out if it is determined that the section contains image data suitable for recognizing content on the terminal device. In other words, the monitoring data can be used to determine the section of the image data relevant for recognition. For example, coordinates can be used to determine where a display of the terminal device is located.For example, the image data may not be suitable for recognizing content if the device's display is not switched on at all. This can be determined, for example, by brightness differences in the image data. The lighting conditions on the device may also be unsuitable for recognizing content, for example, because the display appears too dark. This can be done, for example, by determining pixels and their saturation, for example, if the number of pixels with a certain (minimum saturation) is less than a specified value.

[0030] In some embodiments, the image data is captured by the image capture device in the visible light range. This can be done using an RGB camera. This method is most suitable for recognizing content appropriated on a terminal device's display.

[0031] In some embodiments, the monitoring data is acquired by the monitoring device, which is configured to acquire monitoring data for at least part of the vehicle interior, wherein the monitoring data preferably comprises image data in the infrared light range (IR) and / or three-dimensional model data. Driver monitoring systems typically operate with such IR cameras. This method is suitable for monitoring a vehicle interior because image data acquired in the IR range is robust to different lighting conditions. In other words, this method delivers reliable results both in strong sunlight and in the dark (especially compared to image data in the visible light range). Data on a three-dimensional model, recorded, for example, with a time-of-flight camera, allows a detailed three-dimensional view of the vehicle interior.

[0032] A second aspect of the invention relates to a method for training a machine learning model for recognizing a mobile terminal in a vehicle, in particular in a method according to the first aspect. In the method according to the second aspect, training data is provided, wherein the training data comprises image data containing at least one view of a terminal, as well as an assignment of the image data to a function of the vehicle. Content displayed on the terminal is recognized (from the training data). For this purpose, a design of the displayed content is detected by evaluating the image data, the function of the vehicle assigned to the displayed content is determined on the basis of the detected design using the machine learning model, and a function parameter in the displayed content is recognized.By including a mapping of the image data to a corresponding vehicle function in the training data, effective training can be achieved. In other words, the image data is labeled according to its associated vehicle function (domain), allowing a training process to be developed based on it. The machine learning model can, for example, be a neural network trained using the training data.

[0033] In some embodiments of the training method, the training data defines a sub-area of ​​the displayed content in which the function parameter can be recognized. In other words, not only is the appropriate domain known for the image data of the training data, but also the location (the sub-area) in the image data where the function parameter, e.g., relevant text, can be found. The sub-area can, for example, be defined at an anchor point in the form of a bounding box, as also described above.

[0034] In some embodiments, the training data contains image data for different views of the terminal device, which are assigned to the same function with the same function parameter. In this way, the model can be better trained and the reliability of the recognition can be increased. In particular, a dependency on an actual view can be reduced, i.e. content is recognized even if the captured design has deviations, e.g. appears larger or smaller or rotated. For example, it can be provided to display content on the terminal device in different views, e.g. zoomed or scrolled, with a different section being visible in each case. In general, the different views can contain any rotations, enlargements, sections and the like.Furthermore, training can also be carried out with training data that represents a view of a specific function, but with variable function parameters, e.g., views of a music app that may be identical in design but with a different song title (as a function parameter), or map views of a navigation app with different addresses as the navigation destination.

[0035] A third aspect of the invention relates to a data processing system comprising at least one processor configured to carry out the method according to the first aspect of the invention.

[0036] In some embodiments, the system comprises at least one image capture device configured to capture image data in the visible range (RGB) of light. As explained above, content of a terminal device, in particular content displayed on a display of a terminal device, can be recognized in this way.

[0037] In some embodiments, the system further comprises at least one monitoring device configured to acquire monitoring data for at least part of the vehicle's interior, particularly in the form of image data in the infrared (IR) range and / or three-dimensional model data. As explained above, it is advantageous to monitor the interior of a vehicle using an infrared camera or a time-of-flight camera.

[0038] A fourth aspect of the invention relates to a computer program comprising instructions which, when executed on a system according to the third aspect, cause the system to carry out the method according to the first aspect.

[0039] The computer program can, in particular, be stored on a non-volatile data carrier. This is preferably a data carrier in the form of an optical data carrier or a flash memory module. This can be advantageous if the computer program as such is to be handled independently of a processor platform on which the one or more programs are to be executed. In another implementation, the computer program can be present as a file on a data processing unit, in particular on a server, and can be downloadable via a data connection, for example the Internet or a dedicated data connection, such as a proprietary or local network. Furthermore, the computer program can have a plurality of interacting individual program modules.

[0040] The system according to the third aspect can accordingly comprise a program memory in which the computer program is stored. Alternatively, the system can also be configured to access an external computer program, for example, available on one or more servers or other data processing units, via a communication connection, in particular to exchange data with it that is used during the execution of the method or computer program or that represents outputs of the computer program.

[0041] The features and advantages explained with reference to the first aspect of the invention also apply accordingly to the further aspects of the invention. This also applies to the features and advantages explained with reference to the second aspect of the invention.

[0042] Further advantages, features and possible applications of the present invention will become apparent from the following detailed description in conjunction with the drawings.

[0043] 1 shows a flowchart of a method according to an embodiment;

[0044] Fig. 2 is a view of a vehicle interior with a terminal held by a user in a detection area; and

[0045] Fig. 3 a view of a terminal with displayed content.

[0046] Throughout the figures, the same reference numerals are used for the same or corresponding elements of the invention.

[0047] Fig. 1 shows a method 100 for detecting a mobile terminal 2 in a vehicle, which in the example shown leads to the control of a vehicle function. The method 100 can be carried out in particular in a data processing system of the vehicle (not shown). The method 100 is explained below with reference to Figs. 2 and 3. Fig. 2 shows a section of a vehicle interior 1 with a driver as user 10 of a smartphone as an exemplary mobile terminal 2. Fig. 3 shows a smartphone 2 in the hand of a user with content 22 displayed on the display 21.

[0048] In a step S1, monitoring data for at least part of the vehicle interior 1 can be recorded. The vehicle can be equipped with a driver monitoring system (DMS), which monitors the driver and possibly also other parts of the vehicle interior 1. This monitoring can be active permanently while driving (i.e., in particular, beginning and ending with the ignition being switched on or off). For this purpose, a monitoring device 3 can be arranged in the vehicle interior 2, for example, in the base 5 of the interior mirror 6. From this position, the driver 10 and large parts of the vehicle interior 1 can be clearly seen.

[0049] An infrared (IR) camera (e.g., a near-infrared (NIR) camera) can be used. An IR image is well-suited for surveillance because it is robust against changing light conditions, for example, even in strong sunlight. Recordings in the dark are also possible. It goes without saying that a suitable IR light source may be present. It is also conceivable to create a three-dimensional model of the driver 10 or the vehicle interior 1, for example, using a time-of-flight (ToF) camera. Using the three-dimensional model, distances to the camera, and thus distances between objects and their movements, can be calculated.

[0050] From the recorded monitoring data, it can be determined whether a mobile terminal 2 is detected that is currently in use. In principle, this can include any use of the terminal 2 that a user 10 can make in the vehicle, without the terminal 2 having to be held in a special area (see explanations below), for example. It may be sufficient if the user 10 picks up their terminal 2 or the like. For security or convenience reasons, for the purpose of detecting use, it can also be determined, for example, who is using the terminal 2, in particular to distinguish between driver and passenger, and / or whether the vehicle is stationary or moving.For example, it may be provided that use by the driver is only recognized as such when the vehicle is stationary, or that use during travel is only recognized when the terminal 2 is held in the area 7 described below, which is indicated in Fig. 2 by dashed lines in Fig. 2.

[0051] This area 7 is an excerpt from the surveillance data; in the case of (IR) image data, this can simply be a two-dimensional area that has been determined to be suitable for recognizing content displayed on a terminal device. In the vehicle interior, this can correspond to an area near camera 3 (monitoring device) (or camera 4 (image capture device), see below). The arrangement in the base 5 of the interior mirror 6 at the top center of the vehicle interior 1 is also easily accessible for a user to position a terminal device 2 for recognition in area 7 in a suitable manner. It is understood that the area 7 indicated in Fig. 2 is merely an example. The position, size, and shape of area 7 can be adjusted according to desired requirements.

[0052] The following steps of method 100 are advantageously only carried out when terminal device 2 enters area 7. On the one hand, the method can be carried out more reliably, since no evaluation of all surveillance data takes place, even if terminal device 2 is, for example, not in the field of view of camera 3 (or camera 4, see below).

[0053] Before image data is captured for evaluation in step S2, it can first be checked whether the terminal device 2 is suitably positioned in front of the camera 4. For example, for reliable recognition of displayed content, the terminal device 2 should be held in an orientation that is neither too tilted nor too twisted. Furthermore, the terminal device 2 should not be obscured, for example, by the hand of the user 10, which may be an indication that the display 21 of the terminal device 2 is not in the field of vision or at least that large parts of the display 21 are obscured by fingers or parts of the hand of the user 10. Furthermore, excessively rapid movement of the terminal device 2 may result in the content 22 not being recognized. Firstly, recognition of content 22 works more reliably if the terminal device 2 is not moved too much.On the other hand, a rapid movement of the terminal device 2 in area 7 may indicate that it has unintentionally entered area 7. Any reflections 23 on the display 21 can then be removed as explained below.

[0054] The image capture device 4 can, in particular, be a camera 4 that captures image data in the visible light range and can therefore be referred to as an RGB camera. While an IR camera or ToF camera as a monitoring device 3 delivers a suitable image under very different lighting conditions, the monitoring data 3 then does not contain any (RGB) colors. However, these are necessary for the recognition of content 22 displayed (in the visible light range) on a terminal device 2, such as on a display 21 of a smartphone (see Fig. 3). In an IR image, a terminal device 2, such as a smartphone, can only be recognized as a (usually) essentially rectangular object, on which the display 21 can at best be distinguished as a rectangular monochrome surface.

[0055] While in Fig. 2 the monitoring device 3 and the image capture device 4 are shown as separate units, they can also be combined in one unit, which in particular can be an RGB-IR camera, ie a camera which delivers image data in both the RGB and IR ranges.

[0056] The image capture device 4 can, in particular, also be arranged in the base 5 of the interior mirror 6. Therefore, the fields of view of the monitoring device 3 and the image capture device 4 can be assumed to be largely congruent, at least in area 7 (or even substantially congruent if the monitoring device 3 and the image capture device 4 are designed as a single unit). Therefore, in particular, image data located in area 7 are captured. In step S2, image data containing a view of the terminal 2 with the displayed content 22 is captured. The image data is then evaluated by means of image analysis, wherein the design 24, i.e., the layout of the displayed content 22, is captured (step S3). Based on the design 24, a function 9 of the vehicle can now be assigned to the content 22 by applying a machine learning model (step S4).In other words, the design 24 is analyzed, wherein the method 100 is able to classify the displayed content 22 based on the model, i.e., to assign it to a corresponding vehicle function 9 (also called a domain). For example, the displayed design of a music app and a navigation app differ on the terminal device 2, so that the reliability of text recognition can be improved because the method 100 already "knows" whether the recognized text belongs, for example, to an entertainment application or to a navigation application. In particular, the text is recognized as a function parameter for the specific function 9 in a sub-area 23 of the design 24 of the displayed content 22, which is predetermined by the design 24.

[0057] Before the actual application of the machine learning model in method 100, the model is trained with training data. The model is trained with training data containing a series of image data with a labeling of the domain and a marking of the relevant text, e.g., in the form of bounding boxes. In other words, in the training data, it is known for each image data which vehicle function (domain) it belongs to, e.g., music or navigation, and where in the respective view the corresponding function parameter can be found, e.g., text describing the song title or a navigation destination. In this way, method 100 can later easily identify the relevant sub-area 23 in which the text can be recognized and can ignore other texts, words, labels, and the like.To further improve recognition reliability, the model is trained using different views of the same content 22. For example, the content 22 can be scrolled, zoomed, or rotated in the image data, with a correct assignment being trained in each case. Designs belonging to a vehicle function 9 are also trained using different function parameters, e.g., different views of a music app with different song titles.

[0058] To recognize the functional parameter, known methods for optical character recognition (OCR) or for reading machine-readable codes, if the content is a code such as a QR code, can be used. Recognition can be performed quickly and reliably, not least because not the entire content 22 needs to be searched for text, for example, but only the partial area 23.

[0059] In a step S6, the correspondingly assigned vehicle function is then controlled depending on the recognized function parameter. Controlling a vehicle function can, in particular, comprise calling or starting a vehicle function assigned to a recognized content, in particular depending on the recognized content, more precisely the function parameter, i.e. the vehicle function is configured with the recognized function parameter. For example, if the domain is "navigation" and the text (function parameter) is an address, the vehicle function 9 "navigation system" can be called directly with the address as the set destination address, thus eliminating the need to manually enter the address in the vehicle's navigation system. If the recognized function parameter is a song title, the entertainment application (e.g., a streaming provider) with the recognized piece of music can be called directly, etc.

[0060] However, controlling a vehicle function can also simply comprise displaying data associated with recognized content, for example on a screen 8 of the vehicle, in particular a display of the infotainment system. Then, if necessary, provision can be made to offer a query for confirmation by the user, for example, whether the recognized content is correct and / or whether a function associated with the recognized content should be executed. For example: "The following address was recognized. Should route guidance be started?", "The following piece of music was recognized. Should it be played?", etc.

[0061] While at least one exemplary embodiment has been described above, it should be appreciated that a wide variety of variations exist. It should also be understood that the described exemplary embodiments are merely non-limiting examples and are not intended to limit the scope, applicability, or configuration of the devices and methods described herein. Rather, the foregoing description will provide a guide to implementing at least one exemplary embodiment, with the understanding that various changes in the operation and arrangement of the elements described in an exemplary embodiment may be made without departing from the subject matter as defined in the appended claims, as well as their legal equivalents.

[0062] LIST OF REFERENCE SYMBOLS

[0063] 100 Methods for detecting a mobile device in a vehicle

[0064] 1 Vehicle interior 2 End device (smartphone)

[0065] 3 Monitoring device (IR camera)

[0066] 4 Image capture device (RGB camera)

[0067] 5 Interior mirror base

[0068] 6 Interior mirror 7 Area

[0069] 8-display infotainment system

[0070] 9 Vehicle function

[0071] 10 users (drivers)

[0072] 21 Display device (smartphone) 22 Displayed content

[0073] 23 sub-area with function parameters

[0074] 24 Design of the displayed content

Claims

CLAIMS 1. A method (100) for detecting a mobile terminal (2) in a vehicle, the method comprising: - capturing image data by means of at least one image capturing device (4) in the interior (1) of a vehicle, wherein the image data contain at least one view of a terminal device (2); - recognizing a content (22) displayed on the terminal (2), wherein recognizing the content (22) comprises: - detecting a design (24) of the displayed content (22) by evaluating the image data; - Determining a function (9) of the vehicle associated with the displayed content (22) based on the detected design (24) by means of a trained machine learning model; and - Detecting a function parameter in the displayed content (22).

2. Method according to claim 1, wherein the recognition of the function parameter takes place in a partial area of the displayed content (22) which is determined on the basis of the detected design (24) and as a function of the specific function (9).

3. The method of claim 2, further comprising: - Determining an anchor point in the captured design (24) of the displayed content, wherein the anchor point defines the partial area (23) of the displayed content (22) in which the functional parameter can be recognized.

4. The method of claim 3, wherein the anchor point is a position in the design (249) of the displayed content (22) determined by means of the machine learning model.

5. The method according to claim 3 or 4, wherein determining the anchor point comprises detecting a marker in the detected design (24), wherein the detected marker is determined as the anchor point.

6. Method according to one of the preceding claims, wherein detecting the design (24) comprises detecting a pattern of the displayed content (22) and a comparison of the pattern with known patterns, in particular by means of the machine learning model.

7. The method according to any one of the preceding claims, further comprising: - Controlling the specific function (9) of the vehicle depending on the detected functional parameter.

8. Method according to one of the preceding claims, wherein the function parameter comprises text recognized by means of text recognition.

9. A method for training a machine learning model for recognizing a mobile terminal (2) in a vehicle, the method comprising: - Providing training data, wherein the training data comprises image data containing at least one view of a terminal (2), as well as an assignment of the image data to a function (9) of the vehicle; - recognizing a content (22) displayed on the terminal (2), wherein recognizing the content (22) comprises: - detecting a design (24) of the displayed content (22) by evaluating the image data; - Determining the function (9) of the vehicle associated with the displayed content (22) based on the detected design (24) by means of the machine learning model; and - Detecting a function parameter in the displayed content (22).

10. The method according to claim 9, wherein the training data define a sub-area (23) of the displayed content (22) in which the functional parameter can be recognized.

11. The method according to claim 9 or 10, wherein the training data contain image data on different views of the terminal (2) which are assigned to the same function (9) with the same function parameter.

12. A data processing system comprising at least one processor configured to carry out the method according to any one of claims 1 to 8.

13. A computer program comprising instructions which, when executed on a system according to claim 12, cause the system to carry out the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Application state and activity transfer between devices

    US20110275358A1

  • Vehicle head unit and method for setting screen of vehicle head unit

    US20150138045A1

  • Information entry via in-vehicle camera

    US20170213098A1