Method and system for controlling human-machine interface in vehicle

By detecting and distinguishing the gesture control scheme of the seat in the vehicle passenger compartment, the problem of indistinguishable gestures between the driver and the passenger is solved, safe and personalized user interface control is achieved, and driving safety and passenger experience is improved.

CN120447725APending Publication Date: 2025-08-08VALEO COMFORT & DRIVING ASSISTANCE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510145599.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-08
Filing Date
2025-02-10
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Gesture control schemes in existing vehicles are difficult to distinguish between driver and passenger gestures, resulting in interruption or interference in control, especially in multi-person passenger cabin environments, affecting driving safety and passenger experience.

Method used

By setting up a detection device in the vehicle passenger compartment, detecting and distinguishing the monitoring data of at least two seats, identifying gestures and determining the seat that performs gestures, thereby controlling the graphical user interface to ensure interaction is performed by only the user of the seat.

Benefits of technology

It realizes reducing control interruptions in a multi-person passenger cabin environment, improving driving safety and passenger experience, ensuring that the interaction between different users does not interfere, and enhancing personalized control of the user interface.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120447725A_ABST
    Figure CN120447725A_ABST
Patent Text Reader

Abstract

The invention relates to a method for controlling a human-machine interface (200) in a vehicle (100), the human-machine interface (300) comprising an input device (104, 105) for receiving a user input and a display device (103) for displaying a graphical user interface. The method comprises: obtaining monitoring data of at least a portion of a passenger compartment of the vehicle (100) by means of an input device, the input device comprising a detection device (104, 105) having a detection area (106) overlapping the portion of the passenger compartment, the portion of the passenger compartment comprising at least two seats (101) of the vehicle (100); processing the monitoring data to obtain gesture data from the monitoring data by detecting in the monitoring data a gesture performed by a passenger (102) of the vehicle (100) as a user input; determining one of the at least two seats (101) on which the passenger (102) sits, one seat performing a gesture for which gesture data is obtained; and controlling the graphical user interface according to the detected gesture and the determined seat.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and system for controlling a human-machine interface in a vehicle. In particular, the method and system detect gestures as user input in the vehicle. Background Art

[0002] Modern vehicles, such as cars, typically have a wide range of functions, which typically include various vehicle or comfort functions, such as navigation system settings, air conditioning, seat settings, lighting settings, etc. It is also possible to operate various functions of the infotainment system, such as playing music, making calls, etc. A user, in particular the driver or another passenger, such as a co-pilot, can interact with the vehicle via a human-machine interface, which can have an input device for receiving user input and a display device for displaying a graphical user interface (GUI). In particular, at least one display is usually provided as part of the user interface, in order to display and control functions centrally, for example in the dashboard. Individual functions, menus, etc. can be displayed here. These displays are usually touch-sensitive, so that the desired function can be controlled by touching the display. Other displays are known, such as a head-up display that displays a GUI in the windshield of the vehicle.

[0003] In vehicles, it is known to control functions contactlessly using gestures, in particular hand gestures. To do this, the user performs certain predefined gestures in a defined spatial region, such as in the vehicle cabin above the center console or in front of the instrument panel (in particular, the corresponding display), and these gestures are detected by corresponding detection devices (e.g., cameras or other sensors). This may require calibrating the system in the vehicle's global coordinates, which specify the defined spatial region relative to the display in which the gestures are to be performed.

[0004] Typically, gesture control solutions help prevent car accidents because they tend to help drivers or passengers keep their eyes on the road while controlling some features of the car, thereby reducing their attention span required to control vehicle functions. However, in vehicles with both a driver and a passenger, both individuals' hand movements can be interpreted as gestures, potentially leading to interruptions or interference with gesture control. Furthermore, if a passenger, such as the front passenger, is performing gestures, the driver may be distracted by the display of a graphical user interface that the other passenger is interacting with. Summary of the Invention

[0005] It is an object of the present invention to provide an improved method for controlling a human-machine interface in a vehicle. In particular, it is desirable to improve the detection of user input in a vehicle by gestures.

[0006] The solution to this problem is provided by the teaching of the independent claims. Various preferred embodiments of the invention are provided by the teaching of the dependent claims.

[0007] A first aspect of the invention relates to a method for controlling a human-machine interface in a vehicle, in particular a computer-implemented method. The human-machine interface comprises an input device for receiving user input and a display device for displaying a graphical user interface. The method comprises obtaining monitoring data of at least a portion of a passenger compartment of the vehicle by means of the input device. The input device comprises a detection device having a detection area that overlaps with a portion of the passenger compartment, wherein the portion of the passenger compartment comprises at least two seats of the vehicle. The method further comprises processing the monitoring data by detecting a gesture performed by a passenger of the vehicle in the monitoring data to obtain gesture data from the monitoring data as user input, determining one of the at least two seats in which the gesture is performed by the passenger, obtaining the gesture data, and controlling the graphical user interface in dependence on the detected gesture and the determined seat. Furthermore, a function of the vehicle associated with the gesture may be controlled.

[0008] Thus, the method of the first aspect takes into account not only the detected gesture, but also the location in the passenger compartment where the gesture is performed. More specifically, the method can determine the seat of the vehicle and thereby determine the person performing the gesture, such as the driver or another passenger. The HMI is controlled depending on the gesture and the determined seat. In other words, the method monitors not only one person (e.g. the driver) or only a single detection area (e.g. in the center of the vehicle), but also a part of the passenger compartment comprising at least two seats (e.g. the two front seats). This allows for improved control because more than one user is enabled to take over control, wherein control also depends on the user's location in the vehicle. For example, specific functions can be enabled for different users, which can reduce distraction for the driver while providing full control for the passenger.

[0009] As used herein, the term "vehicle" refers particularly to automobiles, including any type of motor vehicle, hybrid electric vehicles and battery electric vehicles, as well as other vehicles such as trucks, vans, or buses. A vehicle may have a passenger compartment (also called a "cabin") having one or more seats for vehicle occupants (including the driver and possibly a front passenger).

[0010] As used herein, the term "human-machine interface" (HMI) refers to a system that enables interaction between a user and a machine. In the context of the present invention, an HMI refers to an interface through which a driver or passenger can interact with a vehicle. Interactions can in particular be performed to control functions of the vehicle. More specifically, gestures can be used to control functions, in particular via a graphical user interface (GUI). Thus, an HMI includes an input device, in particular a detection device configured to detect a user's hand, such as a camera. In addition, an "output device" (in particular a display device) is provided to display the GUI.

[0011] The term "user interface" or "graphical user interface" as used herein refers in particular to a graphical representation of control elements that are linked to a specific function and allow a user to control that function. A user interface (UI) or graphical user interface (GUI) may comprise control elements, such as input surfaces, buttons, symbols, buttons, icons, sliders, toolbars, selection menus, etc., that a user can actuate without touching them, in particular within the meaning of the present invention. In particular, a GUI may be displayed on a display device such as a display, screen, monitor, etc.

[0012] As used herein, the term "gesture" particularly refers to a posture or movement of a user, particularly a part of the user's body, such as the left hand, right hand, or both hands. Therefore, a "gesture" may also be referred to as a "hand pose". A gesture may include the position and orientation of a hand in three-dimensional space, including movement, as well as the position or movement of one or more fingers of the hand. In particular, a gesture may consider a pointing gesture, wherein the direction of a "pointing finger" (usually the index finger) is determined. Gestures are detected and processed as user input.

[0013] As used herein, the term "user input" refers in particular to user interaction with an HMI (or graphical user interface). This may be a simple movement of a pointer (also called a "cursor") on a graphical user interface or control of functions, such as selecting and activating a control element (in particular by "clicking" or "double-clicking"), navigating through the user interface (e.g., "scrolling"), changing the display of objects or control elements, including moving objects (in particular, "drag and drop").

[0014] As used herein, the term "input device" particularly refers to a device for receiving user input. Although any type of user input may be received, the "input device" of the present invention may particularly be a "detection device".

[0015] The term "detection device", as used herein, particularly refers to a device that can detect objects in three-dimensional space and determine their position without contact. In particular, the detection device can detect the hand of a user. For example, optical methods can be used to detect the hand of a user in space. The detection device can consist of one or more parts, depending on which detection area is to be covered. For example, a (2D) camera or a 3D sensor can be provided. The detection area is the area within which events or changes can be perceived by the detection device, i.e. in the context of the present disclosure in particular the area (or more precisely the three-dimensional space area) in which the hand can be detected. In the case of a camera or other optical detection device or sensor, this can in particular also be referred to as the "field of view" (FOV).

[0016] The term "function", as may be used herein, refers in particular to technical features that may be present in a vehicle, for example in the interior, so as to be controlled by a corresponding control system. In particular, these may be functions of the vehicle and / or infotainment system, such as lighting, audio output (e.g. volume), air conditioning, telephone, etc.

[0017] Where applicable, the terms "first," "second," "third," and the like in the description and claims are used to distinguish between similar elements and not necessarily to describe a sequential or chronological order. It is to be understood that the terms so used are interchangeable under appropriate circumstances, and that the embodiments of the invention described herein are capable of operation in other sequences than described or illustrated herein.

[0018] Where the term "comprises" or "comprising" is used in the present description and claims, it does not exclude other elements or steps. Where an indefinite or definite article is used when referring to a singular noun (e.g. "a", "an", "the"), this includes a plural of that noun unless specifically stated otherwise.

[0019] Furthermore, unless expressly stated to the contrary, "or" refers to an inclusive or and not to an exclusive or. For example, condition A or B is satisfied by any of the following: A is true (or exists) and B is false (or does not exist), A is false (or does not exist) and B is true (or exists), and both A and B are true (or exist).

[0020] In the following, preferred embodiments of the method are described, which can be arbitrarily combined with one another or with other aspects of the invention, unless such a combination is explicitly excluded or technically impossible.

[0021] In some embodiments, the method further includes processing the monitoring data to obtain seat occupancy data, the seat occupancy data specifying a corresponding seat occupancy state for each of the at least two seats, wherein the seat occupancy data is obtained before processing the monitoring data to obtain the gesture data. Specifically, for each seat, a bounding box may be used to identify whether the seat is occupied. Considering the seat occupancy state for each state, including detecting and tracking seat occupancy, may improve the consistency of uninterrupted control of the HMI. Knowing the seat occupancy state for each seat may also be used to prioritize driver interactions over passenger interactions, or to enable driver-only control.

[0022] In some embodiments, gesture data is obtained only for seats whose seat occupancy status is determined to be occupied, wherein processing the monitoring data to obtain gesture data includes determining a bounding box for one or both hands of a passenger in the occupied seat. This can improve the stability of the control method because gesture data can only be presented for seats that are actually occupied. Vice versa, monitoring data in the area of unoccupied seats logically cannot convey gesture data. For each occupied seat, the right palm and the left palm can be detected, for example, using a mobile real-time deep learning object detection model. This can then also result in detected hand bounding boxes for the right and left sides of each occupied seat, which can facilitate the detection of gestures when limited to the relevant bounding boxes of the hands.

[0023] In some embodiments, when a gesture is detected by one passenger, any gestures by any other passenger are ignored. Specifically, using seat occupancy status, specifically by detecting and tracking seat occupancy, interaction gestures can be tracked without interruption from other hand movements within the vehicle, which can typically occur with the hand movements of multiple passengers. In other words, once a gesture is detected, only interactions from the same seat and hand that initiated the interaction can be allowed. The remaining seat detections can be ignored, ensuring that gesture control is not interrupted by hand movements of other passengers.

[0024] In some embodiments, controlling the graphical user interface includes displaying the graphical user interface at a position on the display device that is dependent on the position of the seat that the passenger performing the gesture is sitting in. This can improve interaction with the HMI because the HMI can be designed to not affect the driving experience when the position of a GUI (such as a widget or icon) changes based on the seat from which the gesture control is initiated, thereby enhancing the user experience.

[0025] In some embodiments, processing the surveillance data to obtain gesture data includes detecting multiple landmarks of the corresponding hand, where the detected landmarks are used to detect gestures performed by the passenger. By using landmarks of the hand rather than, for example, an entire hand image, gesture detection is less sensitive to background noise and requires less processing time. Landmarks can be prominent key points on the hand, such as fingertips and joints. Starting from the wrist, the hand can be described specifically by 21 landmarks. In this way, gestures, hand postures, or pointing directions can be easily detected or classified. Specifically, using a landmark model rather than, for example, image recognition that relies on the entire image can reduce computational resources. In particular, the aforementioned bounding box of the hand can describe the area where the landmarks can be found. A model such as a mobile deep learning regression network can be used to detect the 21 landmarks for each detected hand's bounding box. This allows for a compact and real-time gesture classifier (see below) that is invariant to any lighting conditions, as the gesture classifier now relies solely on the detected landmarks as input, without the image itself or background interference.

[0026] In some embodiments, the detected landmarks are input into a gesture classification model, where the output of the gesture classification model specifies the gesture performed by the passenger. In other words, for gesture classification, the landmarks (particularly the occupied tracked seat right hand landmark and the left hand landmark) can be passed to a classifier network that takes as input the 21 landmarks and classifies the gesture. The use of a gesture classification model (which uses the detected hand landmarks) allows for more robust predictions without hand background image interference, as well as real-time predictions for hand movement classification. In contrast, processing gestures directly from the captured images without any prior processing would have very low performance because multiple hands may be present in the captured images within the vehicle cabin.

[0027] In some embodiments, the gestures specified by the gesture classification model include at least one of a hand pose classification and a pointing finger motion classification. This means that the model may be able to distinguish between two types of gestures, namely (stationary) hand poses and (moving) pointing motions. The pointing finger may specifically be an index finger.

[0028] In some embodiments, the method further includes determining a projection of the position of the pointing finger onto the display device, wherein the projection is determined by determining the finger vector direction using a landmark of the corresponding finger. The projection can provide a conversion from 3D coordinates of the hand in the cabin of the vehicle to 2D coordinates on the display device, which can be, for example, a windshield when the display device is provided as a heads-up display. This can be accomplished by using camera calibration parameters and the index finger coordinates of the pointing finger to obtain a projection on the display device. This allows the pointing finger to be used to move a cursor or other item of the GUI.

[0029] In some embodiments, if a gesture is classified as a pointing finger motion classification, the gesture is determined by processing monitoring data for at least two subsequent time points ("frames"). This allows tracking of the movement of the hand, and more specifically the movement of the pointing finger. For example, the index finger / pointing finger landmark positions can be recorded and a queue of short timestamps of the processed frames can be filled. Once the queue is full, the pointing finger motion classification deep learning model can be evaluated with the index finger / pointing finger positions of all entries in the queue. This may result in different classifications such as "not moving", "moving", "clockwise motion", "counterclockwise motion". It should be understood that if any gesture other than a pointing finger gesture is detected during the queue fill time, the queue can be emptied so as to not carry information from the interrupted gesture.

[0030] In some embodiments, processing the surveillance data to obtain gesture data includes determining the 3D coordinates of detected landmarks. Gesture recognition can be improved if the landmark's 3D world coordinates are known. Processing can be faster and more cost-effective, particularly compared to extensive image recognition or detection of gestures from a large number of points in a radar point cloud, for example. A deep learning model can be used that takes the coordinates as input.

[0031] A second aspect of the present invention relates to a data processing system configured to perform the method of the first aspect. The data processing system can be configured to perform the method of the first aspect using one or more computer programs. Additionally or alternatively, this configuration can be implemented in whole or in part by corresponding hardware. The system includes at least one display device configured to display a graphical user interface and an input device configured to receive user input.

[0032] In some embodiments of the system, the display device includes at least one display, monitor, screen, etc., which is arranged in the vehicle, for example, as part of an infotainment system. In particular, the display device can be a head-up display (HUD) configured to display an image in the windshield of the vehicle.

[0033] In some embodiments of the system, the input device includes a detection device that includes at least one image detection device, in particular a camera and / or at least one 3D sensor. A camera can be used to easily determine the position of a user's hand in three-dimensional space. One camera or multiple cameras can be provided as the detection device. In particular, in a vehicle, a rear camera and a front camera can be provided to cover a large detection area in the passenger compartment of the vehicle. The camera can be an infrared camera, or can capture images in the visible spectrum. The camera can be a time-of-flight camera (ToF camera). By using such a 3D sensor device, the position of the hand in three-dimensional space and its movement can be directly detected. 2D sensors can also be combined to detect the position of the hand in three-dimensional space. As described above, the field of view of at least one camera is in a portion of the passenger compartment of the vehicle having at least two seats (preferably the two front seats) of the vehicle, or extending to all seats of the vehicle. In this way, gestures made by any passenger of the vehicle can be detected.

[0034] A third aspect of the invention relates to a computer program or computer program product comprising instructions which, when executed on a data processing system according to the second aspect of the invention, cause the system to perform the method according to the first aspect of the invention.

[0035] The computer program (product) can be implemented in particular in the form of a data carrier on which one or more programs for executing the method are stored. Preferably, this is a data carrier such as a CD, DVD or other optical medium, or a flash memory module. It may be advantageous if the computer program product is intended to be traded as a separate product independent of the processor platform on which the one or more programs are to be executed. In another embodiment, the computer program product is provided as a file on a data processing unit, in particular a file on a server, and can be downloaded via a data connection, for example the Internet or a dedicated data connection, such as a proprietary or local area network.

[0036] The system of the second aspect may accordingly have a program memory storing the computer program. Alternatively, the system may also be arranged to access an externally available computer program via a communication link, for example on one or more servers or other data processing units, in particular to exchange data therewith that is used during the execution of the computer program or represents output of the computer program.

[0037] The explanations, embodiments and advantages described above in conjunction with the method of the first aspect apply analogously to the other aspects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Other advantages, features, and applications of the present invention are provided in the following detailed description and accompanying drawings, in which:

[0039] Figure 1 shows a bird's-eye view of a vehicle with a driver and a human-machine interface;

[0040] Figure 2 The functions of the human-machine interface are schematically shown;

[0041] Figure 3 A flow chart showing a method of controlling a human-machine interface;

[0042] Figure 4 The hand signs are schematically illustrated;

[0043] Figure 5 schematically illustrating the projection of pointing directions onto a display device;

[0044] Figure 6 showing a view through a windshield of a vehicle with a displayed graphical user interface controlled by the driver; and

[0045] Figure 7 A view through the windshield of a vehicle is shown with a displayed graphical user interface controlled by a passenger. DETAILED DESCRIPTION

[0046] Figure 1 A vehicle 100 is shown that is equipped with a human machine interface (HMI) 200 to allow the driver 102 or another passenger of the vehicle 100 (e.g., a co-pilot) to interact with the vehicle 100. In particular, a graphical user interface (GUI) and related functions of the vehicle 100 can be controlled. An embodiment will be described in more detail below. Figure 2 Components or functions of the HMI 200 are schematically shown in FIG.

[0047] The HMI 200 has an input device for receiving user input. More specifically, the input device is configured as a detection device to detect gestures of a user (such as the driver 102). In particular, one or more cameras 104, 105 are provided to monitor the passenger compartment (carriage) of the vehicle 100. Figure 1 As shown, a front camera 104 and a rear camera 105 are arranged to cover a detection area 106 capable of detecting hand gestures. Detection area 106 extends over a portion of the passenger compartment having at least two seats, such as a driver's seat 101 and a passenger seat (not shown). In this manner, HMI 200 can detect gestures by both the driver 102 and the passenger. A display device 103 is provided to display a GUI. In this example, display device 103 is a head-up display (HUD) for the windshield of vehicle 100.

[0048] Before describing the method of controlling the HMI 200 in more detail, Figure 2Briefly explain the basic functional components of HMI 200. Gestures are detected and input to be recognized as movements or actions (process 201). Determine where in the vehicle the gesture is made, i.e., from which seat or which user. In other words, a distinction is made between the driver's hand and the co-pilot's hand (process 202). Based on the identified user (driver or co-pilot), the cursor of the GUI is moved to the desired position (process 203). In particular, if the driver performs the gesture, the cursor may appear on the driver's side of the windshield, while if it is determined that the co-pilot performs the gesture, the cursor may appear on the co-pilot's side of the windshield. In addition, it is determined whether the gesture is a valid gesture (process 204), which may depend on the current application controlled by the corresponding user. Action is then taken based on the gesture, the current user, and the current application (process 205). The corresponding output of the GUI on the display device 103 is provided.

[0049] That being said, the overall process will be explained. This may in particular include certain deep learning techniques. In order to detect or capture gestures made by any occupant of the vehicle 100, front seat images and rear seat images are captured by the front compartment monitoring camera 104 and the rear compartment monitoring camera 105. The front image and the rear image are processed by a seat occupancy object detection deep learning model to determine whether the seat is occupied. This can be done using a corresponding bounding box for each seat. This step is done to ensure consistency in non-interrupted interaction with the HMI. By tracking the occupancy state, it can be ensured that an operation (gesture control) started by one of the users can be completed by that user. During active gesture control, gestures of other passengers can be ignored. Furthermore, gestures can be detected only for seats that are actually occupied. It is also conceivable to assign priority to driver interactions over passengers, or to enable only the driver to take control if configured.

[0050] For each occupied and tracked seat bounding box, the right and left palm bounding boxes are detected, in particular using a mobile real-time deep learning object detection model. This produces right and left detected hand bounding boxes for each occupied detected and tracked seat. Then, in order to robustly obtain a classification of the hand gesture, rather than relying on the hand image itself, an intermediate step is performed, which is to detect hand landmarks 400 (see Figure 4 A mobile deep learning regression network is used to detect 21 landmarks for the bounding box of each detected hand. This allows for a very small and real-time gesture classifier that is invariant to any lighting conditions, as the gesture classifier now only depends on the detected landmarks 400 as input, without the image itself or background interference.

[0051] For gesture classification, the right and left hand landmarks of each occupied tracked seat are passed to a (very small) classifier network that takes 21 landmarks as input and classifies the gesture of interest to control the HMI. In case a pointing finger gesture is classified, the index (pointing) finger landmark positions are recorded (landmarks 5, 6, 7, 8) and the queue is filled with processed frames for a short period of time. Once the queue is full, the pointing finger motion classification deep learning model is evaluated with the index finger positions of all entries in the queue. The model can detect the following classes: "no movement", "movement", "clockwise movement", "counterclockwise movement". If any gesture other than a pointing finger gesture is detected during the queue filling time, the queue is cleared so as not to carry information from the interrupted gesture.

[0052] As will be described in more detail below, the pointing direction of the index finger is projected onto the display device 103. The camera calibration parameters, together with the pointing finger index coordinates (more specifically, the 3D coordinates of markers 5, 6, 7, and 8), can be used to obtain a projection on the (2D) windshield head-up display. The application manager can then transmit the information detected by the described algorithm (projected pointing finger x, y position on the HUD windshield screen, gesture classification, pointing finger motion classification) to the HMI application via shared memory.

[0053] The HMI 200, in particular the GUI, should be designed so as not to interfere with the driving experience. Therefore, the GUI may have a transparent design, i.e., icons, widgets, etc. may appear on the windshield with a certain degree of transparency. In addition, the position of the GUI may be changed based on the seat from which the gesture control is initiated in order to enhance the user experience (see respectively Figure 6 and Figure 7 ). An interaction can be initiated when one of the users in the occupied seats starts with a certain gesture (such as an "open gesture"). The HMI 200 also keeps track of the seat that has initiated the interaction and ignores detection of other seats. Therefore, gesture control is not interrupted by hand movements of other passengers in the cabin of the vehicle 100. When the HMI 200 receives a pointing finger gesture from a tracked and gesture-enabled seat, the cursor is moved based on the projected pointing finger position to enable the gesture interaction initiator to move through and select HMI widgets.

[0054] The following special reference Figure 3 The above process is further described by way of example. In process 301, seat occupancy is determined. Specifically, a captured image 310 is processed using an object detection model that detects the following categories: human body, human hand. Bounding boxes of human body detections are processed 311, and each bounding box is assigned to a predefined and preconfigured seat location using a parametric algorithm that uses maximum and minimum intersection thresholds for the bounding box area. This can reduce the number of false positives in detection.

[0055] Palm detection is then performed in process 302. The output of the previous object detection model (bounding boxes 311) is processed. This can use the same parameter algorithm as in seat occupancy, but now the detected hands are assigned to the assigned occupied seat human bounding box 311, which results in left / right hand bounding boxes 312 for the occupied seat.

[0056] Then, hand landmark regression is performed in process 303. In this process, a batch of detected hands (more specifically, images cropped to left and right hand bounding boxes 312) are passed to a deep learning hand landmark regression model, which detects 21 landmarks 313 in 3D world coordinates (x, y, z). Therefore, for each occupied seat, this results in 21 landmarks for the right and left hands, which are in Figure 4 It is shown as symbol 400 in FIG.

[0057] The detected landmarks 313 are then used in process 304 to perform gesture classification. This component is specifically responsible for classifying the instant hand landmark category without looking at previous frames. For classification, a minimal deep learning model can be used that takes as input the hand landmark relative positions in (x,y) from the hand landmark origin O (wrist). More specifically, the model can be designed to use the hand landmark relative positions instead of the image itself, so that the model becomes independent of changes in image lighting conditions that may affect the classification quality. In addition, it helps to make the gesture classification lightweight because it does not process the entire image space, but only the 2D vectors of the 21 landmarks. The output 314 is the instant gesture / landmark classification for the left and right hands of each occupied seat. The gestures can be, for example, open, closed, pointing finger, click, up, down, right, left, etc.

[0058] Furthermore, a time-dependent pointing finger classification is performed in process 305. The pointing finger gesture can be considered to be of particular importance. Once the gesture is detected by the gesture classification 304, the time-dependent pointing finger classification component 305 starts processing the (x,y) position of the pointing finger (typically the index finger), where the time aspect is classified to detect movement. In particular, a plurality of frames over time are classified for the pointing finger motion category. For the time-dependent classification, an LSTM deep learning model architecture with a minimal implementation can be used. It obtains the position of the pointing finger, in particular the marker 8, and outputs (315) the time-dependent pointing finger motion category for each occupied seat. The following categories can be determined: idle, moving, clockwise and counterclockwise.

[0059] Pointing finger projection can then be performed in process 306, which will be described in further detail below. The pointing finger described by the markers 5, 6, 7, and 8 in the 3D vector is projected onto a 2D position (316) on the display device 103 (i.e., the windshield HUD), which allows for accurate and smooth interaction. The process can use the camera calibration parameters 317 and the 3D world coordinates to estimate the projected 2D image coordinates on the HUD image space.

[0060] In summary, the above method is capable of recognizing two main types of gestures, which can be referred to as "hand markers" (or hand gestures) and time-dependent pointer finger markers (i.e., movements or motions). For hand markers, the gesture classification 304 at one time frame does not require any time dependency. For time-dependent pointer finger markers, the pointer finger motion gesture classification 305 over multiple frames (moving windows) requires motion dependency. It will be appreciated that the instantaneous hand marker classification gestures and the time-dependent pointer finger marker classification gestures are not limited, but rather depend entirely on the number of gestures required for the desired HUD interaction application.

[0061] Now, specifically refer to Figure 4 as well as Figure 5 , describes a specific example of projection of a hand 500 (more specifically, an index or pointing finger 501) onto a display device 103. Figure 1 As mentioned above, the camera arrangement should be configured such that the detection area 106 is large enough to cover hand movements for more than one seat in the vehicle 100. At least two cameras 104, 105 may be provided, wherein a front camera 104 and a rear camera 105 may be provided. Note that since the 3D hand landmark model 400 detects 3D coordinates in world space, the orientation and position depend on the vehicle model, provided that the detection area 106 covers all hand movement scenarios, and therefore does not limit gesture detection.

[0062] Using the pointing finger 3D markers 5, 6, 7, 8, the 3D vector of the pointing direction 502 can be estimated in world coordinates. The finger vector can be defined as a ray vector:

[0063] Ray(t)=(x i t,y i t,t).

[0064] The screen plane of a display device can be defined as:

[0065] Ax+By+Cz+D=0.

[0066] Combining a 3D plane and a 3D vector results in:

[0067]

[0068]

[0069] By substituting t into Ray(t), the intersection point on the screen of the display device 103 is obtained.

[0070] Gestures from the driver 102 or from another passenger (such as the front passenger) are input to the HMI 200 to control various functions of the vehicle, such as the music player, door locks, AC, windows, messages, phone calls, etc. The driver 102 and the passenger (front passenger) can have their own field of view for better visualization and easier control. For example, when the driver 102 takes control, the GUI 600 is shown on the left (in front of the driver), as shown in FIG. Figure 6 As shown, and in the case where the co-pilot takes control, the GUI 700 is shown on the right side (also without interfering with the driver), as shown Figure 7 shown. Figure 6 and Figure 7 A music player app is shown as an example.

[0071] When the vehicle 100 is moving, for greater safety, most controls should be taken away from the driver 102. For example, if the driver 102 attempts to diagnose the car while exceeding a certain speed limit, they will not be able to use it. Furthermore, it can be provided that a passenger, such as a co-pilot, cannot override the driver 102, and vice versa. For example, when the driver 102 is controlling an application, a passenger cannot use gestures to control the application, meaning that the person controlling the HMI must release control so that the other person can take control again.

[0072] Below, provide some examples of panel controls, wherein each application (screen) can have its own unique controls. It should be understood that the following list is only exemplary and not limiting. To switch between applications, the user can use MoveRight and MoveLeft gestures.

[0073] Music: Users can use clockwise and counterclockwise gestures to change the volume, or use MoveTop and MoveDown gestures to go to the next / previous song. Additionally, users can use tap gestures to pause playback, go to the next or previous song.

[0074] AC: User can turn on / off the AC and change the mode between hot and cold using tap gestures. By using clockwise and counter-clockwise gestures, user can change the temperature.

[0075] Lock: Users can lock / unlock the vehicle using a tap gesture or MoveTop and MoveDown gestures.

[0076] Window: Users can use a tap gesture to select a window, and then they can use the MoveTop and MoveDown gestures to lower or raise the window.

[0077] Phone: Users can use gestures to accept, decline, or make phone calls.

[0078] Cluster / Speedometer / Fuel Tank: The HMI may also be able to display a cluster including the car's speed, fuel tank, etc.

[0079] Diagnose: This can show the broken part of the vehicle in red without user interaction. The user can click on the part to diagnose it.

[0080] Help: This may show all available gestures without interaction (which may be refined with instructions).

[0081] While at least one exemplary embodiment of the present invention has been described above, it must be noted that there are numerous variations thereto. Furthermore, it should be understood that the exemplary embodiments described merely illustrate non-limiting examples of how the present invention may be implemented and are not intended to limit the scope, application, or configuration of the apparatus and methods described herein. Rather, the foregoing description will provide those skilled in the art with a framework for implementing at least one exemplary embodiment of the present invention, wherein it must be understood that various changes may be made to the function and arrangement of the elements of the exemplary embodiments without departing from the subject matter defined by the appended claims and their legal equivalents.

[0082] Reference Signs List

[0083] 1-21 Hand Signs

[0084] 100 vehicles

[0085] 101 Driver's Seat

[0086] 102 Driver

[0087] 103 Display Devices

[0088] 104 Camera

[0089] 105 Camera

[0090] 106 Detection Area

[0091] 200 Human Machine Interface (HMI)

[0092] 201-205 Functional components of HMI

[0093] 300 Methods of controlling HMI

[0094] 301-306 Process or steps of method 300

[0095] 310-316 Data or parameters in method 300

[0096] 400 Hands Logo Model

[0097] 500 lots

[0098] 501 Index Finger

[0099] 502 Pointing direction

[0100] 503 Projection point on display device

[0101] 600 GUI controlled by driver

[0102] 700 GUI controlled by co-pilot

Claims

1. A method (300) for controlling a human-machine interface (200) in a vehicle (100), the human-machine interface (300) comprising an input device (104, 105) for receiving user input and a display device (103) for displaying a graphical user interface, the method comprising: - obtaining monitoring data of at least a portion of a passenger compartment of the vehicle (100) by means of the input device, the input device comprising a detection device (104, 105) having a detection area (106) overlapping the portion of the passenger compartment, wherein the portion of the passenger compartment comprises at least two seats (101) of the vehicle (100); - processing the monitoring data by detecting a gesture performed by a passenger (102) of the vehicle (100) in the monitoring data to obtain gesture data from the monitoring data as the user input; - determining a seat of the at least two seats (101) on which the passenger (102) who performed the gesture is seated, and obtaining the gesture data for the gesture; and - controlling the graphical user interface based on the detected gesture and the determined seat.

2. The method according to claim 1 further includes processing the monitoring data to obtain seat occupancy data, wherein the seat occupancy data specifies a corresponding seat occupancy status for each of the at least two seats (101), wherein the seat occupancy data is obtained before processing the monitoring data to obtain the gesture data.

3. The method of claim 2 , wherein gesture data is obtained only for seats whose seat occupancy status is determined to be occupied, wherein processing the monitoring data to obtain gesture data includes determining a bounding box of one or both hands of a passenger in the occupied seat.

4. A method according to any one of the preceding claims, wherein When a gesture by a passenger (102) is detected, any gesture by any other passenger is ignored.

5. A method according to any of the preceding claims, wherein controlling the graphical user interface comprises displaying the graphical user interface at a position on the display device (103), the position being dependent on the position of the seat (102) in which the passenger (101) performing the gesture is seated.

6. A method according to any one of the preceding claims, wherein processing the monitoring data to obtain gesture data includes detecting a plurality of landmarks (400) of respective hands (500), wherein the detected landmarks (400) are used to detect gestures performed by the passenger (102).

7. The method according to claim 6, wherein: The detected landmark (400) is input into a gesture classification model, wherein an output of the gesture classification model specifies a gesture performed by the passenger (102).

8. The method according to claim 7, wherein: The gesture specified by the gesture classification model includes at least one of a hand pose classification and a pointing finger motion classification.

9. The method according to claim 8, further comprising: A projection of a pointing finger (501) to a position (503) on the display device (103) is determined, wherein the projection is determined by determining a finger vector direction using a landmark of the corresponding finger (501).

10. The method according to claim 8 or 9, wherein: If the gesture is classified as a pointing finger motion classification, the gesture is determined by processing monitoring data at at least two subsequent time points.

11. The method according to any one of claims 6 to 10, wherein: Processing the monitoring data to obtain gesture data includes determining 3D coordinates of the detected marker (400).

12. A data processing system comprising at least one processor configured to perform the method according to any of the preceding claims, at least one display device (103) configured to display a graphical user interface, and input devices (104, 105) configured to receive user input.

13. The system according to claim 12, wherein: The display device (103) is a head-up display.

14. The system according to claim 12 or 13, wherein: The input device comprises a detection device comprising at least one camera (104, 105) and / or at least one 3D sensor.

15. A computer program or computer program product comprising instructions which, when executed on one or more processors of a system according to any one of claims 12 to 14, cause the system to perform the method according to any one of claims 1 to 11.