Method and system for capturing user input
The method and system address the challenge of clear and quick user input detection in vehicles by enabling intuitive gesture-based control of large or 3D displays, reducing distraction and improving usability and safety.
Patent Information
- Application Number
- DE102023133482
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-08-28
- Estimated Expiration
- 2043-11-30
AI Technical Summary
Existing methods for detecting user inputs in vehicles, particularly with large displays or 3D displays, face challenges in providing clear and quick representations, leading to user distraction during operation.
A method and system for detecting user inputs through hand gestures in three-dimensional space, where a graphical user interface is manipulated by converting hand postures and movements into actions on a display, allowing for intuitive and precise control of display elements without physical contact.
Enhances user interaction by reducing distraction and improving the usability of large or 3D displays in vehicles by allowing intuitive and accurate manipulation of graphical user interfaces through gesture recognition, thereby enhancing safety during vehicle operation.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The present invention relates to a method and a system for detecting user inputs, particularly in a vehicle. In particular, gestures can be detected as user inputs in a vehicle.
[0002] Modern vehicles often have a wide range of functions, which usually includes various vehicle or comfort functions, such as settings for the navigation system, the air conditioning, seat settings, lighting settings, and the like. Various functions of an infotainment system can also be operated, such as playing music, making phone calls, and the like. To display and control the functions, at least one display is usually provided as part of the user interface, for example, in the center of the dashboard. This is where the individual functions, menus, and the like can be shown. These displays are often touch-sensitive, so that the desired function can be controlled by touching the display. For this purpose, the controls are designed so that they can be reached and operated with a finger.
[0003] However, modern vehicles often have multiple displays and increasingly larger displays, which may no longer be easily accessible by hand. For example, several distributed displays or a large display that extends across the width of the vehicle, possibly curved, may be provided. It is therefore also known to control certain functions contactlessly using gestures. To do this, a user performs certain predefined gestures in free space, for example, in a spatial area of the vehicle cabin in front of the display. Such gestures can be detected by appropriate sensors that are capable of determining hand postures and movements in three-dimensional space.
[0004] However, such gestures often become more difficult the larger the display and the further at least parts of the display may be from the user. In particular, it can become difficult to precisely grasp and control controls using gestures. For example, operating a very large panoramic display using gesture control can be difficult if the gestures require a relatively elaborate execution.
[0005] In addition to conventional 2D displays, 3D displays are also known to expand the presentation of content on a display. These displays are capable of making the displayed content appear three-dimensional. However, three-dimensional representations are often limited to a relatively small display. For example, individual objects on the dashboard directly in front of the driver can be highlighted by creating a 3D effect. With such 3D displays, the three-dimensional effect can be created, for example, by combining two 2D displays. For this purpose, a transparent, reflective surface can be provided in the display to create a 3D effect in the viewer's perception from two coordinated two-dimensional representations. The viewer sees two superimposed representations through the reflective surface.
[0006] US 2012 / 0200494 A1 discloses a computer vision system for controlling electronic devices using hand gestures. The system captures images, detects movements, and applies shape recognition algorithms to identify a user's hand. Confirmation of hand identification is enhanced by combining information from multiple images or analyzing movement patterns. After identification, hand movements are tracked to generate device control commands, often with adjustable sensitivity and user aids for enhanced interaction.
[0007] US 9 009 594 B2 discloses gestures for controlling media playback. A computing device displays controls, including a search bar, on a user interface. Gestures detected by a camera enable interaction with these controls, for example, to navigate through media playback. When approaching a position in the search bar, it automatically enlarges to allow for more precise selection. Additionally, thumbnails of the media content can be displayed. Additionally, thumbnails of the media content can be displayed.
[0008] EP 3 491 493 B1 discloses gesture-based control of autonomous vehicles. A triggering condition, such as a hand gesture, initiates an interaction. A display shows control options, some of which are based on environmental analysis. Occupants select an option, such as changing lanes or parking, using gestures, and the vehicle executes the command. The system can personalize interactions based on user profiles and preferences.
[0009] A challenge with both 2D and 3D displays is to present the displayed content clearly. This is especially relevant in vehicles, as the systems can have a large number of different functions, the operation of which should distract the user from driving as little as possible. A clear and quickly comprehensible display is particularly important for very large displays with areas located far away from the driver.
[0010] The present invention is based on the object of providing an improved method for detecting user inputs. In particular, the operation of a system using gesture control is to be improved.
[0011] This object is achieved according to the teaching of the independent claims. Various embodiments and further developments of the invention are the subject of the dependent claims.
[0012] A first aspect of the invention relates to a method, in particular a computer-implemented method, for capturing user inputs in a system having a display device and a capturing device. A graphical user interface is displayed by means of the display device, in particular continuously during the method, and a user's hand is captured in a three-dimensional spatial area by means of the capturing device. A gesture is captured as user input by means of the capturing device, and at least a hand posture and a hand movement are determined for this purpose.If, when determining the hand posture as a gesture, a picking gesture is recognized which is defined by a transformation of the hand from a first hand posture to a second hand posture, then an action area of the graphical user interface is determined by assigning the picking gesture to at least one area of the graphical user interface, and the action area of the graphical user interface is moved according to the determined movement of the hand such that the action area on the display device is moved starting from a start position, as long as the second hand posture is determined as the hand posture.
[0013] The aforementioned method according to the first aspect is therefore based in particular on user inputs being recorded contactlessly in the form of gestures. For this purpose, hand postures and movements in three-dimensional space are recorded. If a picking up gesture, i.e. a gesture that signals a "pick-up" or a "grasp," is recorded based on the hand postures, this is assigned to an action area. The action area can be determined such that the picking up gesture is assigned to a specific area of the graphical user interface as the "action area." This area is referred to as the action area because an action is performed on it by the recorded gesture. In particular, the action area on the display device is moved from a starting position, i.e. a position in which it is located at the time the picking up gesture is recorded, in accordance with the movement of a user's hand while holding the second hand posture.Through this interaction, the user can change the appearance of the graphical user interface, particularly by rearranging it. This can be advantageous, for example, to move content on the user interface to a specific location, e.g., to bring it closer to the driver's field of vision. This can be particularly beneficial with large displays in vehicles, as it can reduce distraction while driving. The user can personalize the graphical user interface and adapt it to their needs.
[0014] The term “display device” used here refers in particular to a device by means of which content, such as a user interface (GUI), can be graphically represented. This can in particular be a display or a screen, especially in a vehicle. The display device can in particular be configured to show an overall representation which is composed of two partial representations which are displayed in different, in particular spaced-apart, image planes. The distance between the image planes and the superimposition of the corresponding partial representations can create a three-dimensional effect for the user (“3D display”). One of the two image planes can be referred to as “virtual” if the view of the overall representation results from looking through this “virtual” image plane to an image plane behind it. Technically, this can be, for example,This can be achieved by two displays positioned at an angle to each other (e.g., approximately 90 degrees) with a transparent, reflective optical element between them (e.g., at approximately 45 degrees). The display device can be designed as a "panoramic display" (2D or 3D), i.e., it can extend across the entire width of the vehicle or even further along the sides of the vehicle.
[0015] The term "user interface" or "graphical user interface" used here refers in particular to a graphical representation of control elements (operating elements) that are linked to a specific function and allow a user to control the function. The user interface ("UI" or "graphical user interface" - GUI) can contain "control elements" ("UI elements") such as input areas, buttons, symbols, buttons, icons, sliders, toolbars, selection menus, and the like, which a user can activate, particularly within the meaning of the present invention, without touching them, in order to control an associated function. The (graphical) user interface can also be referred to as a (graphical) user interface.
[0016] The term "detection device" used here refers in particular to a device that can detect objects in three-dimensional space and determine their position without contact. In particular, the detection device can detect and locate a user's hand. For example, optical methods can be used to detect a user's hand in space ("gesture control"). The detection device can consist of one or more parts, depending on the detection area to be covered. For example, one or more cameras can be provided. Detection takes place in a detection area, i.e., the field of view of the respective sensor. No detection takes place outside the detection area. This can prevent unintentional actions.
[0017] The term "three-dimensional spatial area" used here refers in particular to an area that can be described by three-dimensional coordinates. A position in the three-dimensional spatial area has unique three-dimensional coordinates. The coordinate system can be chosen arbitrarily. For example, a coordinate system of the detection device or a coordinate system related to the display device can be selected. In particular, there is a relationship between a detection area of the detection device and the display device in order to be able to assign the position of the hand to corresponding locations on the GUI.
[0018] The term “vehicle” as used herein refers in particular to a passenger car, including all types of motor vehicles, hybrid and battery-powered electric vehicles, as well as vehicles such as sedans, vans, buses, trucks, delivery vans and the like.
[0019] The term "function" used here refers in particular to technical features that may be present in a vehicle, for example, in the interior, to be controlled by a corresponding control system. In particular, these can be functions of the vehicle and / or an infotainment system, such as lighting, audio output (e.g., volume), climate control, telephone, etc.
[0020] The terms "comprises," "includes," "includes," "has," "has," "with," or any other variation thereof, as used herein, are intended to cover non-exclusive inclusion. For example, a method or apparatus that includes or has a list of elements is not necessarily limited to those elements, but may include other elements not expressly listed or that are inherent in such a method or apparatus.
[0021] Furthermore, unless explicitly stated to the contrary, "or" refers to an inclusive "or" and not an exclusive "or." For example, a condition A or B is satisfied by one of the following conditions: A is true (or present) and B is false (or absent), A is false (or absent) and B is true (or present), and both A and B are true (or present).
[0022] The terms "a" or "an" as used herein are defined as "one or more." The terms "another" and "another," and any other variations thereof, are defined as "at least one other."
[0023] The term “plurality” as used here shall mean “two or more”.
[0024] The term “configured” or “set up” to fulfil a specific function (and respective modifications thereof) is to be understood within the meaning of the invention that the corresponding device is already in a design or setting in which it can carry out the function or is at least adjustable - i.e. configurable - so that it can carry out the function after being set accordingly. The configuration can be carried out, for example, by appropriately setting parameters of a process sequence or of switches or the like for activating or deactivating functionalities or settings. In particular, the device can have a plurality of predetermined configurations or operating modes, so that the configuration can be carried out by selecting one of these configurations or operating modes.
[0025] Preferred embodiments of the method are described below, which, unless expressly excluded or technically impossible, can be combined with each other as well as with the other aspects of the invention described.
[0026] In some embodiments, the movement of the action area of the graphical user interface is stopped at a target position if, when determining the hand posture as a gesture, a drop gesture is detected, which is defined by a transformation of the hand from the second hand posture to the first hand posture. After the start of the movement, i.e. the movement of the specific action area, has been signaled by the detection of the above-mentioned "pick-up gesture", the movement can be ended by a "drop gesture". This forms the counterpart to the pick-up gesture, since the second hand posture is maintained during the movement. For example, as explained in more detail below, the pick-up gesture can indicate a "grasping" by closing the hand, while the drop gesture signals a "letting go" by opening the hand (while it was kept closed during the movement).The movement can then be stopped at a position which is the target position at the time of the drop gesture.
[0027] In some embodiments, when the movement is stopped, the action area is moved to one of several predefined target areas in which the target position lies. In other words, the movement does not end at any arbitrary point, but rather in predefined target areas. By "snapping" to a target area, the user can achieve the result that the gesture does not have to be executed with less precision, and the action area is always moved to a precisely defined position. For example, individual objects such as controls of the graphical user interface can be easily aligned to a grid without the user having to aim at a specific point. It can be provided that the target area is the target area in which at least half of the action area lies when the movement is stopped. If the movement is then ended by releasing, the action area snaps into this target area.This can be particularly advantageous for large display devices with correspondingly large displays, where a large hand movement may be required to achieve a large shift. This can simplify operation and improve user-friendliness.
[0028] In some embodiments, the action area is moved along a path on the display device which is obtained by smoothing the corresponding movement of the hand. In this way, in particular, any trembling of the user's hand during the movement can be compensated for. The movement on the graphical user interface, i.e. the moving of the action area, does not exactly follow the movement of the hand, but takes a tolerance or a threshold value into account. This means that even with shaky or somewhat inaccurate hand movement, a movement occurs without trembling or outliers, which can improve the overall impression during use. It can also be provided that the movement only starts when the movement of the hand has reached a certain threshold value, e.g. a minimum distance or minimum rotation. Overall, a more stable response to user inputs is achieved. This can facilitate operation, in particular, for example,while driving a vehicle, where it can be particularly difficult to perform clean hand movements without shaking or slipping.
[0029] In some embodiments, detecting the gesture comprises determining a position and / or orientation of the hand in the three-dimensional space relative to the display device, wherein the assignment of the picking gesture to determine the action area is performed based on the determined position and / or the determined orientation of the hand. In addition to hand postures and movements, the position and / or orientation of the hand (detected during the movement or during hand postures and, for example, the picking gesture or the placing gesture is recognized) can play a role in detecting corresponding user inputs. In particular, the action area can be determined based on the position and / or orientation of the hand in three-dimensional space.This can be achieved, as explained below, by recognizing a pointing direction or by projecting the hand position onto the display device or other relative (spatial / positional) relationships between the hand and the display device. Intuitive operation can be achieved in this way, for example, by the user moving their hand within a detection area in front of the display device (e.g., inside the vehicle in front of the dashboard) and performing corresponding gestures. By appropriately assigning the recording gesture, in particular, to an action area, a response behavior can be generated that the user would intuitively expect.
[0030] In some associated embodiments, the graphical user interface has at least one control element, wherein a pointing direction of the hand is determined in the first hand position, and a control element corresponding to the determined pointing direction is assigned to the recording gesture as an action area. As just explained, the position and / or orientation of the hand can be used to assign the recording gesture to an action area. In particular, the action area can be a control element that the user intends to move to another location on the graphical user interface. For this purpose, a pointing direction of the hand can be determined in the first hand position, i.e. before the recording gesture is executed. For example, the user can point to a desired control element, e.g. with the index finger or the entire (open) hand. If the hand is then, e.g.closed, the selected control is “grasped” and can be moved by moving the hand and then released at a desired target position or target area.
[0031] In some related embodiments, the movement is defined as a panning or shifting of the hand in the second hand posture in the three-dimensional spatial area along the display device, and the action area is shifted according to a distance traveled by the hand. This movement can be perceived as the natural movement used after a "point and grab" to shift a control element. To do so, the user can, in particular, move their hand (held in the second hand posture) in space, with the control element on the display device then following this movement (possibly by means of a smoothed movement to compensate for any shaking, as described above). In this way, even larger movements and corresponding larger shifts on the user interface can be easily executed.
[0032] In one embodiment according to the invention, the entire graphical user interface or a group of several areas of the graphical user interface is defined as the action area. Therefore, instead of just moving one control element to a new position, it can also be provided to move several areas as a group, e.g. several controls or in particular the entire content of the user interface. Particularly with relatively large displays, it may be desirable not only to rearrange individual elements, but to "grab" and move the entire user interface. The entire display can move in the process. It can also be provided to move an entirety of movable elements, such as all controls, as a group, with a background remaining stationary.
[0033] In one embodiment according to the invention, a pointing direction of the hand is determined in the first hand position, and the entire graphical user interface or a combination of several areas of the graphical user interface is assigned to the capture gesture as the action area if the pointing direction points to an area outside the user interface. Particularly in comparison to a specific selection of a control element to which the user points, pointing "into the void" can signal that, for example, the entire view should be moved. For example, the user can point upwards and then perform the capture gesture (e.g., closing the hand) in order to "grab" the entire displayed content.
[0034] In some associated embodiments, the movement is defined as a rotation of the hand in the second hand position around the axis of the forearm, wherein the area of action is shifted according to a rotation angle. In particular, with an upward-pointing hand position, shifting "around the user" can thus be achieved in a simple manner. In particular, with a closed hand position pointing upwards, the user can rotate the display as desired after, for example, the entire user interface has been defined as the area of action. Instead of a distance traveled by the hand in space, the shifting occurs according to an angle of rotation, e.g. directly proportional. Proportionality can be used to realize fast or slow shifting; for example, if a smaller angle of rotation causes a large shift, this is perceived as fast rotation.However, operation may be easier for some users if the displacement is smaller while maintaining the same angle of rotation.
[0035] In some embodiments, an open hand posture or a hand posture pointing at least one finger is defined as the first hand posture, and a closed hand posture is defined as the second hand posture. This means that the pick-up gesture is described by closing the hand, and the drop gesture is described by opening the hand. This grasping gesture, with a corresponding grasping, moving, and releasing, is easy to understand and thus advantageous for intuitive operation. For example, with a natural gesture, the user can grasp a control element of the graphical user interface, hold it, and move it simultaneously, and release it again at a desired target position, thereby dropping it.
[0036] In some embodiments, a representation of the action area is adjusted upon detection of the picking gesture and, if applicable, also during movement. This allows the user to be given visual feedback as to which area they have grasped as the action area and also whether they are still grasping the area or have already released it. Since the gestures are performed freely in space and the user cannot receive haptic feedback about grasping and holding an element, at least visual feedback can be given by a corresponding change in the representation. For example, a selected control element can be highlighted, for example by enlarging it and / or changing its color or by an animation. If the display device is designed as a 3D display, this can be exploited, for example, by generating a 3D effect when grasping an element.The grasped element (i.e. the action area) can move towards the user when grasping it (when performing the pick-up gesture) by transitioning from a first image plane to a second image plane and can move back again at the target position when dropping it (when performing the drop gesture).
[0037] A second aspect of the invention relates to a data processing system comprising at least one processor configured to execute the method according to the first aspect of the invention. The system also comprises at least one display device configured to display a graphical user interface and a detection device configured to detect a user's hand in a three-dimensional spatial area, in particular a gesture.
[0038] In some embodiments of the system, the display device is designed as a display device for a vehicle and is configured and dimensioned such that it extends at least partially around seats of the vehicle. The display device can extend at least substantially across the width of the vehicle and, for example, be arranged on the dashboard. The display device can also extend further along the sides of the vehicle and in this way surround the driver and front passenger, for example at least 180 degrees in the front area of the vehicle. The display device can also be referred to as a panoramic display. With such a large extension of the display device, it is particularly advantageous if individual control elements of the graphical user interface can be moved according to personal preferences or depending on the situation.As explained above, the entire view of the graphical user interface can also be rotated around the user (e.g. the driver) using an appropriate gesture.
[0039] In some embodiments of the system, the display device is configured to display a first partial representation in a first image plane and a second partial representation in a second image plane spaced apart from the first image plane, such that the first partial representation and the second partial representation are superimposed to produce an overall representation (appearing three-dimensional). The display device can therefore be a 3D display.
[0040] In some embodiments of the system, the display device comprises a first partial display device, a second partial display device, and an optical element, wherein the first partial display device is configured to display the first partial representation, and the second partial display device is configured to display the second partial representation, wherein the optical element is configured to optically superimpose the first and second partial representations. The first and second partial display devices can in particular be arranged at an angle to one another (e.g., approximately 90 degrees) and spaced apart from one another. One of the two partial display devices, in particular the first, can be arranged such that it lies in the field of view of a user (first image plane). The second image plane can be created by mirroring the second partial representation displayed by means of the second partial display device.The optical element can therefore be designed as a reflective transparent element, which can be arranged at an angle between the partial display devices (e.g., at approximately 45 degrees), so that the user looks through the optical element (through which the second image plane is virtually generated) onto the first image plane. This superposition creates the overall display.
[0041] In some embodiments of the system, the detection device comprises at least one image capture device, in particular a camera. Using a camera, the position of a user's hand in three-dimensional space can be easily determined. One or more cameras can be provided. The camera can be an infrared camera. Advantageously, the at least one camera is a time-of-flight (ToF) camera. By using such a 3D sensor device, the position of the hand in three-dimensional space and its movement can be directly detected. 2D sensors can also be combined to detect the position of the hand in three-dimensional space.
[0042] A third aspect of the invention relates to a computer program comprising instructions which, when executed on a system according to the second aspect, cause the system to carry out the method according to the first aspect.
[0043] The computer program can, in particular, be stored on a non-volatile data carrier. This is preferably a data carrier in the form of an optical data carrier or a flash memory module. This can be advantageous if the computer program as such is to be handled independently of a processor platform on which the one or more programs are to be executed. In another implementation, the computer program can be present as a file on a data processing unit, in particular on a server, and can be downloadable via a data connection, for example the Internet or a dedicated data connection, such as a proprietary or local network. Furthermore, the computer program can have a plurality of interacting individual program modules.
[0044] The system according to the second aspect can accordingly comprise a program memory in which the computer program is stored. Alternatively, the system can also be configured to access an external computer program, for example, available on one or more servers or other data processing units, via a communication connection, in particular to exchange data with it that is used during the execution of the method or computer program or that represents outputs of the computer program.
[0045] The features and advantages explained with respect to the first aspect of the invention also apply accordingly to the further aspects of the invention.
[0046] Further advantages, features and possible applications of the present invention will become apparent from the following detailed description in conjunction with the drawings.
[0047] It shows: Fig. 1 schematically shows a display device with two image planes; Fig. 2 schematically shows a structure of a display device; Fig. 3 schematically shows a system for capturing user inputs with a display device and a capturing device; Fig. 4 schematically shows a graphical user interface with several target positions or target areas; Fig. 5 schematically shows a graphical user interface in which a control element is moved; Fig. 6 schematic gestures and hand movements for moving a control element; Fig. 7 schematically shows a graphical user interface, with several controls being moved; and Fig. 8 schematic gestures and hand movements for moving multiple controls.
[0048] Throughout the figures, the same reference numerals are used for the same or corresponding elements of the invention.
[0049] In Fig. 1 schematically shows a display device 10 with two image planes 11, 12. The upper view shows a view of the display device 10, for example, from a user perspective. The lower view illustrates how a user 1 looks at the display device 10, looking through the "virtual" second image plane 12 to the first image plane 11 located behind it. Depending on the user perspective, a three-dimensional effect can be created. This is done, as shown in Fig. 2 is achieved schematically by an optical element 24 which is arranged between two partial display devices 21, 22 which are designed as conventional two-dimensional displays.
[0050] The optical element 24 is designed as a semi-transparent mirror surface, so that the two partial display devices 21, 22 are perceived as two parallel image planes 11, 12, which are arranged at a short distance 13 from one another and directly one behind the other. In this way, an overall representation is created by superimposing a first partial representation, which is displayed by the first partial display device 21, and a second partial representation, which is displayed by the second partial display device 22, which the user 1 perceives three-dimensionally. Since the two partial display devices 21, 22 are not actually arranged parallel to one another, but only the image planes 11, 12 are perceived as such by the user 1, the second image plane 12 in particular can be referred to as "virtual." The perceived distance 13 between the two partial display devices 21, 22 can also be referred to as the virtual distance.Since this is created by an actual distance 23 between the second partial display device 22 and the semi-transparent optical element 24, this distance nevertheless also exists physically.
[0051] Due to the virtual distance 13 between the two image planes 11, 12, a parallax effect occurs. The strength of the parallax effect depends on the user's perspective, i.e., the position of the eyes of user 1, as well as the positioning or alignment of the second (virtual) image plane 12 with respect to the first (real) image plane 11. It should be noted that the displacement resulting from the design of the display device 10 is always undesirable, whereas the displacement resulting from the eye position of user 1 may be desirable depending on the application.
[0052] In addition, a natural parallax effect occurs, which is caused by the distance between the human eyes (and the brain) and which creates the 3D illusion. However, this effect is not the subject of the present invention and must be distinguished from the aforementioned parallax effect, which is caused by the distance 13 between the image planes 11, 12. It is assumed that both image planes 11, 12 are perfectly parallel and that the 2D coordinates of the projection of any target point of one image plane 12 onto the other image plane 11 are known.
[0053] Such a display device 10 designed as a 3D display can also be used advantageously in connection with the present invention. In this case, as in Fig. 2, an element that has been grasped and moved by a grasping gesture can be displayed in a highlighted representation, in which it appears, for example, to be approaching the user. However, a conventional 2D display can also be used. This can, for example, Fig. 3, have a relatively large extension in the width direction. For example, it can be a display for a vehicle that extends over a large part of the width of the vehicle, e.g., in the area of the dashboard, from one side of the vehicle to the other side of the vehicle, and possibly even into the sides of the vehicle.
[0054] In Fig. 3 illustrates a corresponding system for capturing user inputs. It comprises a display device 10 for displaying a graphical user interface (GUI) and a capturing device 30. The system can be installed, in particular, in a vehicle to interact with the vehicle using gesture control and to control corresponding functions, such as infotainment content, air conditioning, lights, etc. It is possible to move corresponding control elements 41 of the GUI 40 within the user interface to a specific location, for example, from a starting position 42 to a target position 43. In this way, the GUI 40 can be adapted to personal needs. Content can also be brought closer to the driver's field of vision, which in turn can increase driving safety, as the driver is less distracted than if viewing and operating content from a distance.
[0055] The detection device 30 forms a hand recognition system that detects a person's hand in real time with a suitable level of detail within a suitable field of view (the detection area 31). This means that the detection device 30 can be capable of detecting fingertip positions in a defined 3D world coordinate system (which may or may not coincide with the sensor coordinate system). Furthermore, it can be provided that a classification is made as to whether an extended (pointing) finger is present or not, i.e., whether a detected hand is currently making a pointing gesture. A control unit (“controller”; not shown) can also be provided that calibrates the system. In particular, the physical position and size of the input surface 1 (in world coordinates) are known to the controller, so that it can use the detection data of a hand (e.g.Finger positions) can be geometrically related to the surface positions. The controller receives the inputs from the recognition system and converts them into suitable inputs for the GUI application. Gesture recognition is performed by software that processes detected hand postures, hand movements, and the position and / or orientation of the hand in space, e.g., calculating the direction of the pointing finger and thus determining an action area, e.g., a control element 41 of the user interface 40.
[0056] In Fig. Figure 4 illustrates the graphical user interface 40 with various target positions 44 and target areas 45. This division or the definition of specific points on the GUI can simplify operation, as moved content can snap there without having to be hit precisely. For example, a control 41 can snap into the central area 45 at the target position 44 as soon as more than half of the control 41 lies in area 45 when the gesture is released. If another control is already placed at this location, this (and other neighboring controls) can move automatically. To create a better user experience, it is generally advantageous if the system does not react too sensitively. For example, a certain tolerance limit should be exceeded when a grasping gesture is recognized before areas or elements of the GUI 40 move.In order to remove a possible tremor of the hand movement, a movement of an action area, such as a control element 41, is smoothed by an algorithm, for example by taking dropouts into account.
[0057] In Fig. Figure 5 illustrates a first possibility for moving an action area 50 of a graphical user interface 40. In particular, a specific control element of the user interface can be grasped and moved in this way.
[0058] Fig. Figure 6 shows a sequence of hand movements that result in a grasping gesture. The sequence begins with an open hand 2 (or a hand with an outstretched index finger), where the user interface control to be moved is selected based on the pointing direction (i). The control can be grasped by closing the hand (into a fist) (ii). More generally, closing the hand forms a grasping gesture, where an action area in the form of a control is determined based on the pointing direction. As long as the hand is kept closed, the grasped area 50 on the user interface 40 can be moved as desired (iii and iv), in particular to the left and right, by moving the hand (in particular essentially horizontally). The target position is determined by stopping the movement (v) and opening the hand (vi), which signals a release of the grasped area 50.
[0059] A second possibility how an action area 51 of a graphical user interface 40 can be pushed is in Fig. 7. In contrast to the first possibility just described, not just one control element is defined as the action area, but several control elements, up to the entire content of the graphical user interface 40. In this way, it is possible for the user to move several control elements at once in combination (as action area 51) or even the entire user interface 40. If the display is, for example, a panoramic display which extends across the width of the vehicle around the vehicle occupants, the content can be rotated around the user (or users) by means of a corresponding hand movement.
[0060] Fig. Figure 8 shows a sequence of hand movements that effect the movement just described, for example, of the entire content of the GUI 40. Instead of pointing at a specific control element, in this case the user's hand 2 initially points upwards in an open position (i). If the hand is closed (ii and iii), this is recognized as a pick-up gesture. However, since a pointing direction points to an area outside the GUI (namely upwards), the action area 51, for example several or all movable controls with or without the background of the user interface, is determined as the area to be moved. The movement then occurs not by a horizontal movement of the closed hand, but by a rotation of the closed hand, which is still pointing upwards. This is shown in Fig. 8 in several directions (iv to vii). Once the desired position is reached, which is also the case with Fig. 4 using the example of a single control element, the user can open their hand (viii and ix). This is recognized as a drop gesture, and the movement is terminated. The movement of the action area 51 occurs according to a rotation angle that the closed hand travels around the axis of the forearm during the movement just described. In this way, carousel-like navigation through a menu is also possible. For example, the view of the GUI 40 can be rotated until a desired menu item is visible in the center.
[0061] While at least one exemplary embodiment has been described above, it should be appreciated that a wide variety of variations exist. It should also be understood that the described exemplary embodiments are merely non-limiting examples and are not intended to limit the scope, applicability, or configuration of the devices and methods described herein. Rather, the foregoing description will provide a guide to implementing at least one exemplary embodiment, with the understanding that various changes in the operation and arrangement of the elements described in an exemplary embodiment may be made without departing from the subject matter as defined in the appended claims, as well as their legal equivalents. LIST OF REFERENCE SYMBOLS 1 user 2 hands 10 Display device 11 first image plane 12 second image plane 13 Distance 21 first partial display device 22 second partial display device 23 Distance 24 optical element 30 Detection device 40 graphical user interface (GUI) 41 Control 42 starting position 43 Target position 44 Target position 45 Target area 50 Action area 51 Area of action
Claims
[1] Method for detecting user inputs in a system having a display device (10) and a detection device (30), wherein a graphical user interface is displayed by means of the display device (10) and a user's hand is detected in a three-dimensional spatial area (31) by means of the detection device (30), the method comprising: - detecting a gesture by means of the detection device (30) as user input, wherein for this purpose at least a hand posture and a movement of the hand (2) are determined; - if, when determining the hand posture as a gesture, a recording gesture is recognized which is defined by a transformation of the hand (2) from a first hand posture to a second hand posture: - determining an action area (50, 51) of the graphical user interface (40) by assigning the recording gesture to at least one area of the graphical user interface (40); - moving the action area (50, 51) of the graphical user interface (40) according to the determined movement of the hand (2) such that the action area (50, 51) on the display device (10) is moved from a starting position (42) as long as the second hand position is determined as the hand position characterized by , that - the entire graphical user interface (40) or a combination of several areas of the graphical user interface is determined as the action area (51), and - a pointing direction of the hand in the first hand position is determined and the recording gesture is assigned the entire graphical user interface (40) or a combination of several areas of the graphical user interface (40) as the action area (51) if the pointing direction points to an area outside the user interface (40). [2] Method according to claim 1, wherein the displacement of the action area (50, 51) of the graphical user interface (40) is stopped at a target position (43) if, when determining the hand posture as a gesture, a deposit gesture is recognized which is defined by a transformation of the hand (2) from the second hand posture to the first hand posture. [3] Method according to claim 2, wherein the action area (50, 51) is displaced when the movement is stopped to one of several predetermined target areas (44, 45) in which the target position (43) lies. [4] Method according to one of the preceding claims, wherein the movement of the action area (50, 51) takes place along a path on the display device (10) which is obtained by smoothing the corresponding movement of the hand (2). [5] Method according to one of the preceding claims, wherein the detection of the gesture comprises determining a position and / or orientation of the hand (2) in the three-dimensional spatial area (31) with respect to the display device (10), wherein the assignment of the recording gesture in order to determine the action area (50, 51) is carried out on the basis of the determined position and / or the determined orientation of the hand (2). [6] Method according to one of the preceding claims, wherein the graphical user interface (40) has at least one control element (41), and wherein a pointing direction of the hand (2) in the first hand position is determined and a control element (41) corresponding to the determined pointing direction is assigned to the recording gesture as an action area (50). [7] Method according to claim 6, wherein the movement is determined as a pivoting or shifting of the hand (2) in the second hand position in the three-dimensional spatial area along the display device (10) and the action area (50) is shifted according to a distance traveled by the hand (2). [8] Method according to claim 1, wherein the movement is determined as a rotation of the hand (2) in the second hand position around the axis of the forearm and the action area (51) is displaced according to a rotation angle. [9] Method according to one of the preceding claims, wherein an open hand posture or a hand posture pointing with at least one finger is determined as the first hand posture and a closed hand posture is determined as the second hand posture. [10] A data processing system comprising at least one processor configured to carry out the method according to any one of the preceding claims, as well as at least one display device (10) configured to display a graphical user interface, and a detection device (30) configured to detect a hand (2) of a user in a three-dimensional spatial area (31). [11] The system of claim 10, wherein the display device is configured as a display device for a vehicle and is designed and dimensioned to extend at least partially around seats of the vehicle. [12] System according to claim 10 or 11, wherein the detection device (30) comprises at least one camera and / or a 3D sensor. [13] A computer program comprising instructions which, when executed on a system according to any one of claims 10 to 12, cause the system to carry out the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Gesture based control of autonomous vehicles
EP3491493B1
Computer vision gesture based control of a device
US20120200494A1
Content gestures
US9009594B2