Method and system for capturing user inputs
The method and system enhance user input in vehicles by using a dual-image plane display to create a three-dimensional effect for control elements, facilitating intuitive and contactless interaction, particularly through gesture control.
Patent Information
- Application Number
- PCT/EP2024/082172
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-30
- Filing Date
- 2024-11-13
- Publication Date
- 2025-06-05
AI Technical Summary
Modern vehicles with multiple and large displays face challenges in providing easy and accessible user input methods, especially for contactless operations, as traditional touch-sensitive controls may not be easily reachable or intuitive.
A method and system that utilize a display device with two image planes to provide visual feedback by adapting the representation of control elements, allowing them to 'migrate' from one image plane to another, creating a three-dimensional effect that enhances user interaction, particularly through gesture control.
The solution facilitates intuitive and contactless user input by providing clear visual feedback, making it easier to operate complex vehicle systems without distracting the driver, especially in scenarios where traditional controls are not accessible.
Smart Images

Figure EP2024082172_05062025_PF_FP_ABST
Abstract
Description
[0001] Method and system for capturing user input
[0002] The present invention relates to a method and a system for detecting user inputs, particularly in a vehicle. In particular, a display device is controlled in response to user inputs.
[0003] Modern vehicles often have a wide range of functions, which usually includes various vehicle or comfort functions, such as settings for the navigation system, the air conditioning, seat settings, lighting settings, and the like. Various functions of an infotainment system can also be operated, such as playing music, making phone calls, and the like. To display and control the functions, at least one display is usually provided as part of the user interface, for example, in the center of the dashboard. This is where the individual functions, menus, and the like can be shown. These displays are often touch-sensitive, so that the desired function can be controlled by touching the display. For this purpose, the controls are designed so that they can be reached and operated with a finger.
[0004] However, modern vehicles often have multiple displays and increasingly larger displays, which may no longer be easily accessible by hand. For example, several distributed displays or a large display that extends across the width of the vehicle, possibly curved, may be provided. It is therefore also known to control certain functions contactlessly using gestures. To do this, a user performs certain predefined gestures in free space, for example, in a spatial area of the vehicle cabin in front of the display. Such gestures can be detected by appropriate sensors that are capable of determining hand postures and movements in three-dimensional space.
[0005] In addition to conventional 2D displays, 3D displays are also known to expand the presentation of content on a display. These displays are capable of making the displayed content appear three-dimensional. However, three-dimensional representations are often limited to a relatively small display. For example, individual objects on the dashboard directly in front of the driver can be highlighted by creating a 3D effect. With such 3D displays, the three-dimensional effect can be created, for example, by combining two 2D displays. For this purpose, a transparent, reflective surface can be provided in the display to create a 3D effect in the viewer's perception from two coordinated two-dimensional representations. The viewer sees two superimposed representations through the reflective surface.
[0006] A challenge with both 2D and 3D displays is the ability to visually separate and highlight specific content. This is especially relevant in vehicles, as these systems can have a large number of different functions, the operation of which should distract the user from driving for as little time as possible. Although 3D displays offer a more advanced visual experience than 2D displays, feedback when operating functions is often just as limited.
[0007] The present invention is based on the object of providing a method for detecting user inputs. In particular, user inputs are to be facilitated by visual feedback.
[0008] This object is achieved according to the teaching of the independent claims. Various embodiments and further developments of the invention are the subject of the dependent claims.
[0009] A first aspect of the invention relates to a method, in particular a computer-implemented method, for capturing user inputs in a system having a display device and a capturing device, wherein the display device is configured to display a first partial representation in a first image plane and a second partial representation in a second image plane spaced from the first image plane, such that the first partial representation and the second partial representation are superimposed to form an overall representation, and wherein the capturing device is configured to capture user inputs. In the method, a graphical user interface is displayed by means of the display device, wherein the graphical user interface has at least one control element, wherein the control element is displayed in the first image plane.A user input is captured by the capture device, and a control element of the graphical user interface to which the user input is directed is determined. A representation of at least the control element to which the user input is directed is then adjusted such that the control element addressed by the user input is at least partially displayed in the second image plane.
[0010] The aforementioned method according to the first aspect is therefore based in particular on the display device being controlled in response to a user input. In particular, the representation of a control element to which the user input is directed is adapted or changed. In this way, the user receives visual feedback, so that in particular contactless user inputs are facilitated, for example through gesture control. The representation is in particular adapted such that the control element "migrates" from the first image plane (at least partially) to the second image plane. This allows a three-dimensional effect, which is possible through the design of the display device, to be exploited for the visual feedback of the user input.Particularly if the second image plane is positioned closer to the user, this can create the effect of a control element activated by a user input appearing visually closer to the user. This can make operation easier, especially on large display devices where parts of the graphical user interface are relatively far away from the user.
[0011] The term "display device" used here refers in particular to a device by means of which content, such as a user interface (GUI), can be graphically displayed. This can in particular be a display or a screen, especially in a vehicle. For the purposes of the present invention, a display device is in particular configured to display an overall representation composed of two partial representations displayed in different, in particular spaced-apart, image planes. The distance between the image planes and the superimposition of the corresponding partial representations can create a three-dimensional effect for the user ("3D display"). One of the two image planes can be referred to as "virtual" if the view of the overall representation results from looking through this "virtual" image plane to an image plane behind it.Technically, this can be solved, for example, by two displays positioned at an angle to each other (e.g., about 90 degrees) with a transparent, reflective optical element in between (e.g., at about 45 degrees).
[0012] The term "user interface" or "graphical user interface" used here refers in particular to a graphical representation of control elements (operating elements) that are linked to a specific function and allow a user to control the function. The user interface ("UI" or "graphical user interface" - GUI) can contain "control elements" ("UI elements") such as input areas, buttons, symbols, buttons, icons, sliders, toolbars, selection menus, and the like, which a user can activate, particularly within the meaning of the present invention, without touching them, in order to control an associated function. The (graphical) user interface can also be referred to as a (graphical) user interface.
[0013] The term "detection device" used here refers in particular to a device that can contactlessly detect objects in three-dimensional space and determine their position. In particular, the detection device can detect and localize a user's hand. For example, optical methods can be used to detect a user's hand in space ("gesture control"). The detection device can consist of one or more parts, depending on which detection area is to be covered. For example, one or more cameras can be provided. Likewise, the detection device can also, for example, detect ("track") a user's eyes in order to capture the user's line of sight as user input.
[0014] The term "three-dimensional spatial area" used here refers in particular to an area that can be described by three-dimensional coordinates. A position in the three-dimensional spatial area has unique three-dimensional coordinates. The coordinate system can be chosen arbitrarily. For example, a coordinate system of the detection device or a coordinate system related to the display device can be selected. In particular, there is a relationship between a detection area of the detection device and the display device in order to be able to assign the position of the hand to corresponding locations on the GUI.
[0015] The term "vehicle" used here refers in particular to a passenger car, including all types of motor vehicles, hybrid and battery-powered electric vehicles, as well as vehicles such as sedans, vans, buses, trucks, delivery vans, and the like. The term "function" used here refers in particular to technical features that may be present in a vehicle, for example, in the interior, to be controlled by a corresponding control system. In particular, these may be functions of the vehicle and / or an infotainment system, such as lighting, audio output (e.g., volume), climate control, telephone, etc.
[0016] The terms "comprises," "includes," "includes," "has," "has," "with," or any other variation thereof, as used herein, are intended to cover non-exclusive inclusion. For example, a method or apparatus that includes or has a list of elements is not necessarily limited to those elements, but may include other elements not expressly listed or that are inherent in such a method or apparatus.
[0017] Furthermore, unless explicitly stated to the contrary, "or" refers to an inclusive "or" and not an exclusive "or." For example, a condition A or B is satisfied by one of the following conditions: A is true (or present) and B is false (or absent), A is false (or absent) and B is true (or present), and both A and B are true (or present).
[0018] The terms "a" or "an" as used herein are defined as "one or more." The terms "another" and "another," and any other variations thereof, are defined as "at least one other."
[0019] The term “plurality” as used here shall mean “two or more”.
[0020] The term “configured” or “set up” to fulfil a specific function (and respective modifications thereof) is to be understood within the meaning of the invention that the corresponding device is already in a design or setting in which it can carry out the function or is at least adjustable - i.e. configurable - so that it can carry out the function after being set accordingly. The configuration can be carried out, for example, by appropriately setting parameters of a process sequence or of switches or the like for activating or deactivating functionalities or settings. In particular, the device can have a plurality of predetermined configurations or operating modes, so that the configuration can be carried out by selecting one of these configurations or operating modes.
[0021] Preferred embodiments of the method are described below, which can be combined with each other as well as with the other aspects of the invention described, unless this is expressly excluded or is technically impossible.
[0022] In some embodiments, the at least one control element is displayed in a standard view (i.e., without a user input directed thereto or without user interaction) in a first shape, wherein the adaptation of the representation of the control element addressed by the user input comprises a transition from the first shape to a second shape. The transition from the first shape to the second shape provides visual feedback as to which control element is currently being addressed by the user input, thereby facilitating operation. The second shape can be substantially similar to the first shape, with only a transition from the first image plane to the second image plane taking place. However, the second shape can, for example, also be larger than the first or represent a new shape, e.g.the first shape can be a two-dimensional object, whereas the second shape can then be a three-dimensional representation.
[0023] In some embodiments, the first shape is formed (only) by a partial representation in the first image plane, and the second shape is formed either (only) by a partial representation in the second image plane or by an overall representation which results from the superimposition of a first partial representation in the first image plane and a second partial representation in the second image plane. In other words, the transition can be created by the addressed control element merely moving from the first image plane to the second image plane and thereby accommodating the user. By superimposing two partial representations, the second shape can also be a shape which appears three-dimensional, as already mentioned above.
[0024] In some embodiments, a user perspective is taken into account during the transition from the first shape to the second shape such that the transition from the first image plane to the second image plane occurs along a direction of execution of the user input to the addressed control element. In this way, a three-dimensional effect in which an addressed control element comes towards the user can be further improved, in particular with relatively large display devices (with a correspondingly large user interface) which, for example, extend across the width of the vehicle in front of the user. The transition then appears even more natural. For example, this can prevent the control element from moving in an unexpected direction, for example to the left or right, instead of coming closer to the user. The “direction of execution” can in particular be a pointing direction which, for example,by the direction of an outstretched index finger or by the direction of the arm. With gaze-based control, the direction of execution can also be the user's gaze toward the control element.
[0025] In some embodiments, the control is displayed in the second shape as long as user input is detected. This may be referred to as a "hover" effect. For example, a control may be highlighted as long as a user holds their hand over the control. As the hand continues to move across the user interface, this effect may then always be applied to a corresponding control, so that a user always receives feedback about which control is currently being addressed by their user input (e.g., hand position).
[0026] In some embodiments, detecting the user input comprises detecting a user input of a first type and a user input of a second type, wherein detecting a user input of the first type causes the display of the control addressed by the user input to be adjusted, and a user input of the second type causes at least a function associated with the control to be controlled. The user input of the first type can, for example, as explained above, merely cause the control to be highlighted. This can, for example, be a pointing gesture. With a user input of the second type, the control can, for example, be activated by means of a specific gesture and an associated function (e.g., in the vehicle) can be controlled. It should be noted that a user input of the second type can require a user input of the first type.
[0027] In some associated embodiments, the second user input causes a further adjustment of the display of at least the addressed control element, wherein the further adjustment comprises a transition from the second shape to a third shape. This can provide further visual feedback. The third shape can, for example, acknowledge the selection of the control element and successful control through a visual effect.
[0028] In some embodiments, detecting the user input comprises detecting a user's hand as the user input using the detection device, detecting a position of the hand in a three-dimensional spatial region to determine the control element to which the user input is directed, and detecting a hand posture to determine a type of user input. By determining the position of a user's hand, gesture control can be implemented.
[0029] In some related embodiments, detecting the user input includes determining a position of a pointer on the graphical user interface, wherein the position of the pointer is determined by projecting the position of the hand onto the graphical user interface along a pointing direction. Projecting along the pointing direction facilitates operation of larger display devices. Compared to a perpendicular projection, it is possible to address a control element to which the user is pointing and which the user actually intends to address based on their pointing direction.
[0030] In some embodiments, a graphical representation of the pointer is displayed at the specific position of the pointer on the graphical user interface, wherein the graphical representation of the pointer comprises a representation of a hand model, wherein the hand model is displayed at the position of the pointer in the detected hand posture. While operation is independent of whether the pointer is actually displayed, a display of the pointer helps the user to recognize where in the user interface they are currently navigating. The hand model, i.e. a "virtual hand", can be continuously updated and displayed based on a detected hand posture. In this way, the user's hand can be displayed on the user interface to represent, in a sense, an extended arm. The user thus receives clear visual feedback on their operation of the user interface.In particular, gesture control, in which the user initially moves their hand in free space without feedback, can be made considerably easier. The representation of the pointer (the hand model) can appear as soon as the start of user input is detected, e.g. the user raises their hand and moves it within a detection area. If the user withdraws their hand, the representation of the pointer can disappear again. By dynamically displaying one's own hand as a hand model, gestures (hand postures) can also be visually fed back. The hand model can be generated from a point cloud that is captured by the detection device. The hand can then be reproduced using a skeleton or mesh representation in order to display the user's hand movements on the display device.
[0031] In some embodiments, detecting the user's hand includes determining whether at least one finger of the hand is pointing toward the display device. In principle, the position of the hand can be arbitrary, and the position of the hand can be determined as described above in the detection range. For a more precise determination, however, it can be determined whether the user is extending a finger, for example the index finger, toward the display device. A user can intuitively use this hand position for operation. It can then be provided to control a function only when the user makes a pointing movement, for example by extending their index finger. If the user does not make a pointing movement, control of a function can be suppressed, since in such a case the user may not intend to control a function. Determining the hand position can therefore improve the accuracy of operation.In addition, the position of a fingertip can be determined with greater accuracy.
[0032] A second aspect of the invention relates to a data processing system, comprising at least one processor configured to carry out the method according to the first aspect of the invention. The system also comprises at least one display device configured to display a first partial representation in a first image plane and a second partial representation in a second image plane spaced apart from the first image plane, such that the first partial representation and the second partial representation, when superimposed, result in an overall representation, and a detection device configured to detect a user input. The display device is, in particular, configured to display a graphical user interface. The detection device is, in particular, configured to detect a user's hand in a three-dimensional spatial area as the user input (gesture control).
[0033] In some embodiments of the system, the display device comprises a first partial display device, a second partial display device, and an optical element, wherein the first partial display device is configured to display the first partial representation, and the second partial display device is configured to display the second partial representation, wherein the optical element is configured to optically superimpose the first and second partial representations. The first and second partial display devices can in particular be arranged at an angle to one another (e.g., approximately 90 degrees) and spaced apart from one another. One of the two partial display devices, in particular the first, can be arranged such that it lies in the field of view of a user (first image plane). The second image plane can be created by mirroring the second partial representation displayed by means of the second partial display device.The optical element can therefore be designed as a reflective transparent element, which can be arranged at an angle between the partial display devices (e.g., at approximately 45 degrees), so that the user looks through the optical element (through which the second image plane is virtually generated) onto the first image plane. This superposition creates the overall display.
[0034] In some embodiments of the system, the detection device comprises at least one image capture device, in particular a camera. Using a camera, the position of a user's hand in three-dimensional space can be easily determined. One or more cameras can be provided. The camera can be an infrared camera. Advantageously, the at least one camera is a time-of-flight (ToF) camera. By using such a 3D sensor device, the position of the hand in three-dimensional space and its movement can be directly detected. 2D sensors can also be combined to detect the position of the hand in three-dimensional space.
[0035] A third aspect of the invention relates to a computer program comprising instructions which, when executed on a system according to the second aspect, cause the system to carry out the method according to the first aspect.
[0036] The computer program can, in particular, be stored on a non-volatile data carrier. This is preferably a data carrier in the form of an optical data carrier or a flash memory module. This can be advantageous if the computer program as such is to be handled independently of a processor platform on which the one or more programs are to be executed. In another implementation, the computer program can be present as a file on a data processing unit, in particular on a server, and can be downloadable via a data connection, for example the Internet or a dedicated data connection, such as a proprietary or local network. Furthermore, the computer program can have a plurality of interacting individual program modules.
[0037] The system according to the second aspect can accordingly comprise a program memory in which the computer program is stored. Alternatively, the system can also be configured to access an external computer program, for example, available on one or more servers or other data processing units, via a communication connection, in particular to exchange data with it that is used during the execution of the method or computer program or that represents outputs of the computer program.
[0038] The features and advantages explained with respect to the first aspect of the invention also apply accordingly to the further aspects of the invention.
[0039] Further advantages, features and possible applications of the present invention will become apparent from the following detailed description in conjunction with the drawings.
[0040] It shows:
[0041] Fig. 1 shows schematically a display device with two image planes;
[0042] Fig. 2 shows schematically a structure of a display device;
[0043] Fig. 3 - 6 each schematically show a display device with a user operating a user interface;
[0044] Fig. 7 schematic different states of a control element during user input; and
[0045] Fig. 8 schematically shows a user interface in different representations during a user input.
[0046] Throughout the figures, the same reference numerals are used for the same or corresponding elements of the invention. Fig. 1 schematically shows a display device 10 with two image planes 11, 12. The upper view shows a view of the display device 10, for example from a user's perspective. The lower view illustrates how a user 1 looks at the display device 10, looking through the "virtual" second image plane 12 to the first image plane 11 behind it. Depending on the user's perspective, a three-dimensional effect can be created in this way. This is achieved, as schematically shown in Fig. 2, by an optical element 24 which is arranged between two partial display devices 21, 22 which are designed as conventional two-dimensional displays.
[0047] The optical element 24 is designed as a semi-transparent mirror surface, so that the two partial display devices 21, 22 are perceived as two parallel image planes 11, 12, which are arranged at a short distance 13 from one another and directly one behind the other. In this way, an overall representation is created by superimposing a first partial representation, which is displayed by the first partial display device 21, and a second partial representation, which is displayed by the second partial display device 22, which the user 1 perceives three-dimensionally. Since the two partial display devices 21, 22 are not actually arranged parallel to one another, but only the image planes 11, 12 are perceived as such by the user 1, the second image plane 12 in particular can be referred to as "virtual". The perceived distance 13 between the two partial display devices 21, 22 can also be referred to as the virtual distance.Since this is created by an actual distance 23 between the second partial display device 22 and the semi-transparent optical element 24, this distance nevertheless also exists physically.
[0048] Due to the virtual distance 13 between the two image planes 11, 12, a parallax effect arises. The strength of the parallax effect depends on the user's perspective, i.e., the position of the eyes of the user 1, as well as the positioning or alignment of the second (virtual) image plane 12 with respect to the first (real) image plane 11. It should be noted that the displacement resulting from the structure of the display device 10 is always undesirable, whereas the displacement resulting from the eye position of the user 1 may be desirable depending on the application.
[0049] In addition, a natural parallax effect occurs, which is caused by the distance between the human eyes (and the brain) and which creates the 3D illusion. However, this effect is not the subject of the present invention and must be distinguished from the above-mentioned parallax effect, which is caused by the distance 13 between the image planes 11, 12. It is assumed that both image planes 11, 12 are perfectly parallel and that the 2D coordinates of the projection of any target point of one image plane 12 onto the other image plane 11 are known.
[0050] 3 to 6 illustrate a display device 10 on which a graphical user interface (GUI) with control elements 43 is displayed. By configuring the display device 10 as explained above, in particular with a transparent, reflective surface as the optical element 24, a two-dimensional object 41 can appear to the user 1 as a three-dimensional object 42. As explained below, user inputs are detected without contact, and the user 1 can receive corresponding optical feedback from the system, thereby facilitating the operation of the user interface. A detection device 30 is provided as a sensor (see FIGS. 1 and 4), which detects a user's hand and determines its position in three-dimensional space. The detection device 30 can, for example, be a 3D sensor, such as a time-of-flight camera, or can comprise optical 2D sensors.Therefore, the following describes the use of gesture control as user input. However, the detection device 30 can, in principle, also be configured to detect the gaze of user 1 by tracking the eyes as user input. Voice input would also be conceivable.
[0051] In the example shown, the detection device 30 forms a hand recognition system that recognizes the hand 2 of the user 1 in real time with a suitable level of detail within a suitable field of view (the detection area 31). This means that the detection device 30 can be capable of recognizing fingertip positions in a defined 3D world coordinate system (which may or may not coincide with the sensor coordinate system). Furthermore, it can be provided that a classification takes place according to hand postures, e.g., whether an extended (pointing) finger is present or not, i.e., whether a recognized hand is currently making another pointing gesture, which is recognized as user input. A control unit (“controller”; not shown) can also be provided, which calibrates the system.In particular, the controller knows the physical position and size of the display device 10 (in world coordinates), allowing it to correlate the recognition data of a hand (e.g., finger positions) with positions on the GUI. The controller receives the inputs from the recognition system and converts them into appropriate inputs for the GUI application.
[0052] The display device 10 with two image planes 11, 12 enables simplified operation in that a control element 43 addressed by a user input is displayed in different shapes by means of an optical transition, i.e., optical feedback is provided. In particular, a transition from a 2D shape (cf. control element 43) to a 3D body (cf. addressed control element 45 in Fig. 6) can assist the user 1 in operating a GUI menu both from close up and from a distance. Such a transition can be combined with all conceivable forms of technical interaction (secondary controls), in particular gesture control, as well as voice assistance, eye tracking, or a brain-machine interface (BMI).
[0053] The user 1, e.g., the driver of a vehicle, can interact with the vehicle via the GUI. For example, they select a control element, e.g., a sound symbol, using a specific form of interaction (gesture). The sound symbol then transforms from a flat 2D symbol into a 3D body. More precisely, a transition from the first image plane 11 to the second image plane 12 (or both image planes 11, 12) creates a 3D effect that is perceived by the human brain. The objects appear to float in front of the display device 10, e.g., the screen of the infotainment system. This technology can be used, for example, in autonomous vehicles. The example shown shows a display device 10 that extends as a curved display in front of the user, for example, across the entire width of the vehicle.Since such systems have areas that are far away from the user 1, in particular cannot be reached by hand, the optical feedback in the form of a transition from a 2D object to a 3D object can be particularly advantageous.
[0054] By using a transition that is coordinated with the two image planes 11, 12, the effect can be achieved in which 2D elements move seamlessly towards the user and appear three-dimensional. To generate the transition, the position of the user 1 relative to the display device or the execution direction of a user input, e.g. the pointing direction of an outstretched index finger or the arm of the user 1, can therefore also be taken into account. For this purpose, the positions of the object on both image planes 11, 12 must be coordinated. This allows the user 1 to clearly see which object is changing its state. The effect of a seamless transition is thus created without a noticeable jump between the image planes 11, 12 being perceived.
[0055] To further simplify operation, the display of a pointer can be provided, which indicates the position of the GUI that is addressed by a current user input (e.g., hand movement). For this purpose, a three-dimensional virtual hand 44 can be displayed, which is generated from the data of the detection device 30 (see below) and thus imitates the actual hand 2 of the user 1, including its movements and hand postures that the user 1 is currently performing. The virtual hand 44 can be displayed on the front partial display device 21 and then optically "floats" (by means of the reflective surface 24) in the second image plane 12 above the GUI, which shows the user 1 that the gesture control is active.
[0056] For example, if user 1 moves his hand 2 from left to right, the virtual hand 44 moves in the same way, so that it exactly mimics the movement of user 1's hand 2 horizontally and vertically. The virtual hand 44 extends the arm of user 1, so to speak. When the user's hand 2 moves out of the field of view 31 of the sensor 30, the virtual hand 44 disappears (partial representation of the virtual hand in Fig. 4). The display of the virtual hand 44 simplifies operation, both from close up and from a distance. Due to the depth image illusion of the two-layer display system, the virtual hand 44 appears to float realistically in 3D as a pointer in the space in front of the display device 10, e.g., in front of the dashboard.
[0057] The data from the detection device 30 can be used to generate the representation of the virtual hand 44. For example, a ToF camera generates a point cloud from which data about the position and orientation of the hand 2 of the user 1 can be obtained. In particular, points in the point cloud are corresponding points on the surface of the user's hand. Artifacts can be reduced by calculating medians or averages. The point cloud of the hand 2 can be separated from the background and used for orientation and gesture control. One way in which the virtual hand 44 can be represented is as a skeleton of the hand, which is generated from the point cloud, for example using the Randomized Hough Transform or the Iterative Closest Point (ICP). Here, the distance of the hand to the finger joints is estimated and used as the basis for the skeleton.One advantage of creating a skeleton is that the skeleton can be used to calculate the angles of hand 2, allowing hand movement to be tracked more precisely. It is also conceivable to automatically create a mesh based on the user's hand and fingers over the skeleton hand, and add colored material based on the user's skin color. This creates a nice-looking virtual 3D avatar hand instead of a point cloud or skeleton. A scaling algorithm can be applied to adjust the size of the virtual hand to the UI elements (controls) and screen size if necessary. It should be understood that all options used to generate the virtual hand (point cloud hand, skeleton hand, 3D avatar hand) can also be used for gesture control and to display the user's precise hand movement.
[0058] The transitions in the display of a control element described above can be advantageously used for user feedback for contactless interaction, such as gesture control. The visual feedback helps in identifying addressed controls on the user interface. A scheme for transitions is illustrated in Fig. 7. A control element can be displayed in a standard view, i.e. in particular without interaction, in a first shape 51, for example as a 2D symbol as explained above. By means of a specific gesture 60 as user input, for example by simple hand movements with an open hand or with an outstretched index finger, an addressed control element, i.e. for example a control element which the user is pointing at, can appear in a second shape 51, for example by a transition to a 3D object as described above. The 3D object can also be animated, for example it can rotate.
[0059] The second shape 51 can exist as long as the user, for example, holds their hand in the direction of the control. To detect such a hover state, the control in the second shape 51 can be further highlighted compared to the first shape 50, e.g. by enlarging it, changing the color or shape of the element. This allows the user to orientate themselves as to where they are on the GUI. If the user removes their hand (hand gesture 61), i.e., ends the operation, the control is displayed again in the first shape 51. The control can be selected (triggered) by a further gesture 62 in order to control an associated function. The control can then appear in a third shape 52. Upon completion of the operation, e.g., by removing the hand (hand gesture 63), the control is displayed again in the first shape 51. Fig.8 shows an example of how control elements 43 can be addressed and triggered by gesture control in order to control corresponding functions. In doing so, they move through the shapes just described, which is made possible in particular by the design of the display device 10 with two image planes 11, 12. Without interaction, the control elements 43 of the GUI 70 appear in the first (rear) image plane 11, e.g., as 2D symbols (i). If the user 1 raises his hand 2 and points to a control element (ii) (cf. hand movement 60), this element emerges as a 3D object (iii) with the help of the second image plane 12. If he moves his hand or a finger further (e.g., swiping or scrolling), the previously highlighted control element retreats (iv) and the next control element is highlighted by a corresponding transition (v). A trigger gesture 62 selects the addressed control element, whereby the other controls can be hidden (vi).A function assigned to the selected control is controlled, and corresponding information can be displayed (vii). If the user removes their hand (hand gesture 63), the control returns to its original form (viii). The user interface 70 reappears with its controls in the standard view (ix).
[0060] In this way, functions of the infotainment system, the air conditioning, music settings, the ambient lighting in the interior, the vehicle lights, communication with other vehicles, and other functions in the interior / exterior of the vehicle can be controlled, as well as the engine start / stop and doors can be locked or unlocked. Control of non-vehicle functions, such as smart home control, is also conceivable. Some scenarios are described below as examples to illustrate the features explained above.
[0061] A user is sitting in their semi-autonomous vehicle. The seats are more than an arm's length away from the infotainment system display. They want to see the vehicle's status. Using a specific technical interaction type (e.g., gesture control), they select a displayed 2D vehicle icon. This transforms it into a 3D vehicle that hovers in front of the infotainment system display. Using another command (e.g., swipe gestures), they make the vehicle rotate 360 degrees. They see that their lights are not turned on, that their electric vehicle's power is only half full, and the driver's door is not fully closed. With another command, the lights come on and the door closes completely. The 3D vehicle view changes accordingly. The driver of a vehicle has heart problems and wants to check his values.Using a specific form of interaction, such as gesture control, he selects a corresponding 2D heart symbol on the GUI, which then appears in 3D and transmits his health values to him.
[0062] A user wants to hear a specific piece of music through their vehicle's audio system. They move their hand over a sound icon, which then appears in three dimensions and animated from the GUI. They select the piece of music by selecting the sound icon and swiping. A virtual hand hovers in front of the screen as a pointer, helping them orient themselves in the gesture execution space.
[0063] While at least one exemplary embodiment has been described above, it should be appreciated that a wide variety of variations exist. It should also be understood that the described exemplary embodiments are merely non-limiting examples and are not intended to limit the scope, applicability, or configuration of the devices and methods described herein. Rather, the foregoing description will provide a guide to implementing at least one exemplary embodiment, it being understood that various changes in the operation and arrangement of the elements described in an exemplary embodiment may be made without departing from the subject matter as defined in the appended claims, as well as their legal equivalents.
[0064] LIST OF REFERENCE SYMBOLS
[0065] I User
[0066] 10 Display device
[0067] II first image plane
[0068] 12 second image plane
[0069] 13 Distance
[0070] 21 first partial display device
[0071] 22 second partial display device
[0072] 23 Distance
[0073] 24 optical element
[0074] 30 Detection device
[0075] 41 2D symbol
[0076] 42 3D objects
[0077] 43 Control
[0078] 44 virtual hands
[0079] 45 addressed control element
[0080] 50 first figure
[0081] 51 second figure
[0082] 52 third figure
[0083] 60 User input / gesture (hover)
[0084] 61 User input / gesture (exit)
[0085] 62 User input / gesture (trigger)
[0086] 63 User input / gesture (exit)
[0087] 70 User Interface (GUI)
Claims
CLAIMS 1. A method for capturing user inputs in a system having a display device (10) and a capturing device (30), wherein the display device (10) is configured to display a first partial representation (41) in a first image plane (11) and a second partial representation (42) in a second image plane (12) spaced apart from the first image plane (11), such that the first partial representation (41) and the second partial representation (42) are superimposed to form an overall representation (40), and wherein the capturing device (30) is configured to capture user inputs, the method comprising: Displaying a graphical user interface (70) by means of the display device (10), wherein the graphical user interface has at least one control element (43), wherein the control element (43) is displayed in the first image plane (11); Detecting a user input (60) by means of the detection device (30), wherein a control element (43) of the graphical user interface is determined to which the user input (60) is directed; and Adapting a representation of at least the control element (43) to which the user input (60) is directed such that the addressed control element (43) is at least partially displayed in the second image plane (12).
2. The method according to claim 1, wherein the at least one control element (43) is displayed in a standard view in a first shape (50), wherein adapting the display of the control element (43) addressed by the user input (60) comprises a transition of the first shape (50) into a second shape (51).
3. The method according to claim 2, wherein the first shape (50) is formed by a partial representation in the first image plane (11), and the second shape (51) is formed either by a partial representation in the second image plane (12) or by an overall representation which results from superimposing a first partial representation in the first image plane (11) and a second partial representation in the second image plane (12).
4. The method according to claim 2 or 3, wherein during the transition from the first shape (50) to the second shape (51), a user perspective is taken into account such that the transition from the first image plane (11) to the second image plane (12) takes place along an execution direction of the user input (60) on the addressed control element (43).
5. The method according to any one of claims 2 to 4, wherein the control element (43) is displayed in the second form (51) as long as the user input (60) is detected.
6. The method according to any one of the preceding claims, wherein detecting the user input comprises detecting a user input of a first type (60) and a user input of a second type (62), wherein detecting a user input of the first type (60) causes the display of the control element (43) addressed by the user input to be adjusted, and a user input of the second type (62) causes at least one function associated with the control element (43) to be controlled.
7. The method according to claim 6 and one of claims 2 to 5, wherein the user input of the second type (62) causes a further adaptation of the representation of at least the addressed control element (43), wherein the further adaptation comprises a transition from the second shape (51) to a third shape (52).
8. The method according to any one of the preceding claims, wherein detecting the user input comprises detecting a hand (2) of a user (1) by means of the detection device (30) as the user input (60), wherein a position of the hand (2) in a three-dimensional spatial area is detected in order to determine the control element (43) to which the user input (60) is directed, and a hand posture in order to determine a type of user input.
9. The method of claim 8, wherein detecting the user input comprises determining a position of a pointer on the graphical user interface (70), wherein the position of the pointer is determined by projecting the position of the hand (2) onto the graphical user interface (70) along a pointing direction.
10. The method of claim 9, wherein a graphical representation of the pointer is displayed at the determined position of the pointer on the graphical user interface (70), wherein the graphical representation of the pointer comprises a representation of a hand model (44), wherein the hand model is displayed at the position of the pointer in the detected hand posture.
11. The method of claim 10, wherein the hand model (44) is continuously updated and displayed based on a detected hand posture.
12. A system for data processing, comprising at least one processor configured to carry out the method according to one of the preceding claims, and at least one display device (10) configured to display a first partial representation in a first image plane (11) and a second partial representation in a second image plane (12) spaced apart from the first image plane (11), such that the first partial representation and the second partial representation are superimposed to form an overall representation, and a detection device (30) configured to detect a user input (60, 61, 62, 63).
13. The system according to claim 12, wherein the display device (10) comprises a first partial display device (21), a second partial display device (22) and an optical element (24), wherein the first partial display device (21) is configured to display the first partial representation and the second partial display device (22) is configured to display the second partial representation, wherein the optical element (24) is configured to optically superimpose the first and second partial representations.
14. System according to claim 12 or 13, wherein the detection device (30) comprises at least one camera and / or a 3D sensor.
15. A computer program comprising instructions which, when executed on a system according to any one of claims 12 to 14, cause the system to carry out the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Method and system for displaying information
US20100201623A1
Multi-layered vehicle display system and method
US20140282182A1