User interface control

By obtaining user background data and switching input types, the problem of inconvenient interface control in high-pressure environments is solved, and more flexible and secure user interface interaction is achieved, adapting to user input needs in various scenarios.

CN120359488APending Publication Date: 2025-07-22KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380086331.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-16
Filing Date
2023-12-04
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In high-pressure and high-pressure environments, existing user interface interaction methods such as gestures and voice control are not applicable in some scenarios, resulting in inconvenient interface control, especially in sterile environments or noisy environments, which affects the user's decision-making efficiency and safety.

Method used

By obtaining the user's background data, it is determined whether the user can provide the main input type, such as gestures, voice commands, etc. If it is not available, switch to the secondary input type, such as foot posture, leg posture, etc., and use the user tracking device and processor to detect and switch the input type.

Benefits of technology

It provides a flexible interface control method to adapt to user input needs in different contexts, improves the convenience and security of user interaction, and avoids accidental control and user frustration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120359488A_ABST
    Figure CN120359488A_ABST
Patent Text Reader

Abstract

The proposed concepts are directed to providing schemes, solutions, concepts, designs, methods and systems relating to obtaining user input to control an interface. Information describing a context of a user is utilized to determine whether the user can provide a primary input type. If they are possible, a primary type of input from the user is detected. If not, a secondary type of input from the user is detected. Therefore, control input of the user to the interface can be detected more appropriately based on the surrounding environment of the user and / or the state of the user. In this way, an easier and more robust interface control means may be provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of user interfaces, and more particularly to the field of obtaining user interface control inputs. Background Art

[0002] An important consideration in interface design is the ease of interacting with and controlling the interface. Some interfaces are interacted with in high-stress, high-pressure environments where the intuitiveness of the interface is crucial so as not to impede instant decision-making. In particular, for facilitating effective clinical decision-making in a healthcare environment, the interaction must not consume more cognitive effort than absolutely necessary.

[0003] Generally, users interact with the interface using gestures, voice control, eye gaze direction, etc. Such ways of interacting with the interface are intuitive and convenient for users. However, in certain backgrounds / scenarios, some interaction modes may not be appropriate. For example, the use of a touch screen may be excluded in a sterile environment (i.e., an operating room or an intensive care unit), and voice interaction is inconvenient and often ineffective in a noisy environment.

[0004] Therefore, in certain scenarios / backgrounds, concepts for improving user control of the interface are needed. Summary of the Invention

[0005] The present invention is defined by the claims.

[0006] According to an example of one aspect of the present invention, a method for obtaining an input from a user to control an interface is provided, the method comprising:

[0007] obtaining background data describing the background of the user, the background data including information describing an attribute of the user or the environment around the user;

[0008] processing the background data to determine whether the user is capable of providing a primary input type;

[0009] in response to determining that the user is capable of providing the primary input type, detecting the input of the primary type performed by the user; and

[0010] in response to determining that the user is not capable of providing the primary input type, detecting a different secondary type of input performed by the user.

[0011] The primary input type includes at least one of a gesture, an eye gaze direction, a voice command, and a controller interaction. The secondary input type includes at least one of a foot posture, a leg posture, an arm posture, and a head posture.

[0012] The proposed concept aims to provide solutions, concepts, designs, methods, and systems related to obtaining user input for controlling an interface. The background data of the user is utilized to determine whether the user can provide a primary input type. If so, the input of the primary type from the user is detected. If not, the input of the secondary type from the user is detected. Thus, the control input of the user to the interface can be more appropriately detected based on the user's surroundings or the user's state. In this way, a more intuitive and less intrusive way of interface control can be provided.

[0013] The present invention provides for switching between a first interface control input (i.e., gesture, voice command, etc.) and a second interface control input (i.e., foot posture) in response to determining that the user cannot currently perform an interface control input of a first input type. For example, the user may be dealing with another device / item and thus cannot provide the primary input type (i.e., gesture). As another example, the user may be in a noisy environment, in which case it can be determined that the user cannot provide a voice command as the primary input type. In such a case, the present invention determines this based on the background information of the user and switches to detecting the secondary input type that the user can provide.

[0014] Although the primary input type may generally be convenient and intuitive for the user, this may not always be the case. The present invention is based on this recognition, as it is beneficial to utilize the secondary input type in scenarios / backgrounds where the user cannot provide the primary input type. Thus, although the secondary input type may not be as effective and / or intuitive as the primary input type, providing the secondary type of input is more preferable for the user in certain backgrounds.

[0015] In fact, the primary type of input is a typical / conventional mode of interacting with an interface with which the user may be familiar, thus providing an intuitive way of interface control. In fact, if the user is in a background where they can provide an input of one of the above types, this is likely to be a preferred control mode (i.e., due to ease and the level of complexity of the input that the user can provide).

[0016] However, if the user is in a background where they cannot perform the primary type of input, they are more likely to use their primary limb to perform the input. This is generally because the tasks / situations that restrict the user's ability to interact with the interface / provide input typically require the use of the user's hands / fingers, voice, and / or gaze. The primary limb is generally not used for tasks / restricted by tasks that limit the user's interaction with the interface in a traditional way (i.e., gesture, voice control, etc.). Thus, these input types / input modalities may provide a convenient fallback for users who wish to provide input to the interface.

[0017] Accordingly, the present invention can provide a way to obtain input to an interface that is more flexible for the various different contexts in which a user may be. The user may particularly benefit when controlling the interface in high-pressure situations where interface controls need to be adjusted.

[0018] In addition, two categories of information that can determine whether a user can provide input in a certain way include attributes (i.e., the user's status, condition, characteristics) and the surrounding environment (i.e., noise, temperature, placement of items, other people). These factors determine whether and to what extent the user can perform an action that results in input for controlling the interface. Therefore, a more effective determination of whether to detect a first input type or a second input type can be obtained by generating context data that includes this information.

[0019] More specifically, the context data can include information describing at least one of the following: the user's location, the user's condition, the task the user is performing, the user's gaze direction, the ambient noise around the user, the user's location, and other people near the user.

[0020] The foregoing information can be particularly useful for accurately determining whether and to what extent the user can provide input of the primary type (and / or secondary type).

[0021] In some embodiments, obtaining the context data can be performed in response to the user providing at least one of a primary input type and a secondary input type.

[0022] By collecting information when the user performs an action related to an input of the primary type or the secondary type, the context data can be up-to-date.

[0023] In some embodiments, the secondary input type can be user-selectable.

[0024] The user is likely to know the secondary (i.e., fallback) input type, which will be convenient for them in the future (i.e., which input types they may expect to be difficult / impossible for them to perform). Therefore, by allowing the user to be able to select such an input type, the convenience of the user's interaction with the interface is improved.

[0025] More specifically, the method can further include receiving an indication from the user of a body part; and determining the secondary input type based on the indication from the user.

[0026] One way to select the secondary input type is to use a similar system that detects user input to retrieve an indication of the body part (i.e., foot, arm, leg, head, etc.) that the user wishes to select. This can be intuitive and thus inherently easy for the user to perform.

[0027] In some embodiments, obtaining background data may include obtaining one or more images including at least a portion of the user from one or more cameras of a user monitoring system, and analyzing the one or more images to extract the background data.

[0028] A useful way to determine a user's background is through an image (or stream of images) of the user in order to determine their surrounding environment. This can provide much of the information listed above, which may be particularly useful for determining whether a user is able to provide a particular input type. Known image processing techniques can be employed to accurately evaluate such images in order to quickly and accurately generate background information about the user.

[0029] In further embodiments, in response to determining that the user cannot provide a primary input type, the method may further include transmitting a notification to the user, prompting a switch of input type, and receiving confirmation from the user.

[0030] By retrieving user confirmation before continuing to detect input of the secondary type (and thus potentially controlling the interface), unexpected control of the interface can be avoided. Unconfirmed automatic switching may lead to misuse of the interface, user frustration, and potential security and safety risks (if it occurs in the wrong scenario). Therefore, it is desirable to avoid situations where the input type used to control the interface is controlled without the user's knowledge and consent.

[0031] In some embodiments, detecting an input of the primary type or an input of the secondary type may include obtaining a signal including the input of the primary type or the secondary type, and analyzing the signal to extract the input of the primary type or the secondary type. The signal may include at least one of an image signal, a sound signal, a pressure signal, or a motion sensor signal.

[0032] Images, sounds, pressure, and motion all provide information that helps to determine the action performed by the user, which can be considered an input of the primary or secondary type. Thus, obtaining at least one signal including the above information enables accurate detection of the user's input.

[0033] The method may further include determining whether the user is able to provide multiple primary input types by processing the background data, and wherein detecting the input of the secondary type is performed in response to determining that the user cannot provide each of the multiple primary input types.

[0034] There may be many different types of primary (i.e., initial) input types available for controlling an interface. For example, an interface may interact using a controller and voice control, such as the interface of a TV. Thus, it is preferable to switch to detecting input of the secondary type only when the user is unable to use all primary input types.

[0035] In some embodiments, the interface may be based on augmented reality or virtual reality.

[0036] According to a further example of one aspect of the present invention, there is provided a computer program comprising computer program code modules which, when the computer program is run on a computer, are adapted to implement a method for obtaining input from a user to control an interface according to an embodiment.

[0037] According to a further example of one aspect of the present invention, there is provided a system for obtaining input from a user to control an interface, the system comprising:

[0038] a user tracking device configured to monitor the user to generate background data describing the background of the user, wherein the background data includes information describing the attributes of the user or the environment surrounding the user;

[0039] a processor configured to process the background data to determine whether the user is capable of providing a primary input type, and

[0040] wherein the user tracking device is further configured to:

[0041] in response to determining that the user is capable of providing the primary input type, detect the primary type of input performed by the user, wherein the primary input type includes at least one of a gesture, an eye gaze direction, a voice command, and a controller interaction; and

[0042] in response to determining that the user is not capable of providing the primary input type, detect the secondary type of input performed by the user, wherein the secondary input type is different from the primary input type, and wherein the secondary input type includes at least one of a foot pose, a leg pose, an arm pose, and a head pose.

[0043] These and other aspects of the present invention will become apparent and be elucidated with reference to the embodiments described hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to better understand the present invention and to more clearly show how the present invention may be implemented, reference will now be made, by way of example only, to the accompanying drawings in which:

[0045] Figure 1 a flowchart of a method for obtaining input from a user to control an interface according to an embodiment of the present invention is shown;

[0046] Figure 2 a flowchart of a method for determining a secondary input type according to aspects utilized in some embodiments of the present invention is shown; and

[0047] Figure 3 A simplified block diagram of a system for obtaining input from a user to control an interface according to another embodiment is shown; and

[0048] Figure 4 is a simplified block diagram of a computer in which one or more portions of the embodiments may be employed. DETAILED DESCRIPTION

[0049] The present invention will be described with reference to the accompanying drawings.

[0050] It should be understood that the detailed description and specific examples, while indicating exemplary embodiments of the apparatus, system and method, are for illustrative purposes only and are not intended to limit the scope of the invention. These and other features, aspects and advantages of the apparatus, system and method of the invention will be better understood through the following description, claims and drawings. The fact that certain measures are cited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.

[0051] It should be understood that the drawings are schematic and not drawn to scale. It should also be understood that the same reference numerals are used throughout the drawings to indicate the same or similar components.

[0052] The present invention proposes a concept for obtaining input from a user to control an interface. In particular, the present invention provides a determination as to whether a user is able to provide a primary input type. If the user is not able, a secondary type of input is detected instead of the primary type of input. In this way, when the user is in a specific scenario / context where they cannot (or are very inconvenient to) provide a primary type of input, they may be able to control the interface in a simple manner.

[0053] This may be particularly beneficial for high-stress and high-pressure scenarios, where the user may not always be able to provide primary types of input. For example, when the user is a caregiver treating a patient, using gestures (due to sterility issues or equipment handling) or voice commands (due to the large number of people interacting around them) may not be appropriate or possible. Additionally, in some situations, the pilot may not be able to remove their hands, so gestures may not be appropriate (or may sacrifice their control of the aircraft).

[0054] Thus, the proposed invention provides a way to provide input to an interface that is more robust in a variety of different contexts. While a primary input type may be preferred due to the intuitiveness and effectiveness of interacting with the interface, a secondary input type may be used in scenarios that effectively reduce the ease and effectiveness of using the primary input type.

[0055] For a specific example, one challenge in the healthcare field is to design an interface that is easy to use in situations where users (e.g., caregivers) are often under a lot of stress. The interface should be intuitive and should not interfere with users when they need to make instant decisions that may affect the treatment outcome, thus affecting the patient's health. An additional complexity is that in some situations (e.g., operating rooms or intensive care units), maintaining sterility is crucial. Therefore, some traditional input modalities for controlling / interacting with the interface may not be ideal (or even impossible) to use.

[0056] Augmented / virtual reality interfaces address some of the problems in the healthcare field because they provide a way to visualize information that can be interacted with in various different ways. Common interface input types include gesture / hand tracking, voice commands, gaze tracking, and / or remote control.

[0057] However, even the input modalities of these interfaces may not be sufficient. This is especially true in the context of minimally invasive surgical spaces. Voice control can be cumbersome to use and distracting (i.e., in an operating room where many people are interacting with each other). Gesture / hand tracking may not work in scenarios where the user is controlling other devices - the so-called "hands-busy" scenarios (i.e., when using a catheter, scalpel, etc.). Gaze interaction is not ideal for controlling complex menu structures, which can lead to unexpected user actions and / or eye fatigue. In fact, the user may be looking at other objects / systems outside of the interface, so diverting the gaze to interact with the interface can be distracting (i.e., it may be a cross-reference between the patient and the clinical picture).

[0058] Therefore, in acute working conditions, where users are already performing tasks using their hands, gaze, and voice, an additional control interface is needed if the standard hand, gaze, and voice controls are not available.

[0059] Therefore, it is proposed to monitor the users of the interface to collect background information that can be used to determine the input modality to detect and for controlling the interface. Such background data can include information describing the attributes of the user (e.g., the user's location, the user's condition) or the environment around the user (e.g., ambient noise, number of people, proximity of input controls).

[0060] In a specific embodiment, the present invention can provide: detecting the user's foot movement (i.e., a secondary input modality) to control the interface when the user is unable to use the primary input modality. For example, if the user's surrounding environment is too noisy for voice control, the user is operating another device with their hands, and / or the user is looking at a specific medical image, the control input will automatically switch to the user's foot. In other words, the user's foot (or another limb) can be tracked because the original controls of the interface are not available (a fallback scenario where the user is not near a physical button).

[0061] Of course, the advantages associated with the present invention are not limited to the healthcare field. For example, the present invention may be advantageous in any situation where the user is unable / cannot provide input using the primary input modality (e.g., cannot operate a mouse / keyboard / touch screen with their hands). This includes, but is not limited to, the fields of aviation, mobility, and accessibility.

[0062] In some embodiments, the user may be able to select their own secondary input type / modality. For example, during user interaction with the interface, the user can mark a body part as the secondary / backup interaction part.

[0063] Specifically, during interaction with the interface, the user is tracked to determine input using the primary input type(s), e.g., what the user is doing with their gaze, hands, and / or voice. This information is used to evaluate whether the user is using some of these interaction parts to interact with something other than the interface (i.e., talking to another person, looking at another object, or handling a device). If the primary input type is determined to be unavailable based on the user's context, then input of the secondary type(s) is detected. In this case, this can be the marked body part (such as a foot, arm, etc.).

[0064] Advantageously, embodiments can provide the user with a notification that the secondary input type has been activated as an interaction part / input. The user may need to confirm the use of one of the primary input types in a simple way (e.g., a small hand twitch, a brief gaze movement to a certain location, a brief audible sound). This may be important for avoiding unintentional interaction with the interface, which may cause frustration and pose safety issues. For example, when performing a foot movement (e.g., foot movement as the secondary input type) (or before performing a foot movement), the user may need to look at a specific part of the screen (e.g., gaze direction as the primary input type).

[0065] The main aspects of embodiments of the present invention include:

[0066] (i) A system for tracking the user. This can be camera - based and thus capable of tracking the skeletons (and postures) of multiple users in a physical space. Different users can be authenticated to detect who is the primary user using the interface (e.g., wearing an AR headset). Thus, data can be provided that helps to understand the user's context.

[0067] (ii) A system for detecting the input(s) of the primary type(s). For example, the system can be capable of tracking the user's hand and the user's eye gaze and using these as input for (initially) controlling the interface.

[0068] (iii) A system for detecting one or more inputs of a secondary type. For example, the system is capable of tracking foot movements, arm movements, and / or leg movements. This can be in the form of a camera system or a pressure sensor.

[0069] (iv) A memory for recording one or more secondary input types marked by the user.

[0070] (v) A processor that analyzes data from the tracking system and optional additional information to determine when to switch between the primary input type and the secondary input type.

[0071] Continuing, Figure 1 A flowchart of a method for obtaining input from a user to control an interface according to an embodiment of the present invention is presented. The interface can be based on augmented reality (AR) or virtual reality (VR). Alternatively, the interface can be any interface of a device that the user can control / read.

[0072] In step 110, background data describing the background of the user is obtained. The background of the user can be the user's surrounding environment, the user's condition, the user's state, or any other information describing the situation of the user when providing input to the interface. In fact, anything about the user and / or their surrounding environment before or during providing input to the interface can be considered part of the user's background.

[0073] As a result, the background data can include information describing the attributes of the user or the user's surrounding environment. More specifically, the background data can include information describing at least one of the following: the user's location, the user's condition, the task the user is performing, the user's gaze direction, the ambient noise around the user, the user's location, and other people near the user. However, the background data is not limited to this and can include any information that can be used to determine the user's ability and / or expectation to provide a specific type of input. Those skilled in the art will easily understand and appreciate these factors.

[0074] In some embodiments, step 110 can include optional sub-steps 112 and 114. In step 112, one or more images of the user are obtained from one or more cameras of the user monitoring system. These images can be obtained from a storage device or can be directly captured by an imaging device. Such images can include at least a part of the user so that background data can be extracted from the images.

[0075] In fact, in step 114, one or more images are analyzed to extract background data. This can be achieved using known image processing methods capable of extracting data describing the image from the image itself. Those skilled in the art will easily know and understand such processing methods.

[0076] Step 110 can be performed in response to the user providing at least one of a primary input type and a secondary input type. In other words, background data can be obtained when the user provides input to the interface, whether by performing an action of the first type (i.e., a gesture) or an action of the secondary type (e.g., a foot movement). Thus, the background related to the user input is obtained.

[0077] Move to step 120 and then process the background data to determine whether the user is able to provide the primary input type. In essence, this means considering the user / user's surrounding environment to determine whether the user can or is suitable to provide the first type of input.

[0078] As a simplified example, the background data can indicate that the user is in a "hands - busy" scenario (e.g., an image captured and analyzed after / during the user input shows the user performing a task with their hands). If the primary input type is a gesture type, it can be determined in step 120 that the user cannot provide the primary type of input without significantly disturbing the user. In other words, the user cannot provide the primary input type.

[0079] As another example, the background data can indicate that the user is in a noisy environment (e.g., the image can show the user surrounded by other people, or the speaker can pick up excessive background noise). In this case, if the primary input type is a voice command type, it can be determined that the user cannot provide the primary type of input without disturbing the surrounding people, or such an input type may not be detected by the corresponding input peripheral / microphone.

[0080] Of course, there are many other scenarios in which the background data will indicate whether the user can (or cannot) perform an action to provide the primary type of input. Those skilled in the art will readily understand these situations. Thus, for the sake of brevity, an exhaustive list of combinations of the primary input type and the background data, and whether the user can provide the primary type of input given the background data, is omitted here.

[0081] In response to determining that the user is able to provide the initial input type, the method proceeds to step 130. Alternatively, in response to determining that the user is not able to provide the initial input type, the method advances to step 140.

[0082] In step 130, the primary type of input performed by the user is detected. To link to the above, when it is determined that the user is able to perform an action of the primary type and thus perform the primary type of input, the primary type of input is detected. Then, this input can be used to control the interface (i.e., control the interface to present certain information, switch modes, etc.).

[0083] The detection can be performed by any known method of detecting the main type of input, which will depend on which type of main input is selected. For example, if the main input type is the eye gaze direction, a camera will be used to detect the eye gaze direction. When the main input type is a voice command, a microphone will be used to detect the voice command.

[0084] In step 140, different secondary types of input performed by the user are detected. Contrary to step 130, when it is determined that the user cannot (or will be distracted, interfere, or otherwise be inappropriate) provide the main type of input, an input of a secondary input type (different from the main type) is detected. This input can then be used to control the interface instead of the main type of input.

[0085] In some cases, the main input type can include at least one of gestures, eye gaze direction, voice commands, and controller interactions. Additionally, the secondary input type can include at least one of foot postures, leg postures, arm postures, and head postures.

[0086] Of course, the various different combinations of the above exemplary main types and secondary types will be advantageous in different scenarios. In a surgical environment where the user is a surgeon, the main input type may be gestures, but when the surgeon starts the operation (i.e., the hands are busy), this may switch to foot postures. In an aviation environment, controller interaction may be the main input type for a pilot, but when the pilot is away from the controls, this may switch to a secondary input type such as head postures or even voice control.

[0087] Furthermore, there may be multiple main input types and multiple secondary input types. In such a case, the input control method may switch to the secondary input type only when each main input type is unavailable to the user. Taking the surgical scenario as an example again, the main input types can be gestures, eye gaze direction, and voice control. However, if the surgeon is currently performing surgery, looking at the patient, and there are many medical support staff around, it may be determined to switch to a secondary input type such as foot postures.

[0088] Steps 130 and 140 may also include (optional) steps 132 / 144 and 134 / 146. In step 132 / 144, a signal containing the input of the main type or secondary type is obtained. The signal can include at least one of an image signal, a sound signal, a pressure signal, or a motion sensor signal. For example, if the input is a voice command, a sound signal can be obtained. If the input is a gesture, an image signal can be obtained. If the input is a controller input, a motion sensor signal can be obtained.

[0089] At step 134 / 146, the signal is analyzed to extract the input of the primary type or the secondary type. Thus, the extracted input can be used to control the interface.

[0090] Finally, method 100 may further include step 142. Step 142 may be performed in response to determining that the user cannot provide the primary input type. In this case, the user may be notified of the input type switch. In other words, a notification is transmitted to the user, prompting the input type switch. This may be in the form of an audio, visual, or tactile signal, such that the user is aware of the switch.

[0091] In this case, the user may need to confirm the switch in order to obtain the secondary type of input. In fact, this may be to prevent unexpected input from being provided to the interface. Alternatively, the user may simply be notified and the secondary type of input may be obtained.

[0092] Figure 2 A flowchart of method 200 for determining a secondary input type, which utilizes aspects according to some embodiments of the present invention, is presented. In fact, in some embodiments, the secondary input type may be selected by the user. More specifically, the secondary input type may be attributed to a body part, and in the case where the primary type of input is not available, the user may move the body part to provide input.

[0093] Specifically, at step 210, an indication of the user's body part is obtained. This may be achieved by the user manually inputting the body part via the user interface (i.e., via a drop-down list or a text box). Alternatively, the user may tap, touch, or otherwise mark the body part, and the camera may detect this.

[0094] Then, at step 220, the secondary input type may be determined based on the indication from the user. For example, if the user taps, touches, or otherwise marks their body near their foot area, the secondary input type may be determined as a foot posture / movement.

[0095] Of course, alternative methods of selecting the secondary input type or indeed the body part as the secondary input means will be readily understood. This gives the system a degree of flexibility, thus allowing the user to select the secondary input type that they consider most appropriate in the case where the primary input type is not available.

[0096] Figure 3 A simplified block diagram of system 300 for obtaining input from a user to control interface 330 according to another embodiment is presented. The system includes a user tracking device 310 and a processor 320.

[0097] The user tracking device 310 is configured to monitor the user to generate background data describing the background of the user. Thus, the user tracking device can obtain background data through a camera, a microphone, or other sensing devices in order to obtain the background data as described above.

[0098] The processor 320 is configured to process the background data to determine whether the user is able to use the method as described above to provide the primary input type. The processor can be part of the same device as the user tracking device or can be provided remotely.

[0099] In addition, the user tracking device 310 is configured to detect the primary type of input performed by the user in response to determining that the user is able to provide the primary input type. The user tracking device 310 is also configured to detect the secondary type of input performed by the user in response to determining that the user is unable to provide the primary input type, where the secondary input type is different from the primary input type. This can be performed using the same sensors used to monitor the user to obtain background information or can be performed by a separate device.

[0100] Of course, the above system can be adapted to perform any aspect of the method described in Figure 1 including or excluding optional steps.

[0101] Through an illustrative medical scenario, the system 300 can operate in the following manner:

[0102] (i) A doctor (i.e., the user) can enter the operating room and can pick up the head-mounted device worn during the surgery. The user interface and / or another application may ask them to select a body part as their fallback / secondary interaction branch to provide input in case the primary type of input is not available.

[0103] (ii) The doctor can look down, tap, and hold their right leg, for example, to confirm that they want that body part to be the secondary input type. This pose is detected and recorded as the secondary input type.

[0104] (iii) During the surgery, the doctor wants to control the interface to view the x-ray image of an object. The system can detect that they are controlling a catheter with their hand (which means they cannot provide a gesture). The system can also detect that others near the doctor are having a discussion, which means that voice control of the interface may be difficult. Therefore, since both primary input types are not available, the system starts tracking the doctor's right foot to detect the secondary type of input.

[0105] (iv) The head-mounted device can display a hologram to the user, which contains the x-ray image.

[0106] (v) By gazing at the hologram, the application can understand that it is ready to interact with it. The doctor can now use their right foot to perform gestures to invoke user interface controls. For example:

[0107] (a) By pointing their foot to the right, they are able to select the "Next Frame" button;

[0108] (b) By tapping their foot, the button action of "Next Frame" can be activated, causing the x-ray image control to move to the next frame; and

[0109] (c) To return to the previous frame, the doctor can now point their foot to the left.

[0110] (d) Alternatively, the doctor can perform the same action by keeping their toes on the ground but "pointing and clicking" with their heel.

[0111] (e) Additionally, the doctor can cancel the foot interaction mode (when an action is selected) by lifting their foot completely off the ground. This can be used as a cancel action.

[0112] (f) Many additional foot gestures will be readily apparent to those skilled in the art.

[0113] In terms of implementation, to capture foot gestures, a motion sensor can be attached to the foot to measure the orientation and movement that the foot is making. A camera system for tracking the foot is another option.

[0114] Finally, the system is also able to define restricted areas in the user's physical space where the above steps (a - f) do not work. Whenever the camera sees (or the motion sensor detects) that the user is about to step on, for example, a colleague's foot, a footswitch, a cable, or other objects, the action may not be triggered. Additionally, visual cues can be added to the hologram to indicate what the user might step on.

[0115] Continuing, Figure 4 An example of a computer 1000 in which one or more portions of an embodiment can be employed is shown. The various operations discussed above can utilize the capabilities of the computer 1000. For example, one or more portions of the system for obtaining input from a user to control the interface can be incorporated into any of the elements, modules, applications, and / or components discussed herein. In this regard, it should be understood that the system functional blocks can operate on a single computer or can be distributed across several computers and locations (e.g., via an Internet connection).

[0116] The computer 1000 includes, but is not limited to, a PC, a workstation, a laptop computer, a PDA, a handheld device, a server, a storage device, etc. Generally, in terms of the hardware architecture, the computer 1000 may include one or more processors 1010, a memory 1020, and one or more I / O devices 1030 communicatively coupled via a local interface (not shown). The local interface may be, for example but not limited to, one or more buses or other wired or wireless connections, as known in the art. The local interface may have additional elements, such as controllers, buffers (caches), drivers, repeaters, and receivers, to enable communication. In addition, the local interface may include address, control, and / or data connections to enable proper communication among the above components.

[0117] The processor 1010 is a hardware device for executing software that may be stored in the memory 1020. The processor 1010 may actually be any custom or commercially available processor, a central processing unit (CPU), a digital signal processor (DSP), or an auxiliary processor among several processors associated with the computer 1000, and the processor 1010 may be a semiconductor-based microprocessor (in the form of a microchip) or a microprocessor.

[0118] The memory 1020 may include any one or combination of volatile memory elements (e.g., random access memory (RAM), such as dynamic random access memory (DRAM), static random access memory (SRAM), etc.) and non-volatile memory elements (e.g., ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic tape, compact disc read-only memory (CD-ROM), magnetic disk, floppy disk, cassette tape, cartridge tape, etc.). In addition, the memory 1020 may contain electrical, magnetic, optical, and / or other types of storage media. Note that the memory 1020 may have a distributed architecture, where various components are located far from each other but can be accessed by the processor 1010.

[0119] The software in the memory 1020 may include one or more individual programs, each program including an ordered list of executable instructions for implementing a logical function. According to an exemplary embodiment, the software in the memory 1020 includes a suitable operating system (O / S) 1050, a compiler 1060, source code 1070, and one or more application programs 1080. As shown, the application program 1080 includes many functional components for implementing the features and operations of the exemplary embodiment. According to an exemplary embodiment, the application program 1080 of the computer 1000 may represent various application programs, computing units, logics, functional units, processes, operations, virtual entities, and / or modules, but the application program 1080 is not meant to be a limitation.

[0120] The operating system 1050 controls the execution of other computer programs and provides scheduling, input-output control, file and data management, memory management, and communication control and related services. The inventors contemplate that the application program 1080 for implementing the exemplary embodiments may be applicable to all commercially available operating systems.

[0121] The application program 1080 may be a source program, an executable program (object code), a script, or any other entity including a set of instructions to be executed. When it is a source program, the program is typically translated via a compiler (such as compiler 1060), an assembler, an interpreter, etc. (which may or may not be included within the memory 1020) so as to operate properly in conjunction with the O / S 1050. In addition, the application program 1080 may be written in an object-oriented programming language having data and method classes, or a procedural programming language having routines, subroutines, and / or functions, such as but not limited to C, C++, C#, Pascal, BASIC, API calls, HTML, XHTML, XML, ASP scripts, JavaScript, FORTRAN, COBOL, Perl, Java, ADA,...NET, etc.

[0122] The I / O device 1030 may include input devices such as but not limited to a mouse, a keyboard, a scanner, a microphone, a camera, etc. In addition, the I / O device 1030 may also include output devices such as but not limited to a printer, a display, etc. Finally, the I / O device 1030 may also include devices that transfer both input and output, such as but not limited to a NIC or a modem / demodulator (for accessing remote devices, other files, devices, systems, or networks), a radio frequency (RF) or other transceiver, a telephone interface, a bridge, a router, etc. The I / O device 1030 also includes components for communicating via various networks such as the Internet or an intranet.

[0123] If the computer 1000 is a PC, a workstation, a smart device, etc., the software in the memory 1020 may also include a basic input output system (BIOS) (omitted for simplicity). The BIOS is a set of basic software routines that initialize and test the hardware at startup, start the O / S 1050, and support data transfer between hardware devices. The BIOS is stored in a certain type of read-only memory such as ROM, PROM, EPROM, EEPROM, etc., such that the BIOS can be executed when the computer 800 is activated.

[0124] When the computer 1000 is operating, the processor 1010 is configured to execute software stored in the memory 1020, communicate data with the memory 1020, and generally control the operation of the computer 1000 in accordance with the software. The application program 1080 and the O / S 1050 are read in whole or in part by the processor 1010, possibly buffered within the processor 1010, and then executed.

[0125] When the application program 1080 is implemented in software, it should be noted that the application program 1080 can actually be stored on any computer-readable medium for use by or in conjunction with any computer-related system or method. In the context of this document, a computer-readable medium can be an electrical, magnetic, optical, or other physical device or apparatus that can contain or store a computer program for use by or in conjunction with a computer-related system or method.

[0126] The application program 1080 can be embodied in any computer-readable medium for use by or in conjunction with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can obtain instructions from and execute the instructions of the instruction execution system, apparatus, or device. In the context of this document, a "computer-readable medium" can be any device that can store, transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable medium can be, for example but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, device, or propagation medium.

[0127] Regarding Figure 1 and Figure 2 the methods described and regarding Figure 3 the systems described can be implemented in hardware or software or a combination of both (e.g., as firmware running on a hardware device). To the extent that an embodiment is implemented in whole or in part in software, the functional steps shown in the process flow diagrams can be performed by a suitably programmed physical computing device, such as one or more central processing units (CPUs) or graphics processing units (GPUs). Each process and its individual component steps shown in the flowcharts can be performed by the same or different computing devices. According to an embodiment, a computer-readable storage medium stores a computer program including computer program code that is configured to cause one or more physical computing devices to perform the encoding or decoding methods described above when the program is run on the one or more physical computing devices.

[0128] The storage medium may include volatile and non-volatile computer memories, such as RAM, PROM, EPROM, and EEPROM, optical discs (such as CD, DVD, BD), and magnetic storage media (such as hard disks and magnetic tapes). The various storage media may be fixed within the computing device or may be removable, such that one or more programs stored thereon may be loaded into the processor.

[0129] Insofar as embodiments are implemented partially or wholly in hardware, the blocks shown in the block diagrams Figure 3 may be separate physical components or logical subdivisions of a single physical component, or may all be implemented in an integrated manner in one physical component. The function of a single block shown in the drawings may be divided among multiple components in an implementation, or the functions of multiple blocks shown in the drawings may be combined in a single component in an implementation. Hardware components suitable for embodiments of the present invention include, but are not limited to, conventional microprocessors, application specific integrated circuits (ASICs), and field programmable gate arrays (FPGAs). One or more blocks may be implemented as a combination of dedicated hardware for performing some functions and one or more programmed microprocessors and associated circuitry for performing other functions.

[0130] By studying the drawings, the disclosure, and the claims, those skilled in the art can understand and implement variations of the disclosed embodiments when practicing the claimed invention. In the claims, the word "comprising" does not exclude other elements or steps, and the quantifier "a" or "one" does not exclude a plurality. A single processor or other unit may implement the functions of several items recited in the claims. The fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be advantageous. If a computer program is discussed above, it may be stored / distributed on a suitable medium, such as an optical storage medium or a solid state medium provided together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems. If the term "adapted to" is used in the claims or the specification, it should be noted that the term "adapted to" is intended to be equivalent to the term "configured to". Any reference signs in the claims should not be construed as limiting the scope.

[0131] The flowcharts and block diagrams in the figures illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions that includes one or more executable instructions for implementing the specified logical function. For example, two consecutive blocks shown in succession may actually be executed substantially simultaneously, or these blocks may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by a dedicated hardware-based system that performs the specified functions or actions or a combination of dedicated hardware and computer instructions.

Claims

1. A method (100) for obtaining input from a user to control an interface, the method comprising: Obtaining (110) background data describing the background of the user, wherein the background data includes information describing the attributes of the user or the environment around the user; Processing (120) the background data to determine whether the user is able to provide a primary input type; In response to determining that the user is able to provide the primary input type, detecting (130) the input of the primary type performed by the user, wherein the primary input type includes at least one of a gesture, an eye gaze direction, a voice command, and a controller interaction; and In response to determining that the user is not able to provide the primary input type, detecting (140) a different secondary input type performed by the user, wherein the secondary input type includes at least one of a foot posture, a leg posture, an arm posture, and a head posture.

2. The method according to claim 1, wherein The background data includes information describing at least one of the following: the location of the user, the condition of the user, the task the user is performing, the gaze direction of the user, the ambient noise around the user, the location of the user, and other people near the user.

3. The method according to claim 1 or 2, wherein Obtaining (110) the background data is performed in response to the user providing at least one of the primary input type and the secondary input type.

4. The method according to any one of claims 1 to 3, wherein The secondary input type is selectable by the user.

5. The method according to claim 4, further comprising: Receiving (210) an indication from the user of a body part; and And Determining (220) the secondary input type based on the indication from the user.

6. The method according to any one of claims 1-5, wherein Obtaining the background data includes: Obtaining (112) one or more images including at least a portion of the user from one or more cameras of a user monitoring system; and Analyzing (114) the one or more images to extract the background data.

7. The method according to any one of claims 1-6, further comprising in response to determining that the user is not able to provide the primary input type: Transmitting a notification to the user prompting a switch in the input type; and Receiving (142) confirmation from the user.

8. The method according to any one of claims 1-7, wherein, Detecting the input of the primary type or the input of the secondary type includes: Obtaining (132, 144) a signal containing the input of the primary type or the secondary type, the signal including at least one of an image signal, a sound signal, a pressure signal, and a motion sensor signal; and Analyzing (134, 146) the signal to extract the input of the primary type or the input of the secondary type.

9. The method according to any one of claims 1-8 further comprises determining whether the user is capable of providing multiple primary input types by processing the background data, and wherein, Detecting the input of the secondary type is performed in response to determining that the user is not able to provide each of the multiple primary input types.

10. The method according to any one of claims 1-9, wherein, The interface is based on augmented reality or virtual reality.

11. A computer program comprising computer program code modules, which when the computer program is run on a computer, are adapted to implement the method according to any one of claims 1-10.

12. A system (300) for obtaining input from a user to control an interface (330), the system comprising: A user tracking device (310) configured to monitor the user to generate background data describing the background of the user, wherein the background data includes information describing an attribute of the user or an environment around the user; A processor (320) configured to process the background data to determine whether the user is capable of providing a primary input type, and wherein the user tracking device is further configured to: In response to determining that the user is capable of providing the primary input type, detect the primary type of input performed by the user, wherein the primary input type includes at least one of a gesture, an eye gaze direction, a voice command, and a controller interaction; and In response to determining that the user is not capable of providing the primary input type, detect the secondary type of input performed by the user, wherein the secondary input type is different from the primary input type, and wherein the secondary input type includes at least one of a foot posture, a leg posture, an arm posture, and a head posture.