Low latency pose estimation system with user interaction-based prioritization

The interaction system addresses latency issues in low power systems by dynamically adjusting frame rates based on user activity, ensuring efficient and responsive user interaction in applications like digital whiteboards.

WO2025212019A1PCT designated stage Publication Date: 2025-10-09FLATFROG LAB
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
PCT/SE2025/050295
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-03
Filing Date
2025-04-02
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Low power computing systems struggle with high latency in processing live camera data for real-time tracking of user movements and gestures due to computational limitations, impacting user experience in interactive applications like digital whiteboards.

Method used

An interaction system that adjusts frame rates based on user activity, tracking active users at a higher frame rate and inactive users at a lower rate, reducing computational load and latency by using a controller to determine a skeleton model and generate a mapping region for user input.

Benefits of technology

Provides low latency and efficient computational use by prioritizing active users with higher frame rates and reducing load on inactive users, enhancing user interaction experience in systems like digital whiteboards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SE2025050295_09102025_PF_FP_ABST
    Figure SE2025050295_09102025_PF_FP_ABST
Patent Text Reader

Abstract

An interaction system comprises a controller to receive spatial data of the position of a user in front of an interaction display, determine a skeleton model of the user based on the spatial data, and track the skeleton model of the user when a position and / or movement of a feature of the skeleton model is determined to match a predetermined trigger condition to initiate user input in a display coordinate system. The controller is configured to generate a mapping region when a match of the trigger condition has been determined, the mapping region defines an interaction space in a region in front of the user, map the mapping region to the display coordinate system, and track the skeleton model to determine subsequent input coordinates in the mapping region and the associated display coordinates in the display coordinate system of the interaction display as the user provides the user input.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Low Latency Pose Estimation System with User Interaction-Based Prioritization

[0002] Field

[0003] The technology relates to the field of human-computer interaction, specifically to systems and methods for tracking and interpreting user movements and gestures in front of an interaction display. This technology is particularly relevant for interactive presentations, where users interact with digital content through their body movements and gestures.

[0004] Background

[0005] Interactive applications, such as those used when providing input to digital whiteboards often require real-time tracking of user movements and gestures to provide an engaging and intuitive user experience. One common approach to achieve this is by calculating and tracking skeleton models of users based on live camera data. These skeleton models represent the position and movement of various body parts, such as hands, head, and limbs, which are relevant for user interaction with the system.

[0006] However, calculating skeleton models from live camera data involves complex algorithms and computations that require significant processing power. This can be particularly challenging for low power computing systems, which are designed to consume less energy and often have limited processing capabilities. As a result, these systems may struggle to efficiently process live camera data and calculate skeleton models, leading to long latency in interactive applications. This delay between user input and system response can negatively impact user experience and engagement.

[0007] Prior art systems are typically concerned with gaming applications requiring a large amount of computational resources.

[0008] Therefore, there is a need for an improved method and system for calculating and tracking skeleton models in interactive applications, particularly for low power computing systems, that can reduce latency and computational load, in particular when providing input to digital whiteboards or similar interaction systems. Summary

[0009] According to a first aspect of the disclosure, an interaction system is provided that comprises a controller. The controller is configured to receive spatial data of the position of a user in front of an interaction display, determine a skeleton model of the user based on the spatial data, and track the skeleton model of the user when a position and / or movement of a feature of the skeleton model is determined to match a predetermined trigger condition to initiate user input in a display coordinate system of the interaction display. When a match of the trigger condition has been determined, the controller generates a mapping region, which defines an interaction space in a region in front of the user. The controller maps the mapping region to the display coordinate system, whereby an input coordinate in the mapping region has an associated display coordinate in the display coordinate system. The controller then tracks the skeleton model of the user to determine subsequent input coordinates in the mapping region and the associated display coordinates in the display coordinate system of the interaction display as the user provides the user input. The controller is configured to track the skeleton model with a first frame rate in a first mode and with a second frame rate in a second mode, the second frame rate being higher than the first frame rate. The controller is configured to track the skeleton model in the second mode when a match of the trigger condition has been determined. This aspect allows for reducing the computational load since the skeleton model is tracked in the second mode with increased frame rate, when the user triggers the trigger condition and not all the time. This feature allows for reduced latency of the interaction when the user is actively providing input, in which case the need for low latency is the most important.

[0010] Optionally in some examples, the controller is configured to track the skeleton model in the first mode until a match of the trigger condition has been determined, upon which the controller is configured to track the skeleton model in the second mode. This feature allows for reducing the computational load while providing transition between the two modes for a smooth user experience.

[0011] Optionally in some examples, the controller is configured to track the skeleton model in the second mode until a detected input feature providing the input coordinate is removed from the mapping region, whereupon the controller is configured to track the skeleton model in the first mode. This feature allows for efficient use of computational resources, as the higher frame rate tracking is only used when necessary.

[0012] Optionally in some examples, the controller is configured to determine skeleton models for a plurality of users positioned in front of the interaction display, and track a subset of skeleton models in the second mode, of an identified active subset of the users, when a corresponding match of the trigger condition has been determined for the respective users in the active subset of users. This feature allows for multiple users to interact with the system simultaneously, with less demand for computational resources, while still maintaining a high level of responsiveness for the active users.

[0013] Optionally in some examples, the frame rate in the first mode is based on the number of users in an identified inactive subset of the users where a match of the trigger condition has not been determined. This feature allows for efficient use of computational resources, as the frame rate is adjusted based on the number of inactive users.

[0014] Optionally in some examples, the controller is configured to determine a respective skeleton model for one user at a time, according to a cycle, amongst the inactive subset of the users, with each advancing frame in the first mode. This feature allows for efficient use of computational resources, while the positions of the inactive users are still determined over a full cycle.

[0015] Optionally in some examples, the controller is configured to track the skeleton model in each frame in the second mode. This feature allows for a high level of responsiveness when the user is actively providing input.

[0016] Optionally in some examples, the controller is configured to track a subset of skeleton models in a third mode, of an identified pre-active subset of the users. The controller is configured to track the skeleton model with a third frame rate in the third mode, wherein the third frame rate is higher than the first frame rate and lower than the second frame rate. This feature allows for a more responsive interaction for users who are about to become active, providing a smoother user experience. Optionally in some examples, the controller is configured to continuously receive the spatial data to track the skeleton model and the associated motion of the user in the mapping region over a duration of time, when a match of the trigger condition has been determined. The controller is configured to determine associated input coordinates for the motion and generate a control signal to display an output of the motion in the display coordinate system. This feature allows for continuous tracking of the user's movements, providing a more natural and intuitive interaction with the system.

[0017] Optionally in some examples, the controller is configured to generate a control signal to stop a display of the output in the display coordinate system as a detected input feature providing the input coordinate is removed from the mapping region. This feature allows for a clear end of the user's input, providing a more intuitive interaction with the system.

[0018] Optionally in some examples, the controller is configured to detect a combination of a plurality of the features of the skeleton model and / or movements thereof so that a match of the trigger condition is determined once the combination of features matches a predetermined user command to trigger user input in the display coordinate system. This feature allows for reducing the occurrence of false triggers for user input.

[0019] Optionally in some examples, the user input is triggered based on detecting a combination of features from the spatial data comprising at least a first feature of the skeleton model associated with the user’s body and a second feature associated with an input object such as a stylus. This feature allows for reducing the occurance of false triggers for user input while allowing a wider range of user inputs.

[0020] Optionally in some examples, the user input is triggered based on detecting a combination of features in a defined sequence, and / or the user input is triggered based on detecting a combination of features in a defined spatial relationship. This feature allows for further reducing the occurance of false triggers for user input

[0021] According to a second aspect of the disclosure, a method in an interaction system is provided. The method comprises receiving spatial data of the position of a user in front of an interaction display, determining a skeleton model of the user based on the spatial data, and tracking the skeleton model when a position and / or movement of a feature of the skeleton model is determined to match a predetermined trigger condition to initiate user input in a display coordinate system of the interaction display. When a match of the trigger condition has been determined, the method comprises generating a mapping region, wherein the mapping region defines an interaction space in a region in front of the user, mapping the mapping region to the display coordinate system, whereby an input coordinate in the mapping region has an associated display coordinate in the display coordinate system, and tracking the skeleton model of the user to determine subsequent input coordinates in the mapping region and the associated display coordinates in the display coordinate system of the interaction display as the user provides the user input. The method comprises tracking the skeleton model with a first frame rate in a first mode and with a second frame rate in a second mode, the second frame rate being higher than the first frame rate. The method comprises tracking the skeleton model in the second mode when a match of said trigger condition has been determined. This method allows for reducing the computational load since the skeleton model is tracked when the user triggers the trigger condition and not all the time.

[0022] Brief Description of the Drawings

[0023] Examples are described in more detail below with reference to the appended drawings. Figure 1 is an interaction system 100 according to an example of the disclosure.

[0024] Figure 2 shows an example of tracking skeleton models in the interaction system 100 with an image series of three frames 0 - 2, illustrating the identification of active and inactive subsets of users and their corresponding skeleton models. The active subset of users is tracked with each frame, while the skeleton models of the inactive subset of users are cycled through over three frames, and thereby calculated with a lower frame rate.

[0025] Figure 3 shows a series of three frames where each frame shows which skeleton models are calculated in that image. Here each skeleton model is calculated in each frame, resulting in a very high computational load and poor framerate and latency for an interacting user. Figure 4 illustrates a typical skeleton model of a user with each key point marked with a unique ID.

[0026] Figure 5 depicts a method in an interaction system according to an example of the disclosure.

[0027] Detailed Description

[0028] The detailed description set forth below provides information and examples of the disclosed technology with sufficient detail to enable those skilled in the art to practice the disclosure.

[0029] Figure 1 shows an interaction system 100 according to an example of the disclosure. The interaction system 100 comprises a controller 101 configured to receive spatial data of the position of a user in front of an interaction display 201 . The controller 101 determines a skeleton model 103 of the user based on the spatial data and tracks the skeleton model 103 when a position and / or movement of a feature 106 of the skeleton model 103 (as exemplified in figure 2 and 4) matches a predetermined trigger condition to initiate user input in a display coordinate system 203 of the interaction display 201 . The controller 101 generates a mapping region 104 that defines an interaction space 105 in a region (A) in front of the user when a match of the trigger condition has been determined. The mapping region 104 is mapped to the display coordinate system 203, whereby an input coordinate (xi, yi) in the mapping region 104 has an associated display coordinate (xd, yd) in the display coordinate system 203. The controller 100 is configured to track the skeleton model 103 of the user to determine subsequent input coordinates (xi,yi) in the mapping region and the associated display coordinates (xd,yd) in the display coordinate system 203 of the interaction display 201 as the user provides the user input.

[0030] The controller 101 is configured to track the skeleton model 103 with a first frame rate in a first mode, and to track the skeleton model 103 with a second frame rate in a second mode. The second frame rate is higher than the first frame rate. The controller 101 is configured to track the skeleton model 103 in the second mode when a match of the trigger condition has been determined, as described further below. Figure 2 shows an example of tracking skeleton models in the interaction system 100 with an image series of three frames 0 - 2, illustrating the identification of active subsets 205 and inactive subsets 206 of users and their corresponding skeleton models. The user marked by a short-dashed bounding box is identified as interacting with the system 100 since in this example a feature 106 of the skeleton model 103 match a trigger condition. The skeleton model 103 of the triggering user is tracked in the second mode with the associated second frame rate, whereas the remaining user are tracked in the first mode with the first frame rate, being lower than the second frame rate. The triggering user may thus be identified as an active subset 205 of the users and the associated skeleton model 103 is calculated and tracked each frame (0 - 2). The remaining users with long-dashed bounding boxes may be identified as an inactive subset 206 and their skeleton models are cycled through over three frames, and thereby calculated with a lower frame rate.

[0031] Figure 3 shows a series of three frames where each frame shows which skeleton models are calculated in that image. In this example, each skeleton model is calculated in each frame, resulting in a very high computational load and poor frame rate and latency for an interacting user.

[0032] Figure 4 illustrates a typical skeleton model of a user with each key point marked with a unique ID. The skeleton model represents the position of the user's body parts, such as hands, head, or other body parts that are relevant for user interaction with the system.

[0033] Figure 5 depicts a method 300 in an interaction system 100 according to an example of the disclosure. The method 300 comprises receiving 301 spatial data of the position of a user in front of an interaction display 201 , determining 302 a skeleton model 103 of the user based on the spatial data, tracking 303 the skeleton model when a position and / or movement of a feature 106 of the skeleton model matches a predetermined trigger condition to initiate user input in a display coordinate system 203 of the interaction display, generating 304 a mapping region 104, mapping the mapping region to the display coordinate system, and tracking 306 the skeleton model of the user to determine subsequent input coordinates (xi, yi) in the mapping region and the associated display coordinates (xd, yd) in the display coordinate system of the interaction display as the user provides the user input. The method 300 comprises tracking 307 the skeleton model with a first frame rate in a first mode and with a second frame rate in a second mode, the second frame rate being higher than the first frame rate. The method 300 further comprises tracking 308 the skeleton model in the second mode when a match of said trigger condition has been determined. The method 300 thus provides for the advantageous as described for the interaction system 100. The method 300 provides a low latency interaction experience for active users, while reducing the computational load by tracking inactive users at a lower frame rate.

[0034] 1 . Interaction System Details

[0035] The interaction system 100 provides facilitated user interaction with a display. The system 100 tracks the user's movements and translate them into inputs for the display 201. The system 100 is responsive and efficient, reducing the computational load by tracking the movements of users who are actively interacting with the display 201 with greater precision. This is achieved by using a controller 101 that receives spatial data of the user's position and creates a skeleton model 103 of the user. The controller 101 then tracks the skeleton model when a position and / or movement of a feature 106 of the skeleton model matches a predetermined trigger condition to initiate user input in a display coordinate system 203 of the interaction display 201 .

[0036] 1.1. Receiving Spatial Data and Determining the Skeleton Model

[0037] In some implementations, the controller 101 is a component of the interaction system 100. The controller 101 is configured to receive spatial data of the position of a user in front of an interaction display 201. This spatial data can be obtained from various sources, such as cameras or sensors 107, and can include information about the user's position, orientation, and movements. Based on this spatial data, the controller 101 determines the skeleton model 103 of the user. The skeleton model 103 is a representation of the user's body, including key points such as the head, hands, and other body parts that are relevant for user interaction with the system.

[0038] 1.2. Tracking the Skeleton Model Based on Trigger Condition The controller 101 tracks the skeleton model when a position and / or movement of a feature 106 of the skeleton model matches a predetermined trigger condition to initiate user input in a display coordinate system 203 of the interaction display 201 .

[0039] 1.3. Generating and Mapping the Interaction Space

[0040] The controller 101 is also configured to generate a mapping region 104, which defines an interaction space 105 in a region (A) in front of the user, as schematically illustrated in figure 1. The mapping region is then mapped to the display coordinate system 203, whereby an input coordinate (xi, yi) in the mapping region 104 has an associated display coordinate (xd, yd) in the display coordinate system 203. The controller 101 tracks the skeleton model 103 of the user and monitors the user's movements in the interaction space 105. As the user moves, the controller 101 determines the input coordinates (xi, yi) in the mapping region 104 and the associated display coordinates (xd, yd) in the display coordinate system 203. This allows the user's movements in the interaction space 105 to be translated into inputs for the interaction display 201. The mapping region 104 is designed to be flexible and adaptable, allowing the interaction space 105 to be adjusted based on the user's position and movements.

[0041] In some examples, the interaction display 201 is a component of the interaction system

[0042] 100. The interaction display 201 can be any type of display that is capable of receiving and processing user inputs. In some implementations, the interaction display 201 can be a touch-sensitive display that can detect the presence and location of a touch within the display area. In other implementations, the interaction display 201 can be a non- touch-sensitive display that is used in conjunction with other input devices, such as a stylus or a mouse. The interaction display 201 is in communication with the controller

[0043] 101.

[0044] 1.4. Triggering Feature of Skeleton Model

[0045] The triggering feature 106 is a part or combination of parts of the skeleton model 103 that is used to determine when a position and / or movement of the feature 106 matches a predetermined trigger condition to initiate user input in a display coordinate system 203 of the interaction display 201. The triggering feature 106 can involve any part of the skeleton model 103, such as the user's hand, arm, head, or other body parts that are relevant for user interaction with the system 100. The present disclosure refers to such part or parts of the skeleton model 103 as feature 106 for brevity. For example, the feature 106 can trigger the trigger condition when a combination of parts of the skeleton model 103 are placed in a particular relative position to each other. E.g. a user who desires to provide input may raise the arm and may trigger the trigger condition when the forearm and the upper arm, of the user’s left or right arm, are identified as being arranged in a defined relative position to each other, such as withing a defined angular interval to each other. In other examples, the trigger condition may be satisfied when the forearm and / or upper arm is placed in a defined position relative to the user’s head, such as within a defined distance from the user’s head. In other examples, the trigger condition may be met when the triggering feature 106 follows a defined path of movement. E.g. the user may raise the arm, and subsequently move the arm relative to the body, such as extending the arm away from the body, to trigger the trigger condition.

[0046] The interaction system 100 may thus be used with different trigger conditions. As discussed, the trigger condition can be a specific position, such as a particular relative position of different parts of the skeleton model 103 or movements thereof, or a combination of positions or movements of multiple parts of the skeleton model 103. The controller 101 tracks the skeleton model and monitors the position and / or movement of the feature 106. When the feature 106 matches the predetermined trigger condition, the controller 101 initiates user input in the display coordinate system 203 and the skeleton model 103 of the active user is tracked with higher accuracy than the skeleton models of inactive users in order to reduce the computational load, as described further below.

[0047] 1 .5. Frame Rate Adjustment Based on User Activity

[0048] The controller 101 is configured to track the skeleton model 103 with a first frame rate in a first mode and with a second frame rate in a second mode, the second frame rate being higher than the first frame rate. The controller 101 tracks the skeleton model 103 in the second mode when a match of the trigger condition has been determined, such as illustrated in Fig. 2 where a feature 106 of a user’s skeleton model 103 match the trigger condition. This allows the system 100 to provide a low latency interaction experience for active users, while reducing the computational load by tracking inactive users at a lower frame rate.

[0049] The controller 101 may be configured to track the skeleton model 103 in the first mode until a match of said trigger condition has been determined upon which the controller 101 is configured to track the skeleton model 103 in the second mode. This saves computational resources when the users are not active in providing input. The controller 101 may be configured to track the skeleton model in the second mode until a detected input feature 202, providing the input coordinate (xi,yi), is removed from the mapping region 104, whereupon the controller 101 is configured to track the skeleton model 103 in the first mode. The input feature 202 may be the user’s hand or finger, or a stylus. Thus, the controller 101 may return to the tracking with lower resolution in the first mode when the user jumps out of input mode.

[0050] The controller 101 may be configured to determine skeleton models 103 for a plurality of users positioned in front of the interaction display 201 , and track a subset of skeleton models in the second mode, of an identified active subset 205 of the users, when a corresponding match of the trigger condition has been determined for the respective users in the active subset 205 of users. Fig. 2 shows an example where one user among a group of four users has been identified as an active subset 205 of the users, since a position and / or movement of a feature 106 of the aforementioned user’s skeleton model 103 match the trigger condition. The active subset 205 is tracked in the second mode, and the skeleton model 103 of the user in the active subset 205 is calculated for each frame in the example in Fig. 2. The controller 101 may thus be configured to track the skeleton model 103 in each frame in the second mode.

[0051] Skeleton models of the inactive subset 206 of users in the example of Fig. 2 are cycled through over three frames, and thereby calculated with a lower frame rate. The frame rate in the first mode may thus be based on the number of users in the identified inactive subset 206 of the users where a match of the trigger condition has not been determined. The controller 101 may be configured to determine a respective skeleton model for one user at a time, according to a cycle, amongst the inactive subset 206 of the users, with each advancing frame in the first mode, as illustrated in the example in Fig. 2.

[0052] The controller 101 may be configured to track a subset of skeleton models in a third mode, of an identified pre-active subset of the users. The controller 101 may be configured to track the skeleton model with a third frame rate in the third mode. The third frame rate is higher than the first frame rate and lower than the second frame rate. This feature allows for a more responsive interaction for users who are about to become active, providing a smoother user experience.

[0053] The controller 101 may be configured to continuously receive said spatial data to track the skeleton model 103 and the associated motion of the user in the mapping region over a duration of time, when a match of said trigger condition has been determined. The controller 101 may thus be configured to determine associated input coordinates (xi,yi) for said motion and generate a control signal to display an output 204 of the motion in the display coordinate system (xd,yd), as exemplified in Fig. 1 . The controller 101 may be configured to generate a control signal to stop a display of the output in the display coordinate system as a detected input feature 202 providing the input coordinate (xi,yi) is removed from the mapping region 104.

[0054] The controller 101 may be configured to detect a combination of a plurality of said features of the skeleton model 103 and / or movements thereof so that a match of said trigger condition is determined once said combination of features matches a predetermined user command to trigger user input in the display coordinate system 203. This feature allows for reducing the occurrence of false triggers for user input. The user input may be triggered based on detecting a combination of features from the spatial data comprising at least a first feature 106 of the skeleton model associated with the user’s body and a second feature associated with an input object such as a stylus. The user input may be triggered based on detecting a combination of features in a defined sequence, and / or wherein the user input is triggered based on detecting a combination of features in a defined spatial relationship.

[0055] In some configurations, the interaction system 100 can be applied in video conferencing systems, home entertainment systems such as TV’s, or in digital whiteboards. Digital whiteboards are interactive displays with a touch sensor that allow users to write or draw on the display using their fingers or a stylus. The interaction system 100 can enhance the user experience of digital whiteboards by providing a low latency interaction experience and reducing the computational load. The controller 101 receives spatial data of the position of a user in front of the digital whiteboard and determines a skeleton model 103 of the user. The controller 101 then tracks the skeleton model when a position and / or movement of a feature 106 of the skeleton model matches a predetermined trigger condition to initiate user input in a display coordinate system 203 of the digital whiteboard. The controller 101 generates a mapping region 104, which defines an interaction space 105 in a region (A) in front of the user. The mapping region is then mapped to the display coordinate system 203, whereby an input coordinate (xi, yi) in the mapping region has an associated display coordinate (xd, yd) in the display coordinate system 203. This allows the user's movements in the interaction space 105 to be translated into inputs for the digital whiteboard, providing a seamless and intuitive user experience. The controller 101 is also configured to track the skeleton model with a first frame rate in a first mode and with a second frame rate in a second mode, the second frame rate being higher than the first frame rate. The controller 101 tracks the skeleton model in the second mode when a match of the trigger condition has been determined. This allows the system 100 to provide a low latency interaction experience for active users, while reducing the computational load by tracking inactive users at a lower frame rate. This application of the interaction system 100 can enhance the user experience of digital whiteboards and touch sensing apparatuses by providing a more responsive and intuitive interaction experience.

[0056] The terminology used herein is for the purpose of describing particular aspects only and is not intended to be limiting of the disclosure. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items. It will be further understood that the terms "comprises," "comprising," "includes," and / or "including" when used herein specify the presence of stated features, integers, actions, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, actions, steps, operations, elements, components, and / or groups thereof.

[0057] It will be understood that, although the terms first, second, etc., may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and, similarly, a second element could be termed a first element without departing from the scope of the present disclosure.

[0058] Relative terms such as "below" or "above" or "upper" or "lower" or "horizontal" or "vertical" may be used herein to describe a relationship of one element to another element as illustrated in the Figures. It will be understood that these terms and those discussed above are intended to encompass different orientations of the device in addition to the orientation depicted in the Figures. It will be understood that when an element is referred to as being "connected" or "coupled" to another element, it can be directly connected or coupled to the other element, or intervening elements may be present. In contrast, when an element is referred to as being "directly connected" or "directly coupled" to another element, there are no intervening elements present.

[0059] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. It will be further understood that terms used herein should be interpreted as having a meaning consistent with their meaning in the context of this specification and the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0060] It is to be understood that the present disclosure is not limited to the aspects described above and illustrated in the drawings; rather, the skilled person will recognize that many changes and modifications may be made within the scope of the present disclosure and appended claims. In the drawings and specification, there have been disclosed aspects for purposes of illustration only and not for purposes of limitation, the scope of the disclosure being set forth in the following claims.

Claims

Claims1. An interaction system (100) comprising a controller (101 ) configured to receive spatial data of the position of a user in front of an interaction display (201 ), determine a skeleton model (103) of the user based on the spatial data, track the skeleton model of the user when a position and / or movement of a feature (106) of the skeleton model is determined to match a predetermined trigger condition to initiate user input in a display coordinate system (203) of the interaction display, wherein the controller is further configured to, when a match of said trigger condition has been determined, generate a mapping region (104), wherein the mapping region defines an interaction space (105) in a region (A) in front of the user, map the mapping region to the display coordinate system, whereby an input coordinate (xi,yi) in the mapping region has an associated display coordinate (xd,yd) in the display coordinate system (203), and track the skeleton model of the user to determine subsequent input coordinates (xi,yi) in the mapping region and the associated display coordinates (xd,yd) in the display coordinate system of the interaction display as the user provides the user input, wherein the controller is configured to track the skeleton model with a first frame rate in a first mode and with a second frame rate in a second mode, the second frame rate being higher than the first frame rate, and wherein the controller is configured to track the skeleton model in the second mode when a match of said trigger condition has been determined.

2. Interaction system according to claim 1 , wherein the controller is configured to track the skeleton model in the first mode until a match of said trigger condition has been determined upon which the controller is configured to track the skeleton model in the second mode.

3. Interaction system according to claim 1 or 2, wherein the controller is configured to track the skeleton model in the second mode until a detected input feature (202) providing the input coordinate (xi,yi) is removed from the mapping region, whereupon the controller is configured to track the skeleton model in the first mode.

4. Interaction system according to any of claims 1 - 3, wherein the controller is configured to determine skeleton models for a plurality of users positioned in front of the interaction display, and track a subset of skeleton models in the second mode, of an identified active subset (205) of the users, when a corresponding match of the trigger condition has been determined for the respective users in the active subset of users.

5. Interaction system according to any of claims 1 - 4, wherein the frame rate in the first mode is based on the number of users in an identified inactive subset (206) of the users where a match of the trigger condition has not been determined.

6. Interaction system according to claim 5, wherein the controller is configured to determine a respective skeleton model for one user at a time, according to a cycle, amongst the inactive subset of the users, with each advancing frame in the first mode.

7. Interaction system according to any of claims 1 - 6, wherein the controller is configured to track the skeleton model in each frame in the second mode.

8. Interaction system according to any of claims 1 - 7, wherein the controller is configured to track a subset of skeleton models in a third mode, of an identified preactive subset of the users, wherein controller is configured to track the skeleton model with a third frame rate in the third mode, wherein the third frame rate is higher than the first frame rate and lower than the second frame rate.

9. Interaction system according to any of claims 1 - 8, wherein the controller is configured to continuously receive said spatial data to track the skeleton model and the associated motion of the user in the mapping region over a duration of time, when a match of said trigger condition has been determined,whereby the controller is configured to determine associated input coordinates (xi,yi) for said motion and generate a control signal to display an output (204) of the motion in the display coordinate system (xd,yd).

10. Interaction system according to claim 9, wherein the controller is configured to generate a control signal to stop a display of the output in the display coordinate system as a detected input feature (202) providing the input coordinate (xi,yi) is removed from the mapping region.11 . Interaction system according to any of claims 1 - 10, wherein the controller is configured to detect a combination of a plurality of said features of the skeleton model and / or movements thereof so that a match of said trigger condition is determined once said combination of features matches a predetermined user command to trigger user input in the display coordinate system.

12. Interaction system according to claim 1 1 , wherein the user input is triggered based on detecting a combination of features from the spatial data comprising at least a first feature (106) of the skeleton model associated with the user’s body and a second feature associated with an input object such as a stylus.

13. Interaction system according to claim 1 1 or 12, wherein the user input is triggered based on detecting a combination of features in a defined sequence, and / or wherein the user input is triggered based on detecting a combination of features in a defined spatial relationship.

14. A method (300) in an interaction system (100) comprising receiving (301 ) spatial data of the position of a user in front of an interaction display (201 ), determining (302) a skeleton model (103) of the user based on the spatial data, tracking (303) the skeleton model (103) when a position and / or movement of a feature (106) of the skeleton model is determined to match a predetermined trigger condition to initiate user input in a display coordinate system (203) of the interaction display,wherein, when a match of said trigger condition has been determined, the method comprises generating (304) a mapping region (104), wherein the mapping region defines an interaction space (105) in a region (A) in front of the user, mapping (305) the mapping region to the display coordinate system, whereby an input coordinate (xi,yi) in the mapping region has an associated display coordinate (xd,yd) in the display coordinate system (203), and tracking (306) the skeleton model of the user to determine subsequent input coordinates (xi,yi) in the mapping region and the associated display coordinates (xd,yd) in the display coordinate system of the interaction display as the user provides the user input, wherein the method further comprises tracking (307) the skeleton model with a first frame rate in a first mode and with a second frame rate in a second mode, the second frame rate being higher than the first frame rate, and tracking (308) the skeleton model in the second mode when a match of said trigger condition has been determined.

Citation Information

Patent Citations

  • Physical interaction zone for gesture-based user interfaces

    US20110193939A1

  • Interacting with user interface via avatar

    US20110304632A1

  • Recognition system for sharing information

    US20170131781A1

  • Information processing apparatus, position information generation method, and information processing system

    US20180011543A1

  • Motion tracking apparatus and system

    US20180239421A1