Input system enabling direct manual input to displayed video

The 3D camera-based input system allows participants to interact with large-screen display devices from their seats, addressing the need for seat-bound operation by recognizing hand gestures, thus enhancing user interaction.

WO2026004411A1PCT designated stage Publication Date: 2026-01-02INTERMAN CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/018324
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-23
Filing Date
2025-05-21
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Participants in large meetings using large-screen display devices must physically move to operate content on a whiteboard or screen, lacking the ability to interact from their seats.

Method used

An input system utilizing a 3D camera with a depth sensor to identify the three-dimensional positions of a user's viewpoint and hand, allowing operation of displayed content through hand gestures without the need for special devices.

Benefits of technology

Enables participants to interact with displayed content from their seats, mimicking touch panel operations using fingers, supporting multi-touch and gesture recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025018324_02012026_PF_FP_ABST
    Figure JP2025018324_02012026_PF_FP_ABST
Patent Text Reader

Abstract

Provided is an input system enabling direct manual input to a displayed video from a spaced-apart position. This input system comprises: a display device; a 3D camera that captures a user facing the display device to look at a display screen; and an information processing device connected to the display device and the 3D camera. The information processing device outputs a video signal to the display device in order to display a video, detects a viewpoint and a position of the hand of a user looking at the display screen from a video of the 3D camera, identifies a position on the display screen that is an extension of a straight line connecting the viewpoint and the position of the hand of the user, as an operation position, and recognizes the operation position and the movement of the hand of the user as an input to the information processing device.
Need to check novelty before this filing date? Find Prior Art

Description

An input system that allows direct manual input into the displayed image

[0001] The present invention relates to an input system that allows direct manual input from a remote position relative to a displayed image.

[0002] When holding meetings at companies or research institutes, it is common to use a whiteboard if the number of participants is small, and a projector if the number of participants is larger. Whiteboards are convenient for situations where ideas can be freely exchanged, as they are easy to write on and correct. On the other hand, projectors are effective when the agenda is decided in advance and materials are prepared. A common projector is one that connects a computer and projects the screen onto a screen.

[0003] Recently, interactive whiteboards that combine a large-screen LCD display with a touch panel have also come into use. In these cases, users can use a touch pen or their fingers to manipulate the displayed content and write directly on the screen, and the content can be easily saved and shared over a network, making them highly convenient (Patent Document 1).

[0004] JP 2013-113916 A

[0005] When multiple participants are holding a meeting while looking at a traditional whiteboard or interactive whiteboard, participants who are standing while giving a presentation can easily write on the whiteboard or screen.

[0006] However, if other participants want to operate the content displayed on the whiteboard or screen or write on the screen, they must get up from their seats, go to the screen, and borrow a stylus.

[0007] SUMMARY OF THE INVENTION It is therefore an object of the present invention to provide an input system that allows each participant to operate the displayed content from their own seat when a large screen display device is used in a meeting with a large number of participants.

[0008] In order to solve the above problem, an input system according to one aspect of the present invention comprises a display device, a position identification means for identifying the three-dimensional positions of the viewpoint and hand of a user facing the display device and looking at the display screen of the display device, and an information processing device connected to the display device and the position identification means, wherein the information processing device outputs a video signal to the display device, causes the display device to display an image, identifies a position on the display screen on an extension of a straight line connecting the viewpoint and the position of the hand of the user as an operation position, and recognizes the operation position and the movement of the user's hand as an input to the information processing device.

[0009] In one embodiment, the position specifying means is a 3D camera equipped with a depth sensor.

[0010] Furthermore, in one embodiment, the position of the user's hand identified by the 3D camera is a fingertip.

[0011] Furthermore, in one embodiment, the display device is a liquid crystal display or a projector.

[0012] The input system according to the present invention allows participants to operate the displayed content without leaving their seats when using a large-screen display device in a meeting with many participants. At this time, participants do not need to use special devices, and can operate the content just by pointing with their fingers, just as if they were using a touch panel in front of them.

[0013] Fig. 1 is a diagram showing a meeting using an input system according to an embodiment of the present invention. Fig. 2 is a diagram showing an image of a user's field of view using the input system according to an embodiment of the present invention. Fig. 3 is a diagram for explaining calculation of an operation position on the screen of a digital whiteboard 1 in the input system according to an embodiment of the present invention. Fig. 4 is a diagram for explaining how a user performs an operation using gestures in the input system according to an embodiment of the present invention. Fig. 5 is a diagram for explaining how a user performs an operation using gestures in the input system according to an embodiment of the present invention.

[0014] An embodiment of the input system according to the present invention will now be described with reference to the accompanying drawings. Fig. 1 shows a meeting using the input system according to the present invention. This input system comprises a digital whiteboard (electronic whiteboard) 1 using a liquid crystal display, a personal computer (information processing device) 3 that supplies a video signal to the digital whiteboard 1, and a 3D camera 5 equipped with a depth sensor that is connected to the personal computer 3.

[0015] What is required of the digital whiteboard 1 here is the functionality of a large-screen display device that displays images from the personal computer 3. In other words, it is sufficient if it is a large-screen LCD monitor that displays the image signal output from the personal computer 3. Therefore, there is no problem in using a projector that can input the image signal from the personal computer 3 instead. The connection to the personal computer 3 can be wireless, such as Wi-Fi, or wired, such as HDMI (registered trademark).

[0016] The 3D camera 5 is installed in a position that allows it to capture images of the faces and hands of meeting participants looking at the screen of the digital whiteboard 1. In this embodiment, the 3D camera 5 is installed above the digital whiteboard 1. An example of such a 3D camera is Kinect (registered trademark).

[0017] In particular, the 3D camera 5 is used to acquire three-dimensional coordinates of each part (face, eyes, hands (fingertips), etc.) of a user (usually multiple users) looking at the digital whiteboard 1. Normally, values ​​can be acquired in a coordinate system with the position of the 3D camera as the origin, but here, an affine transformation is used to convert the coordinates to be expressed as (x, y, z) with the center of the screen of the digital whiteboard 1 as the origin, the vertical direction as the z axis, the horizontal direction as the y axis, and the depth direction as the x axis. The unit is cm.

[0018] Although the user is seated at a distance from the digital whiteboard 1, the user can use his / her hands, especially his / her fingers, to interactively operate the image displayed on the digital whiteboard 1 as if it were a touch panel.

[0019] For example, to click on an icon displayed on the screen of the digital whiteboard 1, one looks at the screen and places one's finger in the air over the icon in one's field of vision. Figure 2 shows an image of the user's field of vision, which corresponds to the image on the user's retina. One clicks on the pen icon P in this field of vision. That is, one places one's finger over the icon in the image of one's field of vision, moves the finger slightly (about a few centimeters) toward the screen, and then moves it back.

[0020] This switches the finger-based input system into drawing mode. Therefore, by moving your finger slightly forward toward the screen again, you can write directly on the screen. For example, by moving your finger forward and then drawing a circle in the air, you can create a marking M as shown in the figure.

[0021] In other words, moving your finger slightly (a few centimeters) toward the screen corresponds to touching the finger to the touch panel, and moving it back corresponds to lifting your finger from the touch panel. If your finger is within the field of view of the screen, the cursor moves in accordance with the movement of your finger's position within the field of view. Also, if your finger is outside the screen, for example, when your hand is down, the cursor does not move and no input is made.

[0022] To enable the above operations, perspective projection transformation is performed on the fingertip position on the screen as seen from the user's viewpoint. In this case, the position on the screen after perspective projection transformation differs depending on whether the viewpoint is placed at the right eye or the left eye. In reality, just like a "dominant hand" or a "dominant foot," there is a "dominant eye." While many people are right-eye dominant, there are also a certain number of people who are left-eye dominant. Unless it is decided in advance whether the view is aligned with the right eye or the left eye, that is, unless it is decided whether the view is aligned with the right eye or the left eye, perspective projection transformation cannot be performed correctly.

[0023] Therefore, the effective point is registered for each user in advance. Then, when obtaining the coordinates of the user's eyes or fingertips, the user is identified by face authentication, and perspective projection transformation is performed with the effective point position as the viewpoint.

[0024] In Figure 3, the coordinates of the user's viewpoint E (position of the dominant eye) are (e x , e y , e z ), and the coordinate of the fingertip (e.g., the tip of the index finger) is (c x , c y , c z ) The point where the line connecting these two points intersects with the screen of the digital whiteboard 1 (yz plane) is called (p x , p y , p z ), then p y , p z is found (p x is always 0).

[0025] p y = (c y *e x -c x *e y ) / (e x -c x ) p z = (c z *e x -c x *e z ) / (e x -c x )

[0026] The coordinates obtained in this way (p x , p y , p z ) into coordinates on the screen of the personal computer that is outputting the original video. If the screen coordinate system of the display device has the upper left corner as its origin, and the display size is 3840 x 2160 pixels, then the screen coordinates (d x , d y ) is calculated as follows:

[0027] d x = ((W / 2)+p y ) * (3840 / W) d y = ((H / 2)-p z ) * (2160 / H)

[0028] where W is the horizontal width (cm) of the digital whiteboard screen, and H is the vertical height (cm) of the digital whiteboard screen. In the display device's screen coordinate system, the left-right direction is the x-axis, the vertical direction is the y-axis, and x increases from the origin in the upper left corner to the right and y increases downward. If another coordinate system is used, simply modify the above formula as appropriate.

[0029] If the device coordinate position of the conversion result fits within the screen, that is, 0 < d x <3840 and 0 < d y If the value is <2160, the user's hand functions as an input device for the personal computer. That is, the cursor moves to the device coordinate position on the screen of the personal computer. Of course, the display of the personal computer itself does not need to be a touch panel. Operations that can be performed on the digital whiteboard screen include tapping, double tapping, flicking, swiping, scrolling, dragging, and long pressing.

[0030] This allows the digital whiteboard to be operated like a touch panel, but it supports multi-touch. That is, the position of each finger is acquired as a feature point from the image captured by the 3D camera. The device coordinate position is calculated for each finger position, enabling conventional operations on a touch panel using multiple fingers, including pinch-in / pinch-out.

[0031] Alternatively, both hands may be used. For example, moving the fingertips of both hands slightly forward at the same time and then separating (opening) corresponds to pinching in. Similarly, moving the fingertips of both hands slightly forward at the same time and then bringing them closer together corresponds to pinching out.

[0032] Furthermore, gestures linked to the screen display can also be supported. For example, an object displayed on the screen can be pinched and moved with the fingers. In the example of FIG. 4, a gesture of pinching a file icon I with the fingers is shown. In this state, by moving the icon I so that it overlaps with the operation icon PT or DP (shown as "print" in this case) as shown in FIG. 5, the corresponding operation (print) can be performed. FIGS. 4 and 5 are direct illustrations of the image in the user's field of vision.

[0033] In the above explanation, the hand movements of one user (meeting participant) are detected. In reality, multiple users may attempt operations one after the other. In such a situation, the operation of the first user detected is given priority, and the operation of the next user is detected only after that user's operation is completed, i.e., after that user's hand leaves the screen.

[0034] The input system according to the present invention allows participants to operate the displayed content without leaving their seats when using a large-screen display device in a meeting with many participants. At this time, participants do not need to use special devices, and can operate the content just by pointing with their fingers, just as if they were using a touch panel in front of them.

[0035] Although the present invention has been described in detail with reference to the embodiments, it will be apparent to those skilled in the art that the present invention is not limited to the embodiments described herein. The device of the present invention can be implemented in various modified and altered forms without departing from the spirit and scope of the present invention as defined by the claims. Therefore, the description of the present application is intended to be illustrative and explanatory and is not intended to be limiting of the present invention.

[0036] In the above embodiment, a digital whiteboard using a liquid crystal display is used as the display device. However, the present invention is not limited to this. For example, a projector that connects a personal computer and projects its screen onto a screen may also be used.

[0037] In the above embodiment, the depth sensor of the 3D camera is used to detect the three-dimensional positions of the viewpoint and the hand, but the present invention is not limited to this. For example, Lidar, an infrared distance sensor, or an ultrasonic distance sensor may be used.

[0038] Furthermore, in the above embodiment, the 3D camera equipped with a depth sensor is installed above the display device, but the present invention is not limited to this. For example, the camera may be installed below the display device or on the left or right side, as long as it is positioned so that the faces of the participants can be photographed.

[0039] 1. Digital whiteboard 3. Personal computer 5. 3D camera

Claims

1. An input system comprising: a display device; position identification means for identifying the three-dimensional position of the viewpoint and hand of a user facing the display device and looking at the display screen of the display device; and an information processing device connected to the display device and the position identification means, wherein the information processing device outputs a video signal to the display device, causes the display device to display an image, identifies a position on the display screen on an extension of a straight line connecting the viewpoint and the position of the user's hand as an operation position, and recognizes the operation position and the movement of the user's hand as input to the information processing device.

2. The input system according to claim 1, wherein the position specifying means is a 3D camera equipped with a depth sensor.

3. The input system according to claim 2, wherein the position of the user's hand identified by the 3D camera is the fingertip.

4. The input system according to claim 1, wherein the display device is a liquid crystal display or a projector.

Citation Information

Patent Citations

  • Gesture detection device and gesture detection method

    JP2020086913A