3D Object Grasping Program

The three-dimensional object grasping program addresses the challenge of accurately manipulating 3D objects in aerial displays by using a motion sensor and machine learning to recognize gestures, enabling natural and intuitive interaction.

JP7893445B1Active Publication Date: 2026-07-22INTERMAN CORP
View PDF 9 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
INTERMAN CORP
Filing Date
2025-08-28
Publication Date
2026-07-22

AI Technical Summary

Technical Problem

Existing aerial image display devices struggle to allow users to accurately manipulate three-dimensional objects due to the discrepancy between perceived and actual positions, as 3D objects are displayed on a two-dimensional plane, making it difficult to grasp them as intended.

Method used

A three-dimensional object grasping program that uses a motion sensor to detect hand joint and fingertip coordinates, applies a machine learning model to recognize predetermined gestures, and locks the relative position of the 3D object to the user's hand when a valid gesture is detected, allowing natural interaction with the object.

Benefits of technology

Enables easy and natural grasping of three-dimensional objects displayed in mid-air by fixing the object's position relative to the user's hand, enhancing the interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007893445000001_ABST
    Figure 0007893445000001_ABST
Patent Text Reader

Abstract

This program provides a 3D object grasping program that allows users to easily grasp 3D objects displayed in the air on an aerial image display device, and then rotate and move them up, down, left, and right. [Solution] A machine learning model is built in advance, which has learned the general hand shape used when grasping an object as a gesture. Then, using this machine learning model, it is determined whether the coordinates of the user's hand joints and fingertips detected by the motion sensor correspond to a grasping gesture in the vicinity of a 3D object displayed in the air. If it corresponds to a grasping gesture, the relative position of the 3D object to the user's hand is fixed, and the 3D object is moved in conjunction with the user's hand movements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0005]

[0001] The present invention relates to a three-dimensional object gripping program that enables easy and natural gripping of three-dimensional objects.

Background Art

[0002] An aerial video display device is a device that displays an object in the empty space. For example, Patent Document 1 shows an aerial video display device that displays a video in the air by utilizing the reflection of light. When displayed as an aerial image using an aerial video display device, a sense of three-dimensionality can be felt because it appears to float from the surroundings. In particular, when an object in a 3D space is displayed by perspective projection, an illusion can be created as if the three-dimensional object actually exists there.

[0003] Many of the currently commercialized aerial video display devices are provided with a three-dimensional motion sensor composed of an infrared LED and an infrared camera. By detecting the movement of the user's hand near the aerial image using this three-dimensional motion sensor, an input interface via the aerial image can be implemented. For example, Patent Documents 2 and 3 show technologies that enable gesture input using the hand in front of the aerial image of an aerial video display device.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Patent Document 3

Summary of the Invention

Problems to be Solved by the Invention

[0005] As a method of manipulating an aerial image using hand gestures, we consider moving a displayed 3D object. In this case, it is necessary to implement interaction between the 3D object and the hand.

[0006] The most direct implementation would be to detect the user's hand movements and simulate their interaction with 3D objects. However, although 3D objects are displayed three-dimensionally as objects in 3D space using perspective projection, they are actually displayed on a two-dimensional plane, making it difficult for the user to accurately perceive their position. Therefore, it becomes difficult for the user to manipulate 3D objects as intended.

[0007] Therefore, the objective of the present invention is to provide a three-dimensional object grasping program that allows users to directly touch and manipulate three-dimensional objects displayed in the air. [Means for solving the problem]

[0008] To solve the above problems, a three-dimensional object grasping program according to one aspect of the present invention is a program executed on an aerial image display device comprising a computer, an image system controlled by the computer that projects an image into the air, and a motion sensor that detects the coordinates of the user's hand joints and fingertips in the vicinity of the aerial image projected by the image system and transmits the detection results to the computer, wherein the program includes the steps of: controlling the image system to project a three-dimensional object as an aerial image to the computer; determining whether the coordinates of the user's hand joints and fingertips acquired by the motion sensor correspond to a predetermined gesture; and determining whether the user's hand movements correspond to a predetermined gesture If it is determined that the gesture corresponds to a certain action, the relative position of the 3D object to the position of the user's hand is fixed, and the 3D object is moved in conjunction with the movement of the user's hand. If it is determined that the gesture corresponds to a predetermined gesture, the relative position of the 3D object to the position of the user's hand is fixed, and the step of determining whether the coordinates of the joints and fingertips of the user's hand correspond to the predetermined gesture is continued. If it is determined that the coordinates of the joints and fingertips of the user's hand do not correspond to the predetermined gesture, the relative position of the 3D object to the position of the user's hand is released, and whether the gesture corresponds to a certain action is determined. The determination is made by a machine learning model that has been pre-trained by associating different gestures with different 3D objects. It is characterized by the following:

[0009] Furthermore, in one embodiment, the determination that the coordinates of the user's hand joints and fingertips correspond to the predetermined gesture is performed only when the user's hand is within a predetermined determination area that includes the three-dimensional object.

[0010] Furthermore, in one embodiment, the aerial image display device is equipped with a speaker, and when the relative position of the three-dimensional object with respect to the position of the user's hand is fixed, it simultaneously generates a sound effect. [Effects of the Invention]

[0011] According to the three-dimensional object grasping program of the present invention, a three-dimensional object displayed in mid-air can be grasped easily and naturally. [Brief explanation of the drawing]

[0012] [Figure 1] Figure 1 is a perspective view showing an aerial image display device 1 for executing a three-dimensional object grasping program according to an embodiment of the present invention. [Figure 2] Figure 2 is a cross-sectional view along the line A-A in Figure 1, showing the aerial image display device 1 in use. [Figure 3] Figure 3 shows the action of grasping a three-dimensional object. [Figure 4] Figure 4 shows an example of the shape of a hand grasping a three-dimensional object. [Figure 5] Figure 5 shows another example of the shape of a hand grasping a three-dimensional object. [Figure 6] Figure 6 shows yet another example of the shape of a hand grasping a three-dimensional object. [Figure 7] Figure 7 shows yet another example of the shape of a hand grasping a three-dimensional object. [Figure 8] Figure 8 shows yet another example of the shape of a hand grasping a three-dimensional object. [Figure 9] Figure 9 shows yet another example of the shape of a hand grasping a three-dimensional object. [Figure 10] Figure 10 shows another example of the shape of a hand grasping a three-dimensional object. [Figure 11] Figure 11 shows yet another example of the shape of a hand grasping a three-dimensional object. [Figure 12] Figure 12 shows yet another example of the shape of a hand grasping a three-dimensional object. [Figure 13] Figure 13 shows yet another example of the shape of a hand grasping a three-dimensional object. [Figure 14]FIG. 14 is a diagram showing yet another example of the hand posture for gripping a three-dimensional object. [Figure 15] FIG. 15 is a diagram showing yet another example of the hand posture for gripping a three-dimensional object. [Figure 16] FIG. 16 is a flowchart for explaining the operation of the three-dimensional object gripping program according to an embodiment of the present invention.

MODE FOR CARRYING OUT THE INVENTION

[0013] Hereinafter, an embodiment of the three-dimensional object gripping program according to the present invention will be described with reference to the accompanying drawings. In the following embodiment, as an example of the implementation of the three-dimensional object gripping program, an airborne video display device provided with a recursive transmissive optical pixel device is shown.

[0014] FIG. 1 is a perspective view showing an airborne video display device 1 for executing the three-dimensional object gripping program according to an embodiment of the present invention. FIG. 2 is a cross-sectional view taken along line A-A of FIG. 1 showing the airborne video display device 1 in use.

[0015] As shown in FIGS. 1 and 2, the airborne video display device 1 includes a liquid crystal display 10 and an optical plate 20 as essential components. The airborne video display device 1 further includes an information processing device 30 for controlling the liquid crystal display 10, a speaker 40, and the like. Here, the liquid crystal display 10, the speaker 40, and the information processing device 30 are housed inside the lower housing 12 and are connected to each other by signal lines (omitted in the figure).

[0016] The information processing unit 30 is essentially a small computer and consists of a CPU (Central Processing Unit), ROM (Read Only Memory), RAM (Random Access Memory), a storage device for storing various programs and data, and an input / output interface. The input / output interface may include, for example, a USB port or a wireless LAN such as Wi-Fi. The information processing unit 30 outputs a video signal to the liquid crystal display 10, which displays the basis for the aerial image. The information processing unit 30 also outputs an audio signal to the speaker 40, which generates guidance voices and sound effects. This information processing unit 30 can also utilize commercially available general-purpose small personal computers or general-purpose tablets.

[0017] The liquid crystal display 10 is housed in a lower housing 12 that is rectangular in shape when viewed from above and has an open top, and is supported almost horizontally with its display screen facing upwards. The optical plate 20 is fitted into the upper housing 14 with its incident surface 21 facing downwards, so as to face the display screen of the liquid crystal display 10 at an angle. Here, the optical plate 20 and the liquid crystal display 10 are fixed at an angle of approximately 45 degrees.

[0018] As such an optical plate 20, for example, a retrotransmitting optical imaging element (two-plane orthogonal reflector) described in Japanese Patent Application Publication No. 2011-175297 can be used. This optical imaging element is realized by arranging a large number of mutually orthogonal planar light reflecting parts at a constant pitch. In addition, structures such as a two-plane corner reflector, in which reflective surfaces are formed on the sides of a square-shaped hole, as described in Japanese Patent No. 4900618, may also be used.

[0019] Furthermore, the upper housing 14 that constitutes the aerial image display device 1 can be easily separated from the lower housing 12. Separating the upper housing 14 makes it easier to perform maintenance and adjustments on the inside of the lower housing 12, and also reduces the height during transportation.

[0020] As shown in Figure 2, light from the display screen of the liquid crystal display 10 enters the optical plate 20, is reflected twice inside the optical plate 20, and exits to the opposite side. As a result, an aerial image region G is formed as a real image in the space opposite the optical plate 20, with the optical plate 20 as the plane of symmetry. In this case, in order to realize a clearer aerial image region G, it is desirable that no external light is added to the light from the liquid crystal display 10.

[0021] In this context, the aerial image area G represents the image area displayed in the air when an image is displayed across the entire screen of the liquid crystal display 10. Since the screen of the liquid crystal display 10 is typically rectangular, the aerial image area G is also rectangular. In this specification, the aerial image area G is also referred to as the display surface of the aerial image display device 1.

[0022] Furthermore, a three-dimensional motion sensor 7, consisting of an infrared LED 72 and a pair of infrared cameras 74, is provided on the front side of the lower housing 12. This motion sensor 7 precisely detects the user's hand movements in three dimensions. Specifically, it can detect the position of the fingertips and joints of each finger on both the left and right hands.

[0023] The detection range of the motion sensor 7 is defined by the emission angle of the infrared LED 72 and the field of view of the infrared camera 74. Such a motion sensor 7 can be a commercially available product, such as the Leap Motion Controller 2 from Ultraleap.

[0024] In particular, the three-dimensional motion sensor 7 is mounted in a way that allows its tilt angle to be adjusted. That is, the motion sensor 7 is housed in a housing that allows it to rotate within a certain angular range (here, ±20 degrees) with the axis of rotation perpendicular to the plane of the cross-sectional view in Figure 2. Here, the center of the detection range of the motion sensor 7 is adjusted to be shifted towards the user side from the center of the aerial image region G. By making this adjustment, the effective operation detection range can be maximized.

[0025] Traditionally, such aerial image display devices have often been used as administrative input devices for reception systems and the like. That is, by displaying a contactless interface screen in the aerial image area G, users can operate the contactless interface screen using gestures, such as touching it with their fingers. This operation is detected by a motion sensor 7 consisting of an infrared LED 72 and an infrared camera 74, and a corresponding operation signal is sent to an information processing device 30 for predetermined processing. The operation interface can include controls such as buttons, checkboxes, and drop-down menus. Therefore, a completely contactless interface can be achieved with the same ease of use as a conventional touch panel.

[0026] In this invention, aerial image display devices are used to display 3DCG-based three-dimensional objects for entertainment or educational purposes. The three-dimensional objects are created by projecting a computer-defined three-dimensional object onto the screen of the liquid crystal display 10 using perspective projection, resulting in an image with a sense of depth (three-dimensionality). Furthermore, because the screen of the liquid crystal display 10 is projected in front of the user, this computer-defined three-dimensional space corresponds to the three-dimensional space in front of the user.

[0027] Here, a sphere S is used as an example of a three-dimensional object. However, although Figure 2 depicts the sphere S viewed from the side, the aerial image display device 1 would not actually appear this way when viewed from the side. The original liquid crystal display 10 is a plane, and the perspective projection of the sphere S is projected onto the plane's image-forming surface. Here, the sphere S in three-dimensional space is virtually depicted as a perspective projection.

[0028] This 3D object S can be grasped by hand, rotated, and moved up, down, left, and right. If you were to grasp a real, physical 3D object, it would naturally be constrained in your hand and move with your hand's movements. To achieve this with a 3D object S as an aerial projection, the relative position of the 3D object and the hand becomes fixed (locked) the moment the hand grasps the 3D object S. Therefore, it is necessary to clearly determine the transition from a state where the hand is not grasping the 3D object S to a state where the hand is grasping the 3D object S (locked state).

[0029] Here, three-dimensional objects are displayed in 3D space using perspective projection, making them appear as if they are actually there. However, regarding the physiological factors of stereoscopic vision, such as binocular parallax, convergence, and accommodation, an aerial image gives the same perception as a flat image. Therefore, even if the image is perceived as three-dimensional, when manipulating the three-dimensional object with one's hands, a discrepancy between the perceived position in the depth direction and the actual object is likely to occur.

[0030] Therefore, it is not possible to expect the user to grasp the 3D object at the precise location. Thus, the transition from a state where the hand is not grasping sphere S to a state where the hand is grasping sphere S is determined by whether or not a predetermined gesture has been performed. The predetermined gesture corresponds to the overall shape and movement of the hand when a person grasps an object.

[0031] Gestures can be broadly categorized into static gestures, which essentially involve no movement, and dynamic gestures, which require specific movements. Examples of static gestures include the victory (V) gesture and the OK gesture. Examples of dynamic gestures include the movement of the index finger or clicking a finger.

[0032] Generally, grasping an object involves basic actions such as opening and closing the hand, as shown in Figure 3. In other words, a gesture that shows the transition from the open state shown in Figure 3(A) to the closed state shown in Figure 3(B) is considered a dynamic gesture for grasping an object. In this case, video is used as the training data for the dynamic gesture.

[0033] However, as shown in Figure 1, the shape of a user's hand when grasping an object has certain characteristics. Therefore, it is perfectly functional to learn only image data as static gestures, omitting the user's hand movements. In addition to the shapes of a user's hand shown in Figures 1 and 2, other hand shapes that represent grasping actions include, for example, the hand shape that grasps from above as shown in Figure 4, and the hand shape that supports from below as shown in Figure 5.

[0034] There are many other examples of how a person's hand can form when grasping an object. In Figure 6, the thumb and index finger form an O-ring. In Figure 7, the thumb and middle finger form an O-ring. Furthermore, in Figure 8, the thumb and ring finger form an O-ring. Furthermore, in Figure 9, the thumb and little finger form an O-ring. Furthermore, in Figure 10, the thumb and the other four fingers form an O-ring. Furthermore, in Figure 11, the thumb and the three fingers other than the little finger form an O-ring. Furthermore, in Figure 12, the thumb, index finger, and middle finger form an O-ring. Figure 13 shows a clenched fist, which is also a hand shape used when grasping an object.

[0035] Furthermore, when grasping an object, the hand can also be used in a gesture involving both hands. In this case, the motion sensor 7 acquires the position of the fingertips and joints of each finger when grasping with both hands, and uses this as training data for the machine learning model. Figure 14 shows a hand gesture as if grasping a ball with both hands. Figure 15 shows a hand gesture as if holding a rice ball between both hands.

[0036] Here, we will use still images as training data and perform training on static gestures to build a machine learning model. However, the present invention is not limited to this. In fact, it is also possible to train the model with both static and dynamic gestures, in which case a more versatile implementation will be possible.

[0037] These gestures are determined by a machine learning model based on a neural network. In other words, a machine learning model is trained using numerous videos of actual hand postures (shapes = poses) and gestures as training data to build a system that can determine the intended operation. Specific frameworks that can be used to build such machine learning models and for real-time recognition of hand gestures include Tensorflow and Mediapipe. The training data input to the machine learning model consists of numerous videos of actual human gestures (grasping movements).

[0038] The determination of whether a hand is grasping a 3D object S is performed only when the hand is located near the 3D object. For example, consider a spherical determination area that includes the entire 3D object, and perform a gesture determination when either the fingertips or joints of each finger are located inside this spherical determination area. Here, since the 3D object S is a sphere, consider a spherical determination area that shares a center with this sphere and has a radius greater than or equal to the radius of sphere S. For example, the radius of this spherical determination area is set to 1.2 times the radius of sphere S.

[0039] When the predetermined gesture described above is detected in the detection area, that is, when it is determined that the user's hand is grasping the 3D object, the relative position of the 3D object to the position of the hand is fixed, and the display of the 3D object is moved in conjunction with the hand movement detected by the motion sensor 7. The reference position for the relative position here is the center position (center of gravity) of the 3D object and the center position of the hand (for example, the joint position at the base of the fingers detected by the motion sensor 7).

[0040] In parallel with this, it is continuously confirmed that the user's hand is grasping a three-dimensional object. For example, every 0.1 seconds, it is repeatedly checked whether a predetermined gesture is detected. In the case of dynamic gestures, it is repeatedly checked whether the final form of the gesture (for example, the closed form shown in Figure 3(B)) is detected.

[0041] If a specific gesture is no longer detected, it is determined that the hand is no longer grasping the 3D object. For example, if the open state shown in Figure 3(A) is detected, it is determined that the 3D object has been released. The lock on the relative position between the 3D object and the hand is then released. Once the lock is released, the 3D object performs inertial motion; that is, it either stays in place according to its previous movement or moves at a constant speed.

[0042] The above process will be explained with reference to the flowchart in Figure 16. When the system starts up, the motion sensor 7 acquires the coordinate positions of the fingertips and joints of each of the user's fingers (step S1). Next, it is determined whether the relative position of the 3D object to the hand position is fixed (locked) or not (step S2). If the 3D object is not locked (NO in step S2), it is determined whether any of the acquired fingertip or joint positions of each finger fall within a spherical determination area that includes the entire 3D object (step S3). If none of the fingertip or joint positions fall within the determination area (NO in step S3), the process returns to step S1 and the positions of the fingertips and joints of each finger are acquired again.

[0043] If either the fingertip or joint of each finger is within the detection area (YES in step S3) or if the 3D object is locked (YES in step S2), the machine learning model determines whether or not a gesture indicating a grasping action is detected (step S4).

[0044] It is recommended that the detection of a grasping gesture be considered valid only if it persists for a certain period of time (for example, 3 seconds). This is because unintentional actions by the user that are not intended may be detected as grasping gestures for a brief moment.

[0045] If no gesture indicating a grasping action is detected (NO in step S4), the 3D object is unlocked if it is locked (step S5), or if it is not locked, it is kept locked, and the process returns to step S1 to acquire the positions of the fingertips and joints of each finger again.

[0046] If a gesture indicating a grasping action is detected (YES in step S4), the system determines that the user's hand is grasping the 3D object and locks it if it is not already locked (step S6). If the 3D object is already locked, the lock is maintained. Next, the display of the 3D object is moved in conjunction with the hand movement detected by the motion sensor 7 (step S7).

[0047] Furthermore, in step S6, when transitioning from an unlocked state to a locked state, a sound effect is played to indicate this. Also, in step S5, when transitioning from a locked state to an unlocked state, a different sound effect is played. This allows users to confirm, through hearing, that the device has been locked or unlocked. [Industrial applicability]

[0048] The 3D object grasping program according to the present invention allows for easy and natural grasping of 3D objects displayed in mid-air, enabling experiences impossible with conventional displays. For example, it becomes possible to directly pick up and manipulate 3D data of cultural artifacts, or to use the floating sensation of an aerial image display device as input in games where the user directly touches and moves the image.

[0049] Although the present invention has been described in detail above with reference to examples, it will be clear to those skilled in the art that the present invention is not limited to the examples described herein. The apparatus of the present invention can be implemented in modified and altered forms without departing from the spirit and scope of the invention as defined by the claims. Therefore, the description herein is for illustrative purposes only and is not intended to be restrictive in any way to the present invention.

[0050] For example, in the above embodiment, a gesture is detected when either the fingertip or joint of each finger enters the spherical detection area. However, the detection area may be an ellipsoid, a cube, or something other than a sphere. Alternatively, a detection area may not be provided, and a gesture may be detected when either the fingertip or joint of each finger touches a three-dimensional object.

[0051] Furthermore, in the above embodiment, gesture detection is performed independently of the 3D object. However, the gestures detected can be changed depending on the 3D object. In fact, the shape of a hand holding a ball is different from the shape of a hand holding a coffee cup. Therefore, when a ball is displayed, a machine learning model trained on the shape of a hand holding a ball can be used, and when a coffee cup is displayed, a machine learning model trained on the shape of a hand holding a coffee cup can be used.

[0052] Furthermore, in the above embodiment, the lock state can be confirmed by the output of a sound effect, but separately or in addition to this, the lock state may also be confirmed by changing the display of the 3D object.

[0053] For example, when a 3D object is locked, its brightness and color can be changed, and when unlocked, they can be restored to their original state, thereby informing the user of the lock status. In this case, the brightness and color could be changed only in the vicinity of the user's hand.

[0054] Furthermore, in addition to changes in brightness, it may be possible to confirm the locked state by adding wave-like fluctuations to the surface of the three-dimensional object.

[0055] Furthermore, the above embodiment employs an aerial image display device using a retrotransmissive optical imaging element. However, the present invention is not limited to this, and for example, an aerial image display device using a retroreflective optical imaging element may also be employed. [Explanation of symbols]

[0056] 1. Aerial Image Display Device 7. Three-dimensional motion sensor 10 LCD displays 12 Lower enclosure 14 Upper chassis 20 Optical Plates 21 Entrance plane 30 Information Processing Devices 40 Speakers 72 Infrared LEDs 74 Infrared Cameras S 3D object

Claims

1. A program executed on an aerial image display device comprising a computer, an image system controlled by the computer that projects images into the air, and a motion sensor that detects the coordinates of the user's hand joints and fingertips in the vicinity of the aerial image projected by the image system and transmits the detection results to the computer, wherein the computer... The steps include controlling the aforementioned video system to project a three-dimensional object as an aerial image, The steps include determining whether the coordinates of the user's hand joints and fingertips acquired by the motion sensor correspond to a predetermined gesture, If it is determined that the user's hand movement corresponds to a predetermined gesture, the relative position of the three-dimensional object with respect to the user's hand position is fixed, and the three-dimensional object is moved in conjunction with the user's hand movement. The program is characterized in that, even after determining that the object corresponds to a predetermined gesture and fixing the relative position of the three-dimensional object with respect to the position of the user's hand, it continues to perform the step of determining whether the coordinates of the user's hand joints and fingertips correspond to the predetermined gesture, and if it is determined that the coordinates of the user's hand joints and fingertips do not correspond to the predetermined gesture, the fixing of the relative position of the three-dimensional object with respect to the position of the user's hand is released, and the determination of whether or not the object corresponds to the gesture is performed by a machine learning model that has been trained in advance by associating different gestures with different three-dimensional objects.

2. The program according to claim 1, characterized in that the determination that the coordinates of the user's hand joints and fingertips correspond to the predetermined gesture is performed only when the user's hand is within a predetermined determination area including the three-dimensional object.

3. The program according to claim 1, wherein the aerial image display device is equipped with a speaker, and when the relative position of the three-dimensional object with respect to the position of the user's hand is fixed, it simultaneously generates a sound effect.