Information processor, control method for information processor, and program

The information processing apparatus enhances MR usability by arranging operation-compatible and display-only models in detected real-space regions, addressing tactile feedback and fatigue issues in hand tracking.

JP2025094548APending Publication Date: 2025-06-25CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023210168
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-13
Publication Date
2025-06-25

AI Technical Summary

Technical Problem

Existing hand tracking technologies in mixed reality (MR) systems using Head Mounted Displays (HMDs) lack usability improvements, particularly in providing tactile feedback and preventing user fatigue during virtual object interaction.

Method used

An information processing apparatus that detects regions in the real space for fixing virtual objects, arranges operation-compatible and display-only models in corresponding virtual spaces, and determines user interactions based on these models to enhance usability.

Benefits of technology

Improves usability by providing tactile feedback and reducing fatigue through operation-compatible models on real surfaces and display-only models visible without neck strain.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025094548000001_ABST
    Figure 2025094548000001_ABST
Patent Text Reader

Abstract

To provide an information processor, a control method for an information processor and a program which improve usability when a user operates a user interface indicated by a virtual object by hand tracking.SOLUTION: An information processor 102, which is a control device for controlling an HMD 101, includes: a virtual UI fixed region candidate detection part 209 that detects a virtual UI fixed region candidate, which allows fixation of a virtual UI indicating a user interface, from a real space; a virtual object separation arrangement control part 210 that arranges an operation combined-use model, constituted by a hand model corresponding to hand tracking and a virtual UI, in a region of virtual space corresponding to the virtual UI fixed region candidate, and arranges a display-only model, constituted by the hand model corresponding to hand tracking and the virtual UI, in a region of virtual space different from the region where the operation combined-use model is arranged; and a contact determination part 208 that performs determination related to operation of the user interface on the basis of the operation combined-use model.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, a control method for the information processing apparatus, and a program.

Background Art

[0002] One of the technologies for making a user feel a space different from the real space using an HMD (Head Mounted Display) is the technology of Mixed Reality (MR). In mixed reality, it is possible to present a composite reality space in which the real space and the virtual space are fused to a user wearing an HMD. Further, in mixed reality, a user wearing an HMD can perform various operations in the composite reality space, for example, by hand tracking. In hand tracking, a user's hand is recognized by a camera mounted on the HMD, and the position and posture of each joint point of the user's hand are estimated by pose estimation.

[0003] As a result, the CG (Computer Graphics) of the user's hand is arranged in the composite reality space. In the composite reality space, for example, when a user operates the CG of a UI (User Interface) arranged in the air by hand tracking, the user has to operate with the hand floating. Therefore, the user does not obtain feedback such as a keystroke feeling, and the arm may get tired. Hereinafter, the CG of the UI is referred to as a "virtual UI", and the CG of the user's hand is referred to as a "hand model".

[0004] Regarding the above points, one of the technologies applicable to hand tracking is a technology called plane detection. Plane detection is a technology for extracting the same plane by detecting geometric feature points from a camera image and clustering the detected feature points. In the composite reality space, for example, a virtual UI is arranged on the plane of a real object detected by plane detection with respect to the real space, and when a user operates the virtual UI by hand tracking, the user can obtain a keystroke feeling due to the contact between the fingertips and the real object.

[0005] FIG. 9 is a diagram showing a situation in which a user is operating a virtual UI 900 arranged on the plane of a real object table by hand tracking. In FIG. 9, it is assumed that a user wearing an HMD 901 is sitting on a chair and facing the table. In the display viewing angle 902 of the HMD 901, the virtual UI 900 is displayed along the table plane area 903 detected by plane detection. Further, in the display viewing angle 902 of the HMD 901, a hand model 904 is displayed by hand tracking. As a result, the user can operate the virtual UI 900 by hand tracking while placing the hand on the table, so that the fatigue feeling is significantly reduced.

[0006] Furthermore, when the user operates the virtual UI 900, the user can obtain a key pressing feeling by the contact between the fingertips and the table. In FIG. 9, a keyboard is shown in CG as the virtual UI 900, but the same effect can be achieved with the CG of a UI other than the keyboard. However, in the case of FIG. 9, the user cannot see the virtual UI 900 unless the line of sight is directed downward to the table. Therefore, in the composite reality space, for example, if a CG different from the virtual UI 900 is arranged in front of the user, the user cannot see the arranged different CG when operating the virtual UI 900. In addition, if the user continues to direct the line of sight downward to the table in order to operate the virtual UI 900, the neck may get tired.

[0007] On the other hand, Patent Document 1 discloses a technique of reflecting the input content of pointer operation on an electronic device in a part of the display in an HMD. In the technique disclosed in Patent Document 1, since an electronic device held by the user is used, the user during operation can obtain feedback such as vibration from the electronic device.

Prior Art Documents

Patent Documents

[0008]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0009] However, since the technology disclosed in Patent Document 1 assumes the use of an electronic device, it is difficult to directly apply it to hand tracking.

[0010] The present invention has been made in view of the above problems. An object of the present invention is to provide an information processing apparatus, a control method of the information processing apparatus, and a program that can improve the usability of a user when operating a user interface represented by a virtual object by hand tracking.

Means for Solving the Problems

[0011] In order to achieve the above object, an information processing apparatus of the present invention is an information processing apparatus that enables a user to operate a user interface represented by a virtual object by hand tracking in a space where the real space and the virtual space are fused, and includes a detection unit that detects one or more regions in the real space where the virtual object representing the user interface can be fixed, a first arrangement unit that arranges an operation-dedicated model composed of a first hand model corresponding to the hand tracking and a first virtual object representing the user interface in a first region of the virtual space corresponding to any one of the one or more regions detected by the detection unit, a second arrangement unit that arranges a display-dedicated model composed of a second hand model corresponding to the hand tracking and a second virtual object representing the user interface in a second region of the virtual space different from the first region, and an operation determination unit that makes a determination regarding the operation of the user interface based on the operation-dedicated model.

Effects of the Invention

[0012] According to the present invention, it is possible to improve the usability of a user interface represented by a virtual object when the user operates it by hand tracking.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Modes for Carrying Out the Invention

[0014] Hereinafter, each embodiment of the present invention will be described in detail with reference to the drawings. However, the configurations described in the following embodiments are merely examples, and the scope of the present invention is not limited by the configurations described in the embodiments. For example, each part constituting the present invention can be replaced with any configuration that can exhibit the same function. Also, an arbitrary component may be added. Also, any two or more configurations (features) in each embodiment can be combined. Furthermore, in the accompanying drawings, the same or similar configurations are denoted by the same reference numerals, and redundant descriptions are omitted.

[0015] <First Embodiment> Hereinafter, with reference to FIGS. 1 to 6, the first embodiment will be described. FIG. 1 is a diagram showing an information processing system 100 according to the first embodiment. The information processing system 100 presents a composite reality space in which the real space and the virtual space are fused to the user. As shown in FIG. 1, the information processing system 100 includes an HMD 101 and an information processing device 102. The HMD 101 is a head-mounted display device that is worn on the user's head. In the HMD 101, a composite reality space is presented to the user by displaying a composite image. The composite image is an image in which an imaging image (that is, an image of the real space) obtained by the HMD 101 imaging a space similar to the space the user is looking at and an image of a virtual space in which CGs in a form corresponding to the posture of the HMD 101 are arranged are synthesized, and is generated by the information processing device 102. Hereinafter, the image of the virtual space in which CGs in a form corresponding to the posture of the HMD 101 are arranged will be referred to as a "CG image".

[0016] The information processing device 102 is a control device that controls the HMD 101. The information processing device 102 is, for example, a smartphone, a tablet terminal, a personal computer, or a digital camera. The information processing device 102 is connected to the HMD 101 wirelessly or by wire. The information processing device 102 generates a composite image by synthesizing the imaging image and the CG image as described above. The information processing device 102 transmits the generated composite image to the HMD 101. Each component of the information processing device 102 may be included in the HMD 101. In that case, the HMD 101 constitutes the information processing device.

[0017] Figure 2 is a block diagram showing the internal configurations of the HMD 101 and the information processing apparatus 102. First, the internal configuration of the HMD 101 will be described. As shown in FIG. 2, the HMD 101 includes an HMD control unit 201, an imaging unit 202, an image display unit 203, an attitude sensor unit 204, and a hand joint detection unit 205. The HMD control unit 201 controls each component of the HMD 101. The HMD control unit 201 acquires a composite image (an image in which the captured image and the CG image are combined) from the information processing apparatus 102. The HMD control unit 201 displays the acquired composite image on the image display unit 203. Thereby, the user can view the composite image displayed on the image display unit 203 by wearing the HMD 101. Further, the user can experience various composite reality spaces in which the real space and the virtual space are fused.

[0018] The imaging unit 202 has two cameras. In the HMD 101, the two cameras are provided at respective positions where the left and right eyes of the user are close to each other when the HMD 101 is worn, in order to image a space similar to the space the user is looking at. Images of real objects (that is, real objects included in a space similar to the space the user is looking at) captured by the two cameras are output as captured images to the information processing apparatus 102. At this time, in the imaging unit 202, the two cameras can acquire information on the distance from the two cameras to the real object as distance information by distance measurement using a stereo camera. The distance information acquired by the two cameras is output to the information processing apparatus 102.

[0019] The image display unit 203 displays the composite image. The image display unit 203 includes, for example, an organic EL panel or a liquid crystal panel. In the image display unit 203 when the user is wearing the HMD 101, an organic EL panel or the like is arranged in front of each of the user's left and right eyes. The attitude sensor unit 204 detects the attitude and position of the HMD 101. Further, the attitude sensor unit 204 detects the attitude of the user (that is, the user wearing the HMD 101) corresponding to the attitude and position of the HMD 101. The attitude sensor unit 204 includes an inertial measurement unit (IMU) composed of an acceleration sensor, an angular acceleration sensor, and a geomagnetic sensor. The attitude sensor unit 204 outputs information regarding the attitude and position of the HMD 101 and the attitude of the user as attitude information to the information processing device 102.

[0020] The hand joint point detection unit 205 detects the user's hand from the two camera images acquired by the imaging unit 202, and detects the position and attitude of each joint point of the user's hand. For the detection of the user's hand and the detection of the position and attitude of each joint point of the user's hand, for example, known object recognition and pose estimation of machine learning using a convolutional neural network are used. Information regarding the detected user's hand and information regarding the position and attitude of each joint point of the detected user's hand are output to the information processing device 102 for generating a hand model or the like. Note that the position information in the depth direction of each joint point of the user's hand is obtained by calculating the distance from the imaging unit 202 to each joint point of the hand by stereo triangulation using the two camera images acquired by the imaging unit 202. The obtained position information in the depth direction of each joint point of the user's hand is output to the information processing device 102 for arranging the hand model or the like.

[0021] Next, the internal configuration of the information processing apparatus 102 will be described. The information processing apparatus 102 includes a control unit 206, a content DB 207, a contact determination unit 208 (operation determination means), a virtual UI fixed area candidate detection unit 209 (detection means), and a virtual object separation and arrangement control unit 210 (first arrangement means) (second arrangement means). The control unit 206 controls each component of the HMD 101 and the information processing apparatus 102. The control unit 206 receives the captured image and distance information acquired by the imaging unit 202, the attitude information acquired by the attitude sensor unit 204, etc. from the HMD 101. The control unit 206 performs image processing on the captured image to cancel the aberration in the optical systems of the imaging unit 202 and the image display unit 203.

[0022] Furthermore, as described above, the control unit 206 synthesizes the captured image and the CG image to generate a composite image. The control unit 206 transmits the generated composite image to the HMD control unit 201. Note that the control unit 206 controls the position, orientation, and size of the CG in the composite image based on the information (distance information and attitude information) acquired by the HMD 101. For example, when the control unit 206 arranges the CG near a specific real object existing in the real space in the composite reality space shown by the composite image, the closer the distance between the specific real object and the imaging unit 202, the larger the CG. In this way, by controlling the position, orientation, and size of the CG, the control unit 206 can generate a composite image such that a virtual object that does not exist in the real space (i.e., the virtual object shown by the CG) seems to exist in the real space.

[0023] The content DB 207 is a storage unit that stores information such as CG. Note that the control unit 206 can switch the CG read from the content DB 207, that is, switch the CG used for the CG image. The contact determination unit 208 determines whether the CGs arranged in the virtual space are in contact with each other in the composite reality space shown by the composite image. Also, physical calculations such as changes in the position and attitude of the CG associated with the determination are performed by the contact determination unit 213. In this embodiment, the contact determination between the virtual UI and the hand model is performed by the contact determination unit 208.

[0024] The virtual UI fixing area candidate detection unit 209 detects a virtual UI fixing area candidate from the real space. The virtual UI fixing area candidate is an area in which the virtual UI can be fixed in the real space, such as a surface area of ​​a real object. A specific example of a virtual UI fixing area candidate is a flat surface such as a table. However, the virtual UI fixing area candidate does not have to be a flat surface such as a table, and may be an area in which the virtual UI can be fixed as described above, such as a free-form surface such as the surface of a user's arm.

[0025] The virtual object separation and placement control unit 210 generates a virtual UI and a hand model that are dedicated to display and a hand model that are also used for operation. The virtual object separation and placement control unit 210 places the display-only virtual UI and hand model in a region that is, for example, directly in front of the user (for example, away from both eyes of the user) in the virtual space. In this case, the user wearing the HMD 101 can see the display-only virtual UI and hand model when facing forward.

[0026] The virtual UI and hand model also used for operation are used for contact determination by the contact determination unit 208. The virtual object separation and placement control unit 210 places the virtual UI and hand model also used for operation in an area in the virtual space that corresponds to the virtual UI fixing area in the real space. In this case, when the user wearing the HMD 101 faces the direction in which the virtual UI fixing area in the real space is located, the user can see the virtual UI and hand model also used for operation. The virtual UI fixing area is determined as described later from among the virtual UI fixing area candidates detected by the virtual UI fixing area candidate detection unit 209. The virtual object separation and placement control unit 210 will be described in detail together with the description of FIG. 6 described later.

[0027] FIG. 3 is a diagram for explaining an overview of the first embodiment. FIG. 3(A) is a diagram showing a state in which a user wearing an HMD 101 performs a virtual UI operation by hand tracking while sitting on a chair and facing a table. FIG. 3(B) is a diagram for explaining the arrangement of a display-only model (virtual UI 301 and hand model 302) and an operation-compatible model (virtual UI 303 and hand model 304). In the following, when the virtual UI 301 (second virtual object) and the hand model 302 (second hand model) are collectively referred to without distinction, they are referred to as a "display-only model". When the virtual UI 303 (first virtual object) and the hand model 304 (first hand model) are collectively referred to without distinction, they are referred to as an "operation-compatible model". Such notation is also the same in the second embodiment described later.

[0028] In the case of FIG. 3, in the information processing device 102, among the virtual UI fixing area candidates detected by the virtual UI fixing area candidate detection unit 209, the virtual UI fixing area candidate corresponding to the plane of the table shown in FIG. 3(A) or the left diagram of FIG. 3(B) is determined as the virtual UI fixing area 305. In addition, in the information processing device 102, the virtual object separation and placement control unit 210 generates a display-only model and an operation-compatible model. Furthermore, the display-only model is placed in an area directly in front of the user (for example, away from both eyes of the user) in the virtual space by the virtual object separation and placement control unit 210. As a result, as shown in the right diagram of FIG. 3(A) or FIG. 3(B), the display-only model (the virtual UI 301 and the hand model 302) is displayed in the display angle of view 306 of the HMD 101 when the user faces forward.

[0029] On the other hand, the operation-combined model is arranged in a region corresponding to the virtual UI fixing region 305 (i.e., the plane of the table in the real space) in the virtual space by the virtual object separation and arrangement control unit 210. As a result, as shown in the left diagram of FIG. 3(B), the operation-combined model (virtual UI 303 and hand model 304) is in a state of being located in the virtual UI fixing region 305 (i.e., the plane of the table of the real object). Note that the FOV 307 shown in FIG. 3(A) indicates the FOV (Field Of View) of the camera that performs hand tracking (i.e., the two cameras of the imaging unit 202 of the HMD 101). In the virtual UI fixing region 305 and its vicinity, the user performs virtual UI operation by hand tracking. Therefore, the virtual UI fixing region 305 and its vicinity must be included in the FOV 307.

[0030] When a user performs a virtual UI operation by hand tracking in or near virtual UI fixed area 305, control changes such as hand model position changes accompanying detection of a hand or the like are calculated and processed using the operation combined model (virtual UI 303 and hand model 304). Control changes such as hand model position changes calculated and processed in this way are reflected not only in the operation combined model (virtual UI 303 and hand model 304) but also in the display-only model (virtual UI 301 and hand model 302). Therefore, the display-only model (virtual UI 301 and hand model 302) is linked to the operation combined model (virtual UI 303 and hand model 304).

[0031] The operation shared model (virtual UI 303 and hand model 304) may be arranged in a parallel area with a predetermined interval from an area corresponding to the virtual UI fixing area 305 (i.e., the plane of the table of the real object) in the virtual space. This allows the operation shared model (virtual UI 303 and hand model 304) to be in a state of floating slightly parallel to the virtual UI fixing area 305 (i.e., the plane of the table of the real object) toward the HMD 101, for example, thereby improving the contact determination accuracy and operability. The predetermined interval is defined in advance, but may be adjusted by an operation on the HMD 101 or the information processing device 102.

[0032] Reference numerals 308 and 309 denote CG interruption buttons (hereinafter referred to as "virtual interruption buttons"). The virtual interruption button 308 is disposed near the display-only model in the virtual space. The virtual interruption button 309 is disposed near the operation-combined model in the virtual space. The user can issue an instruction to interrupt the execution of a process relating to virtual object separation and arrangement, which will be described later, at any time by pressing the virtual interruption buttons 308 and 309. In other words, when the virtual interruption buttons 308 and 309 are pressed in the mixed reality space of the HMD 101, the user performs a virtual UI operation by hand tracking similar to that of the conventional technology.

[0033] 3, a keyboard is shown in CG as the virtual UI 301, 303, but the same applies when a UI other than a keyboard is shown in CG. For example, the virtual UI 301, 303 may be something other than a keyboard, such as a button that is pressed, a toggle that is slid, or a pull-down menu that is selected by pulling down. This also applies to the virtual UI of the second embodiment described later.

[0034] FIG. 4 is a diagram for explaining the case where the user turns the line of sight from the front to the hands. FIG. 4(A) is a diagram showing a state where the user wearing the HMD 101 is looking ahead. FIG. 4(B) is a diagram showing a state where the user wearing the HMD 101 is looking at the hands. As shown in FIG. 4(A), when the user performing the virtual UI operation by hand tracking turns the line of sight forward, the display dedicated models (virtual UI 301 and hand model 302) are displayed in the display angle of view 306 of the HMD 101. As shown in FIG. 4(B), when the user performing the virtual UI operation by hand tracking turns the line of sight to the hands, the operation-dedicated models (virtual UI 303 and hand model 304) are displayed in the display angle of view 306 of the HMD 101. In this way, the user can visually recognize at which position on the plane of the real object the contact determination for the virtual UI operation is being performed in the composite reality space by looking at the operation-dedicated models (virtual UI 303 and hand model 304). Note that the operation-dedicated model may be displayed in a color different from that of the display-dedicated model, for example, to facilitate the distinction from the display-dedicated model.

[0035] FIG. 5 is a flowchart showing the control method of the first embodiment. The control method (control method of the information processing apparatus) shown in the flowchart of FIG. 5 is realized by the control unit 206 (computer) executing a program in the information processing apparatus 102. At that time, the control unit 206 exchanges data with the HMD 101 as necessary. In step S501, the control unit 206 acquires information regarding the position and posture of each joint of the user's hand wearing the HDM 101 via the hand joint point detection unit 205 of the HMD 101. In step S502, the control unit 206 acquires virtual UI fixed area candidates by the virtual UI fixed area candidate detection unit 209 (detection step). At that time, the virtual UI fixed area candidate detection unit 209 detects virtual UI fixed area candidates from the image acquired by the imaging unit 202 of the HMD 101. As a result, one or more virtual UI fixed area candidates are detected.

[0036] In step S503, the control unit 206 performs a virtual object separation and placement control process. The virtual object separation and placement control process in step S503 will be described in detail with reference to FIG. 6. In step S504, the control unit 206 performs a contact determination between the virtual UI and the hand model by the contact determination unit 208 (operation determination step). At that time, if a process related to virtual object separation and placement described later is performed in step S503, the contact determination unit 208 performs a contact determination between the virtual UI and the hand model using the operation combined model. Note that, as described above, a control change such as a hand model position change accompanying this contact determination is also reflected in the display dedicated model. On the other hand, if a process related to virtual object separation and placement described later is not performed in step S503, the contact determination unit 208 performs a contact determination between the virtual UI and the hand model similar to the conventional technology. Thereafter, the control method shown in the flowchart of FIG. 5 ends. Note that the processes in steps S501 and S502 are in no particular order, and therefore any process may be performed first.

[0037] 6 is a flowchart showing the virtual object separation and placement control process in step S503. As shown in FIG. 6, when the virtual object separation and placement control process is performed, each process from step S601 onward is performed. In step S601, the control unit 206 (first determination means) determines whether or not there is one or more virtual UI fixing area candidates detected in step S502 that are located within the reach of the user. In this step S601, the control unit 206 determines, for example, a virtual UI fixing area candidate whose distance from the HMD 101 is less than the threshold value as a virtual UI fixing area candidate located within the reach of the user. In this way, the determination in this step S601 is performed based on the distance from the user to the virtual UI fixing area candidate.

[0038] When the control unit 206 determines that there is one or more virtual UI fixed area candidates detected in step S502 that are located within the reach of the user's hand, the process proceeds to step S602. On the other hand, when the control unit 206 determines that there is not a single virtual UI fixed area candidate detected in step S502 that is located within the reach of the user's hand, the process proceeds to step S608 described later. That is, when there is not a single virtual UI fixed area candidate located within the reach of the user's hand, the processes related to the virtual object separation and arrangement shown in steps S606 and S607 described later will not be performed.

[0039] In step S602, the control unit 206 (second determination means) determines whether there is one or more virtual UI fixed area candidates determined to be located within the reach of the user's hand in step S601 that are located near the lower front of the user. In this step S602, the control unit 206 discriminates, for example, a virtual UI fixed area candidate included in a predetermined angular region based on the front face of the HMD 101 worn by the user as a virtual UI fixed area candidate located near the lower front of the user. In this way, the determination in this step S602 is made based on the direction from the user to the virtual UI fixed area candidate. Also, a virtual UI fixed area candidate located near the user's hand is discriminated as a virtual UI fixed area candidate located near the lower front of the user.

[0040] When the control unit 206 determines that there is one or more virtual UI fixed area candidates determined to be located within the reach of the user's hand in step S601 that are located near the lower front of the user, the process proceeds to step S603. On the other hand, when the control unit 206 determines that there is not a single virtual UI fixed area candidate determined to be located within the reach of the user's hand in step S601 that is located near the lower front of the user, the process proceeds to step S608 described later. That is, when there is not a single virtual UI fixed area candidate located near the lower front of the user, the processes related to the virtual object separation and arrangement shown in steps S606 and S607 described later will not be performed.

[0041] In step S603, the control unit 206 (third determination means) determines whether there is one or more virtual UI fixed area candidates determined to be located near the lower front side of the user in step S602 and having a specific size. The specific size refers to the size in which the virtual UI can be operated by hand tracking. The specific size takes into account not only the area but also the shape in which the virtual UI that is the target of the operation by hand tracking can be arranged. Therefore, in this step S603, the control unit 206 approximates the virtual UI fixed area candidate to, for example, a rectangle or an ellipse.

[0042] Furthermore, if the control unit 206 approximates the virtual UI fixed area candidate to a rectangle, the virtual UI fixed area candidates whose long side length, short side length, and area are each equal to or greater than their respective threshold values are determined as virtual UI fixed area candidates having a specific size. Also, if the control unit 206 approximates the virtual UI fixed area candidate to an ellipse, the virtual UI fixed area candidates whose major axis length, minor axis length, and area are each equal to or greater than their respective threshold values are determined as virtual UI fixed area candidates having a specific size. In this way, the determination in this step S603 is made based on the size of the virtual UI fixed area candidate.

[0043] When the control unit 206 determines that there is one or more virtual UI fixed area candidates determined to be located near the lower front side of the user in step S602 and having a specific size, the process proceeds to step S604. On the other hand, when the control unit 206 determines that there is not even one virtual UI fixed area candidate determined to be located near the lower front side of the user in step S602 and having a specific size, the process proceeds to step S608 described later. That is, when there is not even one virtual UI fixed area candidate having a specific size, the processes related to the virtual object separation and arrangement shown in steps S606 and S607 described later are not performed.

[0044] In step S604, the control unit 206 (fourth determination means) determines whether an interruption operation by the user has been performed. In this determination, the control unit 206 treats the pressing operations of the virtual interruption buttons 308 and 309 performed by hand tracking or the like as an instruction from the user, that is, an interruption operation by the user. In this way, the determination in this step S604 is made based on an instruction from the user. When the control unit 206 determines that an interruption operation by the user has been performed, the process proceeds to step S608 described later. That is, when an interruption operation by the user has been performed, the processes regarding the virtual object separation and arrangement shown in steps S606 and S607 described later are not performed. On the other hand, when the control unit 206 determines that an interruption operation by the user has not been performed, the process proceeds to step S605.

[0045] In step S605, the control unit 206 (determination means) determines the virtual UI fixed area from among the virtual UI fixed area candidates determined to have a specific size in step S603. In this step S605, the control unit 206 determines, for example, the virtual UI fixed area candidate that is most similar to a predetermined area among the virtual UI fixed area candidates determined to have a specific size in step S603 as the virtual UI fixed area. The predetermined area refers to an area having the optimal position and size for the user to perform virtual UI operations by hand tracking. In this way, among the plurality of virtual UI fixed area candidates, the one that is most similar to the area having the optimal position and size for the user to perform virtual UI operations by hand tracking is determined as the virtual UI fixed area. Note that when there is only one virtual UI fixed area candidate determined to have a specific size in step S603, the control unit 206 determines that virtual UI fixed area candidate as the virtual UI fixed area.

[0046] In step S606, the control unit 206 generates a combined operation model by the virtual object separation and placement control unit 210. At this time, the virtual object separation and placement control unit 210 generates a hand model included in the combined operation model from information on the position and posture of each joint point of the user's hand acquired in step S501. Furthermore, the virtual object separation and placement control unit 210 places the generated combined operation model in a region (first region) corresponding to the virtual UI fixing region determined in step S605 in the virtual space (first placement process). As a result, the contact determination between the virtual UI and the hand model in step S504 is performed using the combined operation model generated in this step S606. Note that, as described above, the virtual object separation and placement control unit 210 may place the combined operation model in a parallel region at a predetermined interval from the region corresponding to the virtual UI fixing region determined in step S605 in the virtual space.

[0047] In step S607, the control unit 206 generates a display-only model by the virtual object separation and placement control unit 210. At this time, the virtual object separation and placement control unit 210 generates a hand model included in the display-only model from information on the position and posture of each joint point of the user's hand acquired in step S501. Furthermore, the virtual object separation and placement control unit 210 places the generated display-only model in a region (second region) directly in front of the user (for example, away from both eyes of the user) in the virtual space (second placement step). In this way, the display-only model is placed in a region different from the region in which the operation-compatible model is placed. Thereafter, the virtual object separation and placement control process shown in the flowchart of FIG. 6 ends, and the process proceeds to step S504.

[0048] In step S608, the control unit 206 generates an operation combined model and a display only model in the same manner as in steps S606 and S607. Furthermore, the control unit 206 arranges the generated operation combined model and display only model in the virtual space, for example, in an area directly in front of the user (for example, away from both eyes of the user) so as to overlap each other. Therefore, the contact determination between the virtual UI and the hand model in step S504 is performed in the same manner as in the conventional technology. Note that the operation combined model and the display only model may be arranged in the virtual space so as to overlap each other in an area where the user can perform virtual UI operation by hand tracking. Therefore, in this step S608, the control unit 206 may arrange the operation combined model and the display only model in the virtual space so as to overlap each other in an area corresponding to the virtual UI fixed area determined in step S605. Also, in this step S608, the control unit 206 may generate and arrange only the operation combined model.

[0049] 6 ends, and the process proceeds to step S504. Note that the processes in steps S601 to S604 are in no particular order, and therefore any process may be performed first. Furthermore, the processes in steps S606 and S607 are in no particular order, and therefore any process may be performed first.

[0050] As described above, in the first embodiment, when a user operates a virtual UI by hand tracking, a dual-use model and a display-only model compatible with hand tracking are used. In this case, the dual-use model is positioned on the plane of a real object detected near the user's hand or on a parallel position with a predetermined distance from the plane, and a determination is made regarding the virtual UI operation by hand tracking based on the dual-use model. As a result, when performing a virtual UI operation by hand tracking, the user can feel a typing sensation by touching the plane of the real object with his / her fingertips, and can prevent arm fatigue by placing his / her hand on the plane of the real object.

[0051] On one hand, the display-only model is placed in front of the user. As a result, the user can type on the operation-cum-display model placed near their hands or the like while looking at the display-only model placed directly in front without looking down, making it difficult to cause neck fatigue. In this way, the information processing apparatus 102 of the first embodiment can improve the usability when the user operates the user interface represented by the virtual object in the HMD 101 by hand tracking. Note that this also applies to the second embodiment described later.

[0052] <Second Embodiment> Hereinafter, the second embodiment will be described with reference to FIGS. 7 and 8. Here, the description will focus on the parts different from the first embodiment. In the second embodiment, the sizes and positions of the display-only model and the operation-cum-display model are changed according to the size of the determined virtual UI fixed area and the size of the CG content different from the virtual UI. FIG. 7 is a diagram for explaining the outline of the first example of the second embodiment. In FIG. 7, the same use case as the case shown in FIG. 3 described in the first embodiment is assumed. In FIG. 7, the cases shown in FIGS. 7(A-1) and 7(A-2) and the cases shown in FIGS. 7(B-1) and 7(B-2) are described side by side.

[0053] FIGS. 7(A-1) and 7(B-1) are diagrams showing the display angle of view 701 of the HMD 101 when the user faces forward. In the display angle of view 701 of the HMD 101 when the user faces forward, the display-only model is displayed. The display-only model is composed of a virtual UI 702 (second virtual object) and a hand model 703 (second hand model). FIGS. 7(A-2) and 7(B-2) are diagrams showing the operation-cum-display model located in the virtual UI fixed area 706 on the tables 704 and 705. The operation-cum-display model is composed of a virtual UI 707 (first virtual object) and a hand model 708 (first hand model).

[0054] As shown in Figures 7(A-2) and 7(B-2), the operation compatible model is generated at a size that fits in the virtual UI fixed area 706. In the cases shown in Figures 7(B-1) and 7(B-2), the size of the virtual UI fixed area 706 is smaller than in the cases shown in Figures 7(A-1) and 7(A-2). Therefore, in the cases shown in Figures 7(B-1) and 7(B-2), the operation compatible model is generated at a smaller size than in the cases shown in Figures 7(A-1) and 7(A-2), and accordingly the display-only model is also generated at a smaller size.

[0055] However, even if the size of the virtual UI fixed area 706 is relatively small as shown in Fig. 7(B-2), in order to improve the visibility of the virtual UI operation by hand tracking, the display-only model may be generated in a relatively large size as shown in Fig. 7(A-1). The size of the operation-compatible model and the display-only model in the cases shown in Fig. 7(B-1) and Fig. 7(B-2) is changed in steps S607 and S608 of Fig. 6.

[0056] FIG. 8 is a diagram for explaining an outline of a second example of the second embodiment. In FIG. 8, a use case similar to that shown in FIG. 3 described in the first embodiment is assumed. FIG. 8(A) is a diagram showing a display angle of view 801 of the HMD 101 when the user faces forward. In addition to CG content 802, a display-only model is displayed in the display angle of view 801 of the HMD 101 when the user faces forward. The display-only model is composed of a virtual UI 803 (second virtual object) and a hand model 804 (second hand model). FIG. 8(B) is a diagram showing an operation-compatible model located in a virtual UI fixed area 806 on a table 805. The operation-compatible model is composed of a virtual UI 807 (first virtual object) and a hand model 808 (first hand model). As shown in FIGS. 8(A) and 8(B), the CG content 802 (another virtual object) is a virtual object different from the virtual UIs 803 and 807, and includes a CG object, a video, and the like.

[0057] The virtual UI fixed area 806 is large enough for the user to operate the virtual UI by hand tracking. On the other hand, when the CG content 802 is placed directly in front of the user in the virtual space, the display-only model is generated in a smaller size than the operation-compatible model and is placed in a position that does not overlap with the CG content 802. As a result, the CG content 802 and the display-only model fit within the display angle of view 801 of the HMD 101 when the user faces forward. The size and position of the display-only model in the case shown in FIG. 8 are changed in step S607 of FIG. 6. Note that, as long as the CG content 802 and the display-only model fit within the display angle of view 801 of the HMD 101, only the size of the display-only model may be changed, or only the position of the display-only model may be changed.

[0058] As described above, in the first example of the second embodiment, as shown in FIG. 7, the size of the display-only model and the operation-combined model are changed depending on the size of the virtual UI fixed region 706. Even in this way, the information processing device 102 in the first example of the second embodiment can improve the usability when the user operates the user interface displayed by a virtual object on the HMD 101 by hand tracking. Also, in the second example of the second embodiment, as shown in FIG. 8, the size and position of the display-only model are changed depending on the size of the CG content 802. This allows the user to operate the virtual UI by hand tracking while looking at the CG content 802. Even in this way, the information processing device 102 in the second example of the second embodiment can improve the usability when the user operates the user interface displayed by a virtual object on the HMD 101 by hand tracking.

[0059] <Example of change> As described above, the preferred embodiments of the present invention have been explained. However, the present invention is not limited to the above-described embodiments, and various modifications and changes are possible within the scope of the gist thereof. For example, although the HMD 101 is video-transmissive, it may be optically transmissive. When the HMD 101 is optically transmissive, smart glasses may be used instead of the HMD 101. Also, although the virtual UIs 301, 303, 702, 707, 803, 807 are keyboards, only the home position keys of the keyboard may be displayed. Examples of the home position keys of the keyboard include the [F] key and the [J] key. Further, the [A] key, [S] key, [D] key, [K] key, [L] key, and [+] key may be included in the home position keys of the keyboard.

[0060] Also, the determination target in step S602 is a virtual UI fixed area candidate located near the lower front of the user. However, for example, it may be a virtual UI fixed area candidate located near the lower right or lower left of the user. Also, the operation combined model may be arranged to move little by little from the area directly in front of the user in the virtual space to the area corresponding to the virtual UI fixed area in the virtual space. In this way, the user wearing the HMD 101 can easily find the operation combined model by facing forward, and further, by continuing to look at the operation combined model, the user can visually recognize the place where the operation combined model is placed.

[0061] Also, during the virtual UI operation by hand tracking, all or part of the virtual UI 301 may be erased. The erasure of all or part of the virtual UI 301 is performed by the control unit 206 (erasure means) in step S607 of FIG. 6. Similarly, during the virtual UI operation by hand tracking, all or part of the virtual UI 303 may be erased. The erasure of all or part of the virtual UI 303 is performed by the control unit 206 (erasure means) in step S606 of FIG. 6. These points are the same for the virtual UIs 702, 707, 803, 807.

[0062] The present invention can also be realized by supplying a program that implements one or more functions of each of the above embodiments to a system or apparatus via a network or a storage medium, and causing one or more processors of a computer of the system or apparatus to read and execute the program. Further, the present invention can also be realized by a circuit (for example, ASIC) that implements one or more functions.

[0063] Note that the disclosure of each embodiment includes the following configurations, methods, and programs. (Configuration 1) An information processing apparatus that enables a user to operate a user interface represented by a virtual object by hand tracking in a space where the real space and the virtual space are fused, comprising: detection means for detecting one or more regions from the real space in which the virtual object representing the user interface can be fixed; first arrangement means for arranging an operation and display combined model constituted by a first hand model corresponding to the hand tracking and a first virtual object representing the user interface in a first region of the virtual space corresponding to any one of the one or more regions detected by the detection means; second arrangement means for arranging a display dedicated model constituted by a second hand model corresponding to the hand tracking and a second virtual object representing the user interface in a second region of the virtual space different from the first region; operation determination means for making a determination regarding the operation of the user interface based on the operation and display combined model, characterized in that the information processing apparatus comprises the same. (Configuration 2) The information processing apparatus according to Configuration 1, further comprising first determination means for determining whether to perform the arrangement by the first arrangement means and the arrangement by the second arrangement means based on the distance from the user to the one or more regions detected by the detection means. (Configuration 3) The information processing apparatus according to Configuration 1 or 2, further comprising second determination means for determining whether to perform the arrangement by the first arrangement means and the arrangement by the second arrangement means based on the direction from the user to the one or more regions detected by the detection means. (Configuration 4) An information processing device according to any one of configurations 1 to 3, further comprising a third determination means for determining whether to perform placement by the first placement means and placement by the second placement means based on the size of one or more areas detected by the detection means. (Configuration 5) An information processing device according to any one of configurations 1 to 4, further comprising a fourth determination means for determining whether to perform placement by the first placement means and placement by the second placement means based on instructions from the user. (Configuration 6) An information processing device according to any one of configurations 1 to 5, further comprising a determination means for determining, among one or more regions detected by the detection means, a region that is most similar to a specified region as the region corresponding to the first region. (Configuration 7) The information processing device according to any one of configurations 1 to 6, wherein the first arrangement means arranges the operation-compatible model at a predetermined distance from the first area. (Configuration 8) The information processing device according to any one of configurations 1 to 7, wherein the first arrangement means changes a size of the operation compatible model based on a size of the first area. (Configuration 9) In the information processing device according to configuration 8, the second arranging means changes a size of the display-only model in accordance with a change in size of the operation-compatible model by the first arranging means. (Configuration 10) An information processing device described in any one of configurations 1 to 7, characterized in that the second placement means changes at least one of the size and position of the display-only model depending on the size of a virtual object other than the operation compatible model and the display-only model placed in the vicinity of the second area. (Configuration 11) The information processing device according to any one of configurations 1 to 10, wherein the first area is located near the user's hand. (Configuration 12) The information processing device according to any one of configurations 1 to 11, wherein the second area is located directly in front of the user. (Configuration 13) The information processing apparatus according to any one of Configurations 1 to 12, wherein the user interface is one on which a pressing operation is performed. (Configuration 14) The information processing apparatus according to Configuration 13, wherein the user interface is a keyboard. (Configuration 15) The information processing apparatus according to Configuration 14, wherein the first virtual object and the second virtual object are keys at the home position of the keyboard. (Configuration 16) The information processing apparatus according to Configuration 14 or 15, further comprising erasing means for erasing all or part of the first virtual object or the second virtual object during an operation of the user interface. (Configuration 17) The information processing apparatus according to any one of Configurations 1 to 12, wherein the user interface is one on which a slide operation is performed. (Configuration 18) The information processing apparatus according to any one of Configurations 1 to 12, wherein the user interface is one on which a selection operation is performed. (Method 1) A control method for an information processing apparatus that enables a user to operate a user interface represented by a virtual object by hand tracking in a space where the real space and the virtual space are fused, a detection step of detecting one or more regions in the real space where the virtual object representing the user interface can be fixed; a first arrangement step of arranging an operation-dedicated model composed of a first hand model corresponding to the hand tracking and a first virtual object representing the user interface in a first region of the virtual space corresponding to any one of the one or more regions detected in the detection step; a second arrangement step of arranging a display-dedicated model composed of a second hand model corresponding to the hand tracking and a second virtual object representing the user interface in a second region of the virtual space different from the first region; and an operation determination step of determining an operation related to the user interface based on the operation-dedicated model. (Program 1) A program for causing a computer to execute each means of the information processing apparatus according to any one of Configurations 1 to 18.

Explanation of Signs

[0064] 102 Information processing apparatus 208 Contact determination unit (operation determination means) 209 Virtual UI fixed area candidate detection unit (detection means) 210 Virtual object separation and arrangement control unit (first arrangement means) (second arrangement means) 301, 702, 803 Virtual UI (second virtual object) 302, 703, 804 Hand model (second hand model) 303, 707, 807 Virtual UI (first virtual object) 304, 708, 808 Hand model (first hand model)

Claims

1. An information processing apparatus that enables a user to operate a user interface represented by a virtual object by hand tracking in a space where the real space and the virtual space are fused, detection means for detecting one or more regions from the real space where the virtual object representing the user interface can be fixed; first arrangement means for arranging an operation and display combined model constituted by a first hand model corresponding to the hand tracking and a first virtual object representing the user interface in a first region of the virtual space corresponding to any one of the one or more regions detected by the detection means; second arrangement means for arranging a display dedicated model constituted by a second hand model corresponding to the hand tracking and a second virtual object representing the user interface in a second region of the virtual space different from the first region; operation determination means for making a determination regarding the operation of the user interface based on the operation and display combined model, the information processing apparatus being characterized by comprising the operation determination means.

2. The information processing apparatus according to claim 1, further comprising first determination means for determining whether to perform the arrangement by the first arrangement means and the arrangement by the second arrangement means based on the distance from the user to the one or more regions detected by the detection means.

3. The information processing apparatus according to claim 1 or 2, further comprising second determination means for determining whether to perform the arrangement by the first arrangement means and the arrangement by the second arrangement means based on the direction from the user to the one or more regions detected by the detection means.

4. The information processing apparatus according to claim 1, further comprising third determination means for determining whether to perform the arrangement by the first arrangement means and the arrangement by the second arrangement means based on the size of the one or more regions detected by the detection means.

5. The information processing apparatus according to claim 1, further comprising fourth determination means for determining whether to perform the arrangement by the first arrangement means and the arrangement by the second arrangement means based on an instruction from the user.

6. The information processing apparatus according to claim 1, further comprising determination means for determining, among the one or more regions detected by the detection means, the region most similar to a predetermined region as the region corresponding to the first region.

7. 2 . The information processing apparatus according to claim 1 , wherein the first arrangement means arranges the dual-operation model at a predetermined distance from the first area.

8. 2 . The information processing apparatus according to claim 1 , wherein the first arrangement means changes a size of the operation-compatible model based on a size of the first area.

9. 9. The information processing apparatus according to claim 8, wherein the second arrangement means changes a size of the display-only model in accordance with a change in size of the operation-compatible model by the first arrangement means.

10. 2. The information processing device according to claim 1, wherein the second placement means changes at least one of a size and a position of the display-only model in accordance with a size of a virtual object other than the operation compatible model and the display-only model that is placed near the second area.

11. The information processing device according to claim 1 , wherein the first area is located near the user's hand.

12. The information processing device according to claim 1 , wherein the second area is located directly in front of the user.

13. The information processing apparatus according to claim 1 , wherein the user interface is one that is operated by pressing a key.

14. The information processing apparatus according to claim 13, wherein the user interface is a keyboard.

15. The information processing apparatus according to claim 14 , wherein the first virtual object and the second virtual object are keys in a home position of the keyboard.

16. 16. The information processing apparatus according to claim 14, further comprising an erasing means for erasing all or a part of the first virtual object or the second virtual object during an operation of the user interface.

17. The information processing apparatus according to claim 1 , wherein the user interface is one in which a slide operation is performed.

18. The information processing apparatus according to claim 1 , wherein the user interface is an interface through which a selection operation is performed.

19. A method for controlling an information processing device that enables a user to operate a user interface represented by a virtual object by hand tracking in a space where a real space and a virtual space are integrated, comprising: a detection step of detecting one or more areas in the real space to which a virtual object representing the user interface can be fixed; An operation-combined model composed of a first hand model corresponding to the hand tracking and a first virtual object indicating the user interface is arranged in a first region of the virtual space corresponding to any one of the one or more regions detected in the detection step. A first arrangement step; A display-dedicated model composed of a second hand model corresponding to the hand tracking and a second virtual object indicating the user interface is arranged in a second region of the virtual space different from the first region. A second arrangement step; An operation determination step of determining an operation related to the user interface based on the operation-combined model. A control method for an information processing apparatus, comprising:

20. A program for causing a computer to execute each means of the information processing apparatus according to claim 1.

Citation Information

Patent Citations

  • Touchscreen hover detection in augmented and / or virtual reality environment

    JP2020113298A