Information processing apparatus, display system, information processing method, and program
Patent Information
- Application Number
- JP2025120171
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-09-19
AI Technical Summary
Existing user input methods in three-dimensional virtual spaces, such as using a mouse and keyboard, are not intuitive and result in low usability due to the need to tilt the head and operate with the other hand, making interactions difficult.
An information processing device that generates user objects in a three-dimensional virtual space based on the position and tilt of a head-mounted display and a controller, allowing for intuitive interactions through assistant and laser objects that can be easily touched or selected using hand movements.
Improves usability in three-dimensional virtual spaces by enabling more natural and intuitive user interactions, such as writing on a whiteboard, through the use of assistant and laser objects that are positioned based on the user's shoulder and hand movements.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device, a display system, and a method for generating an image, and a program for causing a computer to execute a process for generating an image. [Background technology]
[0002] VR (Virtual Reality) technology allows users to experience content presented in a three-dimensional virtual space, and uses a head-mounted display (HMD) worn on a person's head as a display device to display images of the three-dimensional virtual space.
[0003] Even in a 3D virtual space, there are cases where user input, such as writing on a whiteboard, is required. However, there is no standard input method for user input in a 3D virtual space, such as using a mouse and keyboard as in a personal computer (PC).
[0004] A known method for user input in a three-dimensional virtual space is to place an avatar object corresponding to the user in the user's line of sight and display a user interface on the avatar object to accept operations from the user (see, for example, Patent Document 1). Summary of the Invention [Problem to be solved by the invention]
[0005] However, in the above-mentioned conventional method, the user interface is displayed on one arm of the avatar object, so the user needs to tilt their head to view that arm and operate it with the other hand, which makes operation difficult and does not have high usability.
[0006] The present invention has been made in consideration of the above problems, and aims to provide an information processing device, a display system, an information processing method, and a program that can improve usability in a three-dimensional virtual space. [Means for solving the problem]
[0007] In order to solve the above-described problems, one embodiment of the present invention provides an information processing device that generates an image based on virtual space data representing a three-dimensional virtual space, the information processing device comprising: an acquisition means for acquiring position information and tilt information of the display means and position information of the operation means or hand from a detection means for detecting the position of the display means worn by the user, the tilt of the display means relative to a reference direction, and the position of the operation means operated by the user or the user's hand; a first generation means for generating a user object in a three-dimensional virtual space based on the acquired position information of the display means and the position information of the operation means or the hand; and second generating means for generating an image of the tilt direction of the display means in the three-dimensional virtual space based on the acquired tilt information of the display means and the virtual space data. [Effects of the Invention]
[0008] According to the present invention, it is possible to improve usability in a three-dimensional virtual space. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a diagram showing a first configuration example of a display system. [Figure 2] A diagram showing the first hardware configuration of an HMD, a controller, and a PC. [Figure 3] FIG. 1 is a block diagram showing an example of the functional configuration of a PC. [Figure 4] 2 is a sequence diagram showing the flow of processing executed by the display system shown in FIG. 1. [Figure 5] 10A and 10B are diagrams illustrating a first method for estimating the position of a user's shoulder. [Figure 6] 6 is a flowchart showing the flow of processing for estimating the position of the user's shoulders using the first method shown in FIG. 5; [Figure 7] 10A and 10B are diagrams illustrating a second method for estimating the position of the user's shoulders. [Figure 8] 8 is a flowchart showing the flow of processing for estimating the position of the user's shoulders using the second method shown in FIG. 7. [Figure 9] 10 is a flowchart showing the flow of a process for generating an assistant object as a first example of a user object for assisting a user in inputting data into a three-dimensional virtual space and displaying an image thereof. [Figure 10] FIG. 10 is a diagram showing a display example in which an assistant object is generated in a three-dimensional virtual space and an image is displayed. [Figure 11] A diagram showing the assistant object being touched and its function being invoked. [Figure 12] 10 is a flowchart showing the flow of a process for generating a laser object as a second example of a user object for supporting user input in a three-dimensional virtual space and displaying an image thereof. [Figure 13] FIG. 10 is a diagram showing an example of selecting an object using a laser object. [Figure 14] FIG. 10 is a diagram showing the relationship between the length of a laser object and the distance between the position of the user's shoulder and the position of the controller. [Figure 15] FIG. 10 is a diagram illustrating an example of moving a selected object using a laser object. [Figure 16] 10 is a flowchart showing the flow of a process for moving a selected object with a laser object. [Figure 17] 10 is a flowchart showing the flow of voice input as an example of user input. [Figure 18] FIG. 10 is a diagram showing a display example when voice input is being performed. [Figure 19] FIG. 10 is a diagram showing a second configuration example of the display system. [Figure 20] A diagram showing a second hardware configuration of an HMD and a controller. DETAILED DESCRIPTION OF THE INVENTION
[0010] Fig. 1 is a diagram showing a first configuration example of a display system. The display system includes an HMD 10 that functions as a display means worn on a user's head, a controller 11 that functions as an operation means that the user holds in their hand or wears on their hand and operates, and a PC 12 that functions as an information processing device. The display system also includes a position detection sensor 13 that functions as a detection means that detects the position of the HMD 10 and the tilt of the HMD 10 relative to a reference direction, and the position of the controller 11 and the tilt of the controller 11 relative to the reference direction. In the example shown in Fig. 1, the display system includes a server 14 that manages position information, operation information, etc. of multiple users.
[0011] The PC 12 and the server 14 are communicably connected via a network 15. The HMD 10 and the position detection sensor 13 are connected to the PC 12 via a cable or the like. The controller 11 is wirelessly connected to the PC 12 via Bluetooth (registered trademark), WiFi (registered trademark), or the like. The HMD 10 and the position detection sensor 13 may also be wirelessly connected via WiFi (registered trademark) or the like.
[0012] The HMD 10 has a display for displaying images to the user, and displays images on the display that correspond to the position of the HMD 10 and its inclination relative to a reference direction. The images are presented as two images, one for each of the user's eyes, to make the image appear three-dimensional by using the parallax between the user's left and right eyes. For this reason, the HMD 10 is equipped with two displays that display images corresponding to the left and right eyes, respectively. The reference direction is, for example, any direction parallel to the floor. The HMD 10 has a light source, such as an infrared LED (Light Emitting Diode), and emits infrared light.
[0013] The controller 11 is an operating means that the user holds in their hand or wears in their hand, and has buttons, a wheel, a touch sensor, etc., receives input from the user, and transmits the received information to the PC 12. The controller 11 also has a light source such as an infrared LED, and emits infrared rays.
[0014] The position detection sensor 13 is placed at any position in front of the user, detects the position and tilt of the HMD 10 and the controller 11 from infrared rays emitted from the HMD 10 and the controller 11, and outputs position information and tilt information. The position detection sensor 13 is, for example, an infrared camera, and can detect the position and tilt of the HMD 10 and the controller 11 based on captured images. Note that the HMD 10 and the controller 11 are provided with multiple light sources in order to detect the position and tilt of the HMD 10 and the controller 11 with high accuracy. The position detection sensor 13 is composed of one or more sensors, and when multiple sensors are used, they can be provided on the side, rear, etc.
[0015] The PC 12 generates a user object for supporting user input in the three-dimensional virtual space displayed on the display of the HMD 10, based on the position information and tilt information of the HMD 10 output from the position detection sensor 13, the position information of the controller 11, and, if necessary, the tilt information of the controller 11. Then, based on the position information and tilt information of the HMD 10 and the three-dimensional virtual space data, the PC 12 generates images corresponding to the left and right eyes, which are images in the user's field of view in the three-dimensional virtual space (more precisely, the tilt direction of the HMD 10), and executes processing to display the images on the display of the HMD 10.
[0016] The PC 12 communicates with the server 14 via the network 15, obtains location information of other users in the same three-dimensional virtual space, and can execute a process to display avatar objects representing the other users on the display of the HMD 10.
[0017] The display system can be used, for example, to gather the avatar objects of each user in a virtual conference room as a three-dimensional virtual space and hold a conference using a whiteboard, etc. The display system can be used for holding interactive conferences because the participants can actively participate in the conference using the whiteboard, etc.
[0018] In a conference using the display system, a user can operate the controller 11, touch a user object in a displayed image, or the like to invoke the pen input function, pick up the displayed pen, move the pen, and input text on the whiteboard. Note that this is just one form of use, and the use is not limited to this form.
[0019] 1, the HMD 10 and the controller 11 have a light source and the position detection sensor 13 is disposed at an arbitrary position, but the HMD 10 and the controller 11 may have the position detection sensor 13 and a light source or a marker that reflects infrared light may be disposed at an arbitrary position. When a marker is used, the HMD 10 and the controller 11 are provided with the light source and the position detection sensor 13, and infrared light emitted from the light source is reflected by the marker, and the reflected infrared light is detected by the position detection sensor 13, thereby making it possible to detect the position and tilt of the HMD 10 and the controller 11.
[0020] If there is any object between the position detection sensor 13 and the HMD 10 or controller 11, the infrared rays will be blocked, making it impossible to accurately detect the position or tilt. For this reason, when using the HMD 10 and controller 11 to perform operations or displays, it is desirable to do so in an open space.
[0021] In the example shown in Figure 1, a space is provided in which the user can wear an HMD 10, hold a controller 11 in their hand, and stretch or unstretch their arms, and a PC 12 and a position detection sensor 13 are placed outside this space.
[0022] 2 is a diagram showing an example of the hardware configuration of the HMD 10, the controller 11, and the PC 12. The HMD 10 includes an external I / F 20, a CPU 21, a display 22, a memory 23, an HDD 24, a light source 25, and a microphone 26. The CPU 21 controls the entire HMD 10 and executes processes such as light emission by the light source 25, communication with the outside, and display on the display 22. The external I / F 20 is an interface for communicating with the PC 12. The display 22 may be a liquid crystal display or an organic EL (Electro Luminescence) display.
[0023] The memory 23 provides a working area for the CPU 21. The HDD 24 stores image data to be displayed in the three-dimensional virtual space. The light source 25 is an infrared LED or the like, which emits infrared rays. The infrared rays can be emitted in a flashing manner in a predetermined pattern. The microphone 26 is a voice input device that allows the user to input information by voice.
[0024] The controller 11 includes an operation I / F 30, an external I / F 31, and a light source 32. The operation I / F 30 is a button, a wheel, a touch sensor, etc., and is arranged on the outer surface of the controller 11 to enable operation by the user and to accept input of operation information. The external I / F 31 is wirelessly connected to the PC 12 and transmits operation information accepted by the operation I / F 30 to the PC 12. The light source 32 emits infrared light that flashes in a predetermined pattern. The light source 32 can be distinguished from the HMD 10 by flashing in a pattern different from the light source of the HMD 10.
[0025] The PC 12 includes a CPU 40 , a ROM 41 , a RAM 42 , a HDD 43 , an external I / F 44 , an input / output I / F 45 , an input device 46 , and a display device 47 .
[0026] The CPU 40 controls the entire PC 12, generates the above-mentioned user objects, and executes processing to generate an image in the user's line of sight in the three-dimensional virtual space and display it on the display of the HMD 10. The ROM 41 stores a boot program for starting up the PC 12, firmware for controlling the HDD 43, the external I / F 44, etc. The RAM 42 provides a working area for the CPU 40.
[0027] The HDD 43 stores an OS (Operating System), programs for executing the above processes, image data, etc. The external I / F 44 is connected to the network 15 shown in Fig. 1 and communicates with the server 14 via the network 15. The external I / F 44 is also connected to the HMD 10 and the position detection sensor 13 via a cable or the like, and is wirelessly connected to the controller 11, and communicates with the HMD 10, the controller 11, and the position detection sensor 13.
[0028] The input device 46 is a mouse, keyboard, etc., and is used by the user to input information and accept operations. The display device 47 provides a display screen for the user and displays input information, processing results, etc. The input / output I / F 45 is an interface that controls the input of information from the input device 46 and the output of information to the display device 47.
[0029] 3 is a block diagram showing an example of the functional configuration of the PC 12. Here, since the PC 12 functions as an information processing device, the functional configuration of the PC 12 will be described. The CPU 40 executes a program stored in the HDD 43 to generate functional units for realizing each function, and the PC 12 can be equipped with these functional units. Note that each functional unit is not limited to being realized by a program, and may be realized by a device such as an ASIC (Application Specific Integrated Circuit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), or a conventional circuit module designed to execute the function.
[0030] The PC 12 includes at least an acquisition unit 50, a first generation unit 51, and a second generation unit 52. The PC 12 may also include other functional units. The acquisition unit 50 acquires position information and tilt information of the HMD 10 and the controller 11 detected by the position detection sensor 13. The acquisition unit 50 also acquires information (operation information) input by the user from the controller 11. The acquisition unit 50 also acquires three-dimensional virtual space data from the HDD 43 and the server 14.
[0031] The three-dimensional virtual space data is data on a virtual three-dimensional space represented by x-, y-, and z-axes, and includes object data and image data of one or more objects placed in that space. Taking the above-mentioned virtual conference room as an example, the image data is image data including a whiteboard, floor, walls, ceiling, entrances, and exits. Objects are physical objects placed in the three-dimensional virtual space, and include input objects such as pens for realizing user input, objects such as sticky notes for text input, and user objects for supporting user input.
[0032] The first generation unit 51 generates a user object in a three-dimensional virtual space based on the position information and tilt information of the HMD 10 and the controller 11 acquired by the acquisition unit 50. The user object may be an assistant object that can call functions within an application, a laser object that can select objects, or the like. When the display system is used for a conference, the application may be a conference application, and an example of a function within the application may be pen input.
[0033] The second generation unit 52 generates image data to be displayed on the display of the HMD 10 based on the position information, tilt information, three-dimensional virtual space data, and operation information acquired by the acquisition unit 50, and transmits the image data to the HMD 10. The second generation unit 52 uses the three-dimensional virtual space data based on the position information and tilt information of the HMD in the three-dimensional virtual space to generate image data of an image in the field of view to which a tilt is applied, with the position of the HMD in the three-dimensional virtual space as the starting point.
[0034] 4 is a sequence diagram showing the overall processing flow executed by the display system. When the HMD 10, controller 11, PC 12, and position detection sensor 13 are powered on, the PC 12 acquires three-dimensional virtual space data from its own HDD 43 or server 14 (S1). The HMD 10 and controller 11 emit infrared rays from their respective light sources, and the position detection sensor 13 detects the positions and inclinations of the HMD 10 and controller 11 based on the infrared light sources (S2). The position detection sensor 13 transmits the detected positions and inclinations to the PC 12 as position information and inclination information, respectively (S3).
[0035] The controller 11 accepts an operation from the user and transmits operation information of the accepted operation to the PC 12 (S4). The PC 12 generates image data based on the acquired three-dimensional virtual space data, the position information and tilt information of the HMD 10 and the controller 11 received from the position detection sensor 13, and the operation information received from the controller 11 (S5). The PC 12 then transmits the generated image data to the HMD 10 (S6). The HMD 10 displays the received image data on the display (S7). From this point on, the process from detecting the position and tilt in S2 to transmitting the image data in S5 is repeated until the power of the HMD 10, etc. is turned off, and the image data displayed on the display of the HMD 10 is updated.
[0036] The overall processing flow in the display system has been described above, but a method for generating a user object that is displayed together with image data in a three-dimensional virtual space will now be described in detail.
[0037] User objects must be easy for a user to touch to invoke functionality within an application or to select an object.
[0038] Since the image data is generated with the position of the HMD 10 as the origin and the tilt applied, the user object can also be placed in the three-dimensional virtual space based on the position and tilt of the HMD 10.
[0039] However, when the user's head moves, the user object moves accordingly. Therefore, if only the head moves, it may be difficult to touch the user object depending on the position and orientation of the head. For example, this may occur when the user object moves to the side opposite the dominant hand. The same applies when using the user object to select an object. This does not constitute high usability.
[0040] When moving their arms, people move them up and down and left and right, using the shoulder as a fulcrum. The shoulder is one of the parts of the body that exists between the head and the hand that holds the controller 11. When a user wants to touch a user object, they move their arm to change the position of the controller 11. Therefore, if the user object is positioned based on the position of the shoulder, the user object will not move following the head movement, and even if the shoulder moves, the user object can be positioned within easy reach of the dominant hand.
[0041] When locating a user object based on the shoulder position, it is necessary to estimate the shoulder position based on the position information and tilt information of the HMD 10 and the controller 11 acquired from the position detection sensor 13. To estimate the shoulder position, it is necessary to know the relative positional relationship between the HMD 10 and the shoulder position. The relative positional relationship is the shoulder position as seen from the HMD 10 or the position of the HMD 10 as seen from the shoulder position.
[0042] A first method for determining the relative positional relationship between the user's shoulder position and the HMD 10 will be described with reference to Fig. 5. Fig. 5 is a view of the user as seen from above. The user holds or wears the controller 11 in their hands, stretches their elbows horizontally toward the floor, and, with their elbows stretched, repeatedly moves their hands together to close their arms, then separates their hands to open their arms.
[0043] In such a movement, the left and right hands move in a circular arc with the left and right shoulders as their respective fulcrums. The position information of the controller 11 at this time is acquired by the position detection sensor 13. The position information acquired by the position detection sensor 13 is position information of points on the circular arc. The shoulder positions are points equidistant from each point on the circular arc. Therefore, the position information of the shoulders can be calculated from the position information acquired by the position detection sensor 13.
[0044] The position information of the HMD 10 is acquired by the position detection sensor 13 before the opening or closing of the arms is started. Therefore, the shoulder position based on the position of the HMD 10 can be calculated as the relative positional relationship between the HMD 10 and the shoulder position.
[0045] Fig. 6 is a flowchart showing the flow of processing for determining the relative positional relationship between the user's shoulder position and the HMD 10 using the first method shown in Fig. 5. This processing starts from step 100 when the user starts using the HMD 10 and controller 11, or when the positional deviation detected by the position detection sensor 13 is felt to have increased.
[0046] In step 101, the arm is held horizontally with the elbow extended. In step 102, the arm is opened and closed with the elbow extended. In step 103, position information of the controller 11 is continuously acquired. Here, the position information may be position information for a single movement from the closed state to the open state, but in order to improve the accuracy of the position information, it is preferable to repeat this several times.
[0047] In step 104, the acquired position information is averaged. Because the arm is closed and opened repeatedly, multiple pieces of position information are obtained, such as when the arm is closed and when it is opened. For this reason, the position information for the same closed position or the same open position is averaged. This averaging makes it possible to estimate the shoulder position with high accuracy. In step 105, the relative positional relationship between the HMD 10 and the shoulder position is calculated from the averaged position information, and in step 106, this process ends.
[0048] The shoulder positions are calculated as shoulder positions based on the position of the HMD 10. The calculated shoulder position data is shoulder position data that indicates the relative positional relationship with the HMD 10, and is coordinate data expressed by, for example, the x-axis, y-axis, and z-axis. The shoulder position data is calculated for each of the left and right shoulders.
[0049] A second method for determining the relative positional relationship between the user's shoulder position and the HMD 10 will be described with reference to Fig. 7. Fig. 7 is a view of the user as seen from the front. In the second method, to directly determine the shoulder position, controllers 11 worn on the left and right hands are placed on the left and right shoulders, and position information of the controllers 11 is detected by position detection sensors 13.
[0050] The controller 11 worn on either hand may be placed at either shoulder position, and the controller 11 may be placed at the shoulder position in either pose shown in FIGS. 7(a) and 7(b).
[0051] If the controller 11 is on the left side of the position and orientation of the HMD 10, the position information of the controller 11 is estimated to be left shoulder position information, and if it is on the right side, the position information of the controller 11 is estimated to be right shoulder position information, and each position information is used as shoulder position data.
[0052] Fig. 8 is a flowchart showing the flow of processing for determining the relative positional relationship between the user's shoulder position and the HMD 10 using the second method shown in Fig. 7. This processing can be started from step 200 when the user starts using the HMD 10 and controller 11, or when the user feels that the positional deviation detected by the position detection sensor 13 has become large.
[0053] In step 201, position information of each of the left and right controllers 11 is obtained by the position detection sensor 13. In step 202, position information and tilt information of the HMD 10 are obtained by the position detection sensor 13. In step 203, it is estimated which shoulder (left or right) each controller 11 corresponds to, based on the position information and tilt information of the HMD 10. In step 204, the relative positional relationship between the HMD 10 and the positions of each shoulder is obtained, and in step 205, this process ends.
[0054] After creating shoulder position data, which is position data relative to the HMD 10, the PC 12 executes a process to generate a user object based on the shoulder position. A user object is an object for assisting a user with input, such as an assistant object that invokes a function within an application or a laser object that serves as a selection object for selecting one of the objects placed in the three-dimensional virtual space. The process of estimating these shoulder positions is also executed by the first generation unit 51.
[0055] Referring to FIG. 9, a process of generating an assistant object in a three-dimensional virtual space and generating an image to be displayed on the HMD 10 will be described. This process is executed after determining the relative positional relationship between the position of the user's shoulder and the HMD 10. This process starts from step 300, and in step 301, position information and tilt information of the HMD 10 and position information and tilt information of the controller 11 are acquired from the position detection sensor 13. The position information and tilt information of the controller 11 are the position information and tilt information of the two controllers 11 worn on both the left and right hands. Hereinafter, for ease of explanation, the controller 11 will be simply described as a controller 11, but in reality, it means the two controllers 11 on both hands.
[0056] In step 302, shoulder position data, which is position data relative to the HMD 10, is applied to position information of the HMD 10 (coordinates represented by the x-axis, y-axis, and z-axis) to estimate the shoulder position. The shoulder position data is data that indicates the shoulder position with the HMD 10 as a reference using coordinates. In step 303, position information (coordinates) of an assistant object in a three-dimensional virtual space is calculated using the estimated shoulder position as a reference. The assistant object may be an object of any shape, and the shape, size, etc. are set in advance.
[0057] The PC 12 holds assistant object position calculation data that indicates the relative positional relationship between the shoulder position and the assistant object, and calculates the coordinates of the assistant object in the three-dimensional virtual space by applying the assistant object position calculation data (data in the uvw coordinate system in which the tilt of the HMD 10 is applied to the xyz coordinate system based on the reference direction of the HMD 10) to the shoulder position. The assistant object position calculation data is data that has been adjusted and stored in advance to find the optimal position for the user.
[0058] In step 304, the distance between the controller 11 and the assistant object is created as distance information from the position information of the controller 11 and the calculated position information of the assistant object. The PC 12 can be provided with a calculation unit that calculates the distance to create this distance information. The calculation unit also performs other calculation processes such as the arithmetic average described above. In step 305, it is determined whether the created distance information is longer than a predetermined distance. If the distance is equal to or shorter than the predetermined distance, the process proceeds to step 306, where application of the assistant object position calculation data to the shoulder position is stopped, and the coordinates (coordinates expressed in the xyz coordinate system) of the assistant object in the three-dimensional virtual space at the time when application of the assistant object position calculation data was stopped are fixed. The PC 12 can be provided with a determination unit that determines whether the distance information is longer than a predetermined distance, and fixes the coordinates if the distance is equal to or shorter than the predetermined distance. The determination unit also performs other determination processes.
[0059] The coordinates of the fixed assistant object based on the shoulder position do not change each time the HMD 10 and the controller 11 are powered on and started up. Therefore, the coordinates of the fixed assistant object can use the coordinates of the assistant object determined in the process of displaying the previous image.
[0060] If the coordinates of the assistant object are not fixed, when you try to call a function by touching the assistant object and your shoulder moves, the assistant object will move and it will be difficult to touch. This is because the assistant object moves with the movement of your shoulder. Therefore, the coordinates of the assistant object are fixed so that the assistant object will not move with the movement of your shoulder when you try to touch the assistant object.
[0061] If it is determined in step 305 that the distance is longer than the predetermined distance, the process proceeds directly to step 307. In step 307, the assistant object is placed in the three-dimensional virtual space based on the coordinates of the assistant object. In step 308, image data of an image in the field of view centered on the position of the HMD 10 and applying the tilt of the HMD 10 is generated based on the position information and tilt information of the HMD 10. The generated image data is displayed on the display of the HMD 10, and in step 309, the assistant object is generated, and the process of generating the image to be displayed on the HMD 10 is completed.
[0062] FIG. 10 is a diagram showing a display example of a generated assistant object. The image displaying the three-dimensional virtual space includes a whiteboard 60, and a virtual user's hand 61 is shown in front of the whiteboard 60. The assistant object 62 is a polyhedron having a surface that can be touched by the fingers of the hand 61. Since the assistant object 62 is generated and placed based on the position of the shoulder, it is easy to touch the assistant object 62 with a finger and to call up its functions. Information such as letters representing functions is displayed on the surface of the polyhedron, and since the assistant object 62 is placed based on the position of the shoulder, the information is easy to see.
[0063] 10, the user operates the controller 11 and touches an assistant object 62 with the fingers of a hand 61 in a three-dimensional virtual space. If the assistant object 62 is configured as a polyhedron, information such as characters representing different functions may be displayed on each face, and a predetermined function may be invoked by touching a predetermined face. The user can send an instruction to invoke a function to the PC 12 by touching the assistant object 62 with the fingers of the hand 61 in the three-dimensional virtual space and pressing a predetermined button on the controller 11.
[0064] FIG. 11 is a diagram showing a display example when an assistant object 62 is touched with the fingers of a hand 61 to call a function in an application. The function is displayed as a pop-up (a screen that appears to jump out into the foreground) by, for example, generating an object representing the function for the HMD 10 by the PC 12 and invoking the function by touching the object. For this reason, the PC 12 may be provided with a third generation unit for generating objects representing the function. In FIG. 11, a marker 63 for writing on the whiteboard 60, a blackboard eraser 64, etc. are generated and displayed in the image. Other functions may include a function for moving the whiteboard 60 and a function for copying content written on the whiteboard 60.
[0065] The user can write characters on the whiteboard 60 by operating the controller 11, picking up the marker 63 in the hand 61 in the three-dimensional virtual space, moving the marker 63 to the whiteboard 60, and moving the hand 61. The user can also erase characters written on the whiteboard 60 by operating the controller 11, picking up the eraser 64 in the hand 61 in the three-dimensional virtual space, and moving the hand 61.
[0066] 12, a process for generating a laser object in a three-dimensional virtual space and generating an image to be displayed on the HMD 10 will be described. This process is also executed after determining the relative positional relationship between the position of the user's shoulder and the HMD 10. This process starts from step 400, and in step 401, position information and tilt information of the HMD 10 and the controller 11 are acquired from the position detection sensor 13.
[0067] In step 402, shoulder position data is applied to position information (coordinates represented by the x-axis, y-axis, and z-axis) of the HMD 10 to estimate the shoulder position. In step 403, distance information indicating the distance between the shoulder position and the controller 11 is created based on the estimated shoulder position and position information of the controller 11. In step 404, the length (number of pixels) of the laser object is determined based on the created distance information. The laser object is a rod-shaped object, and the longer the distance between the shoulder and the controller 11, the longer the length of the laser object.
[0068] Fig. 13 shows the relationship between the length of the laser object and the distance between the shoulder and the controller 11. Fig. 13(a) is a graph showing the relationship in which the length of the laser object increases quadratically with the distance. Fig. 13(b) is a graph showing the relationship in which the length increases at a constant rate up to a certain distance, and then increases quadratically once the certain distance is exceeded.
[0069] As shown in Fig. 13(c), when the length of the laser object increases linearly according to the distance between the shoulder and the controller 11, the slope of the graph becomes steeper than the slope of the graph up to a certain distance shown in Fig. 13(b). In this case, if the target object is nearby, extending the arm a little may make the laser object too long, and if the target object is far away, the laser object may not reach the target object even if the arm is extended a long way.
[0070] However, as shown in Figures 13(a) and (b), by changing the length of the laser object so that the rate at which the length of the laser object changes increases as the distance increases, it becomes possible to appropriately select the target object whether the target object is close or far away. Note that the length of the laser object is not limited to the examples shown in Figures 13(a) and (b) as long as it can be changed so that the rate at which the length of the laser object changes increases as the distance increases.
[0071] Referring again to FIG. 12, once the length of the laser object is determined from the created distance information, the laser object is placed in the three-dimensional virtual space in a specific direction, starting from the position of the controller 11. The specific direction is, for example, a predetermined tilt direction of the controller 11 or a direction connecting the controller 11 and the shoulder position. If the controller 11 has a grip portion to be held in the hand, the grip portion has a shape that extends in a certain direction. The predetermined tilt direction is the direction in which the grip portion extends (longitudinal direction).
[0072] In step 405, it is determined whether the laser object placed in the three-dimensional virtual space has passed through a specific object in the three-dimensional virtual space. A specific object is an object that can be selected and then moved or otherwise performed when a specific input (for example, a movement instruction via a button input) is received. If it is determined that the laser object has passed through the specific object, the process proceeds to step 406, where selection processing is performed on the specific object that has passed. Specifically, a selection flag is assigned to the specific object that has passed, the laser object that has passed through the specific object is deleted, and the process proceeds to step 407. The PC 12 may be equipped with a selection processing unit that performs this selection processing. On the other hand, if it is determined in step 405 that the laser object has not passed through the specific object, the process proceeds to step 408.
[0073] In step 407, a user operation is accepted. The user operation is, for example, moving an object, and object movement includes input of an instruction to move the object. Object movement will be described later. In step 408, in the three-dimensional virtual space, image data is generated in a field of view direction centered on the position of HMD 10 and applying the tilt of HMD 10 based on the position information (coordinates) of the laser object and the position information and tilt information of HMD 10, and this is displayed on the display. Then, in step 409, a laser object is generated, and the process of generating an image to be displayed on HMD 10 is completed.
[0074] Figure 14 is a diagram showing an example of selecting an object using a laser object. In step 404 of Figure 12, a laser object 65 is generated as a rod-like object extending in the direction of the extension of a virtual user's hand 61 in the three-dimensional virtual space, and is placed in the three-dimensional virtual space. In Figure 14(a), the laser object 65 is placed in the three-dimensional virtual space so as to extend from the hand 61. In Figure 14(b), in step 406 of Figure 12, an object 66 in the three-dimensional virtual space is selected, and the laser object 65 that has passed through is deleted.
[0075] 15 is a diagram showing an example of a user operation on the object 66 selected in step 407 of FIG. 12, illustrating the movement of the object 66. By changing the position or tilt of the controller 11, the user can change the direction in which the object 66 is placed, as shown in FIG. 15(a). Furthermore, by pulling back the arm while wearing the controller 11, the user can pull the object 66 toward the user, as shown in FIG. 15(b).
[0076] 16 is a flowchart showing the flow of processing for moving an object 66 selected by a laser object 65, as an example of specific processing in step 407 in Fig. 12. Here, the movement processing will be described, but the operation on the object 66 selected by the laser object 65 is not limited to moving the object 66.
[0077] 16 is performed after placing the laser object 65 in the three-dimensional virtual space and completing the selection of a specific object. This process starts from step 500, and in step 501, it is confirmed whether or not a selection flag has been assigned to the specific object 66. If a selection flag has been assigned, the process proceeds to step 502, where it is determined whether or not a movement instruction has been issued by the controller 11. The controller 11 is provided with an operation I / F 30 such as a button, and determines whether or not a movement instruction has been issued based on whether or not a button has been pressed.
[0078] If a movement instruction has been given, the process proceeds to step 503, where movement processing of the object 66 to which a selection flag has been assigned is executed based on the user's operation of the controller 11. The PC 12 may be equipped with a movement processing unit that executes movement processing of the object 66. If the movement processing is an attracting process, the coordinates of the object 66 to which a selection flag has been assigned are instantaneously moved to a state where they are closer to the user's hand. If the movement processing is a process other than an attracting process, the object 66 to which a selection flag has been assigned moves up, down, left, right, forward, and backward in the three-dimensional virtual space. Whether to execute the attracting process or another movement process can be selected by changing the button to be pressed among multiple buttons provided on the controller 11. Note that, in addition to changing the button, selection may also be made by a gesture of pulling the hand closer (the position or tilt of the controller 11).
[0079] If the selection flag is not set in step 501, if there is no movement instruction in step 502, or if the execution of the movement process has ended in step 503, the process proceeds to step 504, where it is determined whether or not there is an input instruction from the user. An input instruction from the user may be, for example, a voice input by the user or a character input in the three-dimensional virtual space. Like the presence or absence of a movement instruction, the presence or absence of an input instruction can be determined by whether or not a button on the controller 11 has been pressed. The presence or absence of an input instruction can also be determined by whether or not there has been an operation using a gesture to bring the object 66 closer to the user's mouth.
[0080] If an input instruction is received from the user in step 504, the process proceeds to step 505, where input processing is executed for the object 66 to which the selection flag is attached. A specific description of the input processing will be given later. If no input instruction is received in step 504, or if the input processing is completed in step 506, the movement processing of the object 66 is completed. If the object 66 is to be further moved, the processing can be started from step 500.
[0081] 17 is a flowchart showing the flow of voice input as an example of user input. Processing starts from step 600, and in step 601, position information of the HMD 10, position information of the controller 11, and position information of the laser object 65 are acquired. The position information of the laser object 65 is the position of the tip of the laser object 65, i.e., the object 66 selected by the laser object 65.
[0082] In step 602, distance information indicating the distance between the HMD 10 and the laser object 65 is created from the position information of the HMD 10 and the position information of the laser object 65 in the three-dimensional virtual space. The distance information is constantly created after the movement process is executed.
[0083] In step 603, the position of the user's mouth is estimated based on the position of the HMD 10, and it is determined whether the distance between the user's mouth and the laser object 65 in the three-dimensional virtual space is shorter than a predetermined distance. If the user pulls the object 66 closer and the distance becomes shorter than the predetermined distance, the process proceeds to step 604, where voice input is performed. The HMD 10 receives voice input through the microphone 26, which functions as an input receiving unit, and sends the input voice data to the PC 12. The PC 12 performs voice recognition processing on the voice data acquired from the HMD 10. In the voice recognition processing, what the user says is converted into image data containing text through voice recognition, and the converted image data is pasted on the object 66 that has been moved. The PC 12 may include a data processing unit that performs predetermined processing, such as voice recognition processing, on the input data. Note that the microphone that receives voice input is not limited to the microphone 26 of the HMD 10, but may be a microphone separately connected to the PC 12.
[0084] If the distance is longer than the predetermined distance in step 603, the voice input process ends in step 605 without executing the voice input.
[0085] 18 is a diagram showing a display example when voice input is being performed. The figure shows a state in which an object 66 is moved to the mouth of a user 67 wearing an HMD 10 on his head, and voice input is being performed to the object 66.
[0086] Up to this point, the display system has been described as including the HMD 10, controller 11, PC 12, and position detection sensor 13 shown in Fig. 1, but the configuration of the display system is not limited to this. As shown in Fig. 19, the functions of the PC 12 and position detection sensor 13 may be incorporated into the HMD 10, and the display system may be configured only with the HMD 10 and controller 11.
[0087] When detecting the positions of the HMD 10 and the controller 11, the HMD 10 may be equipped with an imaging device (camera) in order to detect the positions of the HMD 10 and the controller 11. The HMD 10 and the controller 11 may be equipped with a sensor for detecting the tilt of the HMD 10 and the controller 11.
[0088] The position and tilt of the HMD 10 can be calculated from the size, direction, tilt, etc. of an image of a marker or the like placed at a reference position captured by a camera. The position and tilt of the controller 11 can also be calculated from an image captured by the camera. Note that if the HMD 10 is equipped with a camera, the controller 11 does not need to be used, and a marker can be attached to the hand, and an image including the marker can be captured using the camera of the HMD 10, and the position and tilt of the hand can be calculated from the captured image.
[0089] Fig. 20 is a diagram showing the hardware configuration of the HMD 10 and the controller 11 when the configuration shown in Fig. 19 is adopted. Similar to the configuration shown in Fig. 2, the HMD 10 includes an external I / F 20, a CPU 21, a display 22, a memory 23, an HDD 24, and a microphone 26. The HMD 10 further includes a sensor 27 and a camera 28.
[0090] The sensor 27 is a gyro sensor that detects angular velocity, an acceleration sensor that detects acceleration, or the like, and detects the tilt and orientation of the HMD 10. The camera 28 recognizes markers in the three-dimensional virtual space and attached to the controller 11, and detects the position and tilt of the HMD 10 and the controller 11.
[0091] 2, the controller 11 also includes an operation I / F 30 and an external I / F 31. The controller 11 includes a sensor 33 instead of the light source 32. The sensor 33 is a gyro sensor, an acceleration sensor, or the like, similar to the sensor 27 of the HMD 10, and detects the tilt and orientation of the controller 11.
[0092] The tilt information of the HMD 10 and the controller 11 may use information from the sensor 27 and the sensor 33, or may use information detected by the camera 28 of the HMD 10. The position and tilt of the HMD 10 and the controller 11 may be detected from an image captured by the camera 28 of the HMD 10. Furthermore, the information detected by the sensor 27 and the camera 28 of the HMD 10 and the information detected by the controller 11 may be compiled in the HMD 10, and the position and tilt of the HMD 10 and the controller 11 may be detected based on the obtained information.
[0093] Instead of using the controller 11, the shape and position of the hand may be detected by the camera 28, and the tilt may be detected from the shape of the hand.
[0094] 20, the HMD 10 may include a sensor 27 and a camera 28, the controller 11 may include a sensor 33, and may further include the position detection sensor 13 shown in FIG. 1 as an external sensor. In this case, it is possible to select and switch between using the sensor and camera of the HMD 10 and the controller 11 or using the external position detection sensor 13. For example, when highly accurate information is desired for position information and tilt information, the position detection sensor 13 can be used, and when responsiveness is required, the sensors 27, 33 and the camera 28 can be used.
[0095] By providing the device, system, method, and program of the present invention, usability in a three-dimensional virtual space can be improved because there is no need to check one's hands or use both hands when placing a user object. Furthermore, the position of the shoulder is used as the starting point of the arm to determine the placement position of the user object, improving the accuracy of the placement position. Furthermore, by determining the placement position based on the shoulder position, it becomes possible to handle the user object with a sense close to human intuition.
[0096] The present invention has been described above in terms of the embodiments of an information processing device, a display system, an information processing method, and a program. However, the present invention is not limited to the above-described embodiments, and other modifications, additions, changes, deletions, and other changes can be made within the scope of what one skilled in the art can conceive. Furthermore, any embodiment that achieves the functions and effects of the present invention is within the scope of the present invention. [Explanation of symbols]
[0097] 10...HMD, 11...controller, 12...PC, 13...position detection sensor, 14...server, 15...network, 20...external I / F, 21...CPU, 22...display, 23...memory, 24...HDD, 25...light source, 26...microphone, 30...operation I / F, 31...external I / F, 32...light source, 40...CPU, 41...ROM, 42...RAM, 43...HDD, 44...external I / F, 45...input / output I / F, 46...input device, 47...display device, 50...acquisition unit, 51...first generation unit, 52...second generation unit, 60...whiteboard, 61...hand, 62...assistant object, 63...marker, 64...blackboard eraser, 65...laser object, 66...object, 67...user [Prior art documents] [Patent documents]
[0098] [Patent Document 1] Japanese Patent Application Laid-Open No. 2018-190395
Claims
1. a generating means for generating a first object in a three-dimensional virtual space based on position information of a display means worn by a user and position information of an operation means operated by the user or a hand of the user; an input receiving means for receiving a voice input from the user to a second object in accordance with the position of the user's mouth estimated based on the position information of the display means and the position of the first object in the three-dimensional virtual space; An information processing device comprising:
2. An information processing device as described in claim 1, further comprising a voice recognition means that performs voice recognition processing on the voice data of the input voice, converts it into image data including characters, and pastes the converted image data onto the second object.
3. The method further includes determining whether a distance between the user's mouth and the first object in the three-dimensional virtual space is shorter than a predetermined distance; The information processing apparatus according to claim 1 , wherein the input receiving means receives a voice input from the user to the second object when it is determined that the distance is shorter than the predetermined distance.
4. An information processing device as described in Claim 3, wherein the input accepting means does not accept voice input from the user to the second object if it is determined that the distance is longer than the specified distance.
5. An information processing device described in any one of claims 1 to 4, wherein the first object is an object that can select the second object.
6. An information processing device as described in Claim 5, further comprising a movement processing means that executes a process of moving the second object selected using the first object to a position specified by the user.
7. A display system including an information processing device, The information processing device, a generating means for generating a first object in a three-dimensional virtual space based on position information of a display means worn by a user and position information of an operation means operated by the user or a hand of the user; an input receiving means for receiving a voice input from the user to a second object in accordance with the position of the user's mouth estimated based on the position information of the display means and the position of the first object in the three-dimensional virtual space; a display system including:
8. The display system according to claim 7, further comprising a display means that is attached to the user within the information processing device or separately from the information processing device and that displays the first object generated by the information processing device and the second object selected using the first object in the three-dimensional virtual space.
9. A method executed by an information processing device for generating an image based on virtual space data representing a three-dimensional virtual space, generating a first object in the three-dimensional virtual space based on position information of a display means worn by a user and position information of an operation means operated by the user or a hand of the user; receiving a voice input from the user to a second object according to the position of the user's mouth estimated based on the position information of the display means and the position of the first object in the three-dimensional virtual space; A method comprising:
10. A program for causing a computer to execute a process for generating an image based on virtual space data representing a three-dimensional virtual space, generating a first object in the three-dimensional virtual space based on position information of a display means worn by a user and position information of an operation means operated by the user or a hand of the user; receiving a voice input from the user to a second object according to the position of the user's mouth estimated based on the position information of the display means and the position of the first object in the three-dimensional virtual space; A program that executes.