Information processing device, information processing method, and program
The head-mounted display system addresses the impracticality of simultaneous virtual object viewing in MR by implementing display modes that allow users to view virtual objects from shared or optimal viewpoints, eliminating the need for physical or virtual movement.
Patent Information
- Application Number
- JP2021133964
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-08-19
- Publication Date
- 2025-09-08
- Estimated Expiration
- 2041-08-19
AI Technical Summary
Existing mixed reality (MR) technologies require multiple users to physically move within the real space or adjust their viewpoints in virtual space to view virtual objects simultaneously, which is not practical.
A head-mounted display system that includes an instruction acquisition, mode determination, viewpoint determination, and image generation mechanism to allow users to view virtual objects without moving their real-space position or virtual-space viewpoint, using presenter viewpoint sharing and optimal viewpoint sharing modes.
Enables multiple users to view virtual objects without the need for physical movement or viewpoint adjustment, enhancing the practicality and realism of MR experiences.
Smart Images

Figure 0007735121000001 
Figure 0007735121000002 
Figure 0007735121000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing technique for superimposing an image of a virtual object on an image of real space. [Background technology]
[0002] In recent years, mixed reality (MR) technology has become widespread. MR technology allows users wearing a head-mounted display (HMD) to experience a sense of mixed reality by superimposing a computer-generated virtual image onto an image of the real world captured by the HMD's camera. Furthermore, Patent Document 1 discloses a technology that presents virtual objects displayed in a virtual space to users at multiple locations in the same orientation as other users are viewing them. Patent Document 2 discloses a technology that allows users to share MR space images of other users and further operate them on their own terminals. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2017-120650 [Patent Document 2] Japanese Patent Publication No. 2020-74066 Summary of the Invention [Problem to be solved by the invention]
[0004] When multiple users simultaneously view virtual objects placed in an MR space, it may be necessary to move users within the real space or move the viewpoints of other users within a specific user's virtual space. However, this is not realistic.
[0005] Therefore, an object of the present invention is to enable a plurality of users to view a virtual object without the need to move their position in the real space or the viewpoint position in the virtual space. [Means for solving the problem]
[0006] The present invention provides a head-mounted display system including: an instruction acquisition means for acquiring an instruction from a first user wearing a first head-mounted display; a mode determination means for determining a display mode of a second head-mounted display worn by a second user in accordance with the instruction; a viewpoint determination means for determining a viewpoint with respect to a virtual object in accordance with the determined display mode; and an image generation means for generating an image in a mixed reality space, the mode determination means determines whether to use a first display mode in which a viewpoint of a specific user of the first user and the second user with respect to the virtual object is determined as a viewpoint of another user other than the specific user with respect to the virtual object, or to use a second display mode in which a viewpoint set in advance with respect to the virtual object is determined as a viewpoint of the other user with respect to the virtual object; The image generation means generates an image of the mixed reality space to be displayed on the second head-mounted display by superimposing an image of the virtual object seen from the determined viewpoint different from the viewpoint of the second user onto an image of the real space obtained by the second head-mounted display. [Effects of the Invention]
[0007] According to the present invention, a plurality of users can view a virtual object without having to move their positions in the real space or their viewpoints in the virtual space. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 2 is a diagram illustrating an example of a hardware configuration of an information processing device. [Figure 2] FIG. 2 is a functional block diagram showing the functional configuration of the information processing device. [Figure 3] FIG. 1 is a diagram illustrating an example of a mixed reality space. [Figure 4] FIG. 10 is a diagram illustrating an example of a presenter viewpoint sharing mode. [Figure 5] FIG. 10 is a diagram illustrating an example of an optimal viewpoint sharing mode. [Figure 6] 1 is a flowchart showing a series of information processing flows according to the present embodiment. [Figure 7] 4 is a flowchart of viewpoint determination and image generation processing according to the first embodiment. [Figure 8] FIG. 10 is a diagram illustrating an example of a mixed reality space according to the second embodiment. [Figure 9] FIG. 10 is a diagram illustrating an example of avatar placement. [Figure 10] 10 is a flowchart of a viewpoint determination and image generation process according to the second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, embodiments of the present invention will be described with reference to the drawings. The following embodiments do not limit the present invention, and not all of the combinations of features described in the present embodiments are necessarily essential to the solution of the present invention. The configurations of the embodiments may be modified or changed as appropriate depending on the specifications of the device to which the present invention is applied and various conditions (such as usage conditions and usage environment). Furthermore, a configuration may be achieved by appropriately combining parts of each of the embodiments described below. In the following embodiments, the same components will be described with the same reference symbols.
[0010] First Embodiment Fig. 1 is a diagram showing an example of the hardware configuration of an information processing device 201 according to the first embodiment, and Fig. 2 is a functional block diagram showing the functional configuration of the information processing device 201 according to this embodiment. Before explaining the configurations in Fig. 1 and Fig. 2, an example in which multiple users simultaneously view virtual objects placed in a mixed reality (MR) space using head-mounted displays (HMDs) will be explained with reference to Figs. 3 to 5.
[0011] Here, as an example, a specific user acts as a presenter of a presentation, and explains about automobiles to the participating users by superimposing an image of a virtual object representing an automobile on an image of real space and displaying it on an HMD. The participants are assumed to be users other than the specific user. While giving a presentation about automobiles, the specific user who is the presenter may change the location or direction in which the participating users should pay attention. In this case, with existing MR technology, each user must move within the real space, and the participants must move their viewpoint within the presenter's virtual space, which is not realistic. Therefore, the information processing device 201 of this embodiment performs information processing as described below.
[0012] The information processing device 201 of this embodiment allows other users to change their viewpoints of virtual objects in response to instructions from a specific user regarding the virtual object. In this embodiment, the viewpoint of a virtual object includes the position and orientation (direction) of the viewpoint. As a result, the information processing device 201 of this embodiment allows multiple users to view a virtual object without having to move the position in real space or the viewpoint position in virtual space. Furthermore, the information processing device 201 of this embodiment has a first display mode in which the viewpoint position and orientation of a specific user are shared with other users, and a second display mode in which a preset optimal viewpoint position and orientation for a virtual object are shared with each user. In this embodiment, the first display mode in which the viewpoint position and orientation of a specific user are shared with other users is called a "presenter viewpoint sharing mode." Furthermore, the second display mode in which each user shares a preset optimal viewpoint position and orientation for a virtual object is called an "optimal viewpoint sharing mode." Note that in this embodiment, in the optimal viewpoint sharing mode, the preset optimal viewpoint position and orientation for a virtual object are shared with users other than the specific user, but may be shared with all users, including the specific user.
[0013] When either the presenter viewpoint sharing mode or the optimal viewpoint sharing mode is set, the information processing device 201 of this embodiment enables each user to share the viewpoint position and orientation relative to the virtual object according to the set display mode. Note that which display mode, the presenter viewpoint sharing mode or the optimal viewpoint sharing mode, is to be used may be set in advance, or if not set in advance, it may follow a setting associated with the virtual object to be displayed. Also, which display mode is to be used can be arbitrarily determined by, for example, a specific user or another user via operation of a mouse, controller, or the like. In other words, which display mode is to be used can be arbitrarily determined for each user.
[0014] FIG. 3 is a diagram showing an example of multiple users 101-104 in a real space 100 and a virtual object 105 placed in a virtual space corresponding to the real space 100. The virtual object 105 represents a typical automobile having two headlights 106. Each of the users 101-104 wears an HMD, and each HMD is equipped with an imaging device (camera). The HMD of each of the users 101-104 displays the virtual object 105 superimposed on a real space image captured of the real space 100. In other words, each of the users 101-104 is viewing the virtual object 105 virtually placed in the real space through the HMD. Note that FIG. 3 also shows images 107-110 displayed on the HMDs of the users 101-104. For example, image 107 shows an image displayed on the HMD of user 101, and similarly, image 108 shows an image displayed on the HMD of user 102, image 109 shows an image displayed on the HMD of user 103, and image 110 shows an image displayed on the HMD of user 104. In this example, it is assumed that user 101 is a specific user and the presenter of the presentation, while the other users 102 to 104 are participants in the presentation.
[0015] Here, it is assumed that presenter user 101 indicates headlight 106 of virtual object 105 as a location that he / she wants participant users 102 to 104 to pay attention to. In this case, information processing device 201 enables each user to share the viewpoint position and orientation with respect to the virtual object, depending on the display mode that is set to either presenter viewpoint sharing mode or optimal viewpoint sharing mode.
[0016] For example, in the presenter viewpoint sharing mode, the information processing device 201 acquires the viewpoint position and orientation in the virtual space of the presenting user 101 with respect to the virtual object 105. Then, the information processing device 201 generates an image in which the virtual object 105 visible from the viewpoint position and orientation is superimposed on the real space captured by the cameras of the HMDs of the participant users 102 to 104, and displays the image on the HMDs of the participant users 102 to 104. As a result, an image of the headlights 106 of the virtual object 105 seen from the viewpoint position and orientation of the specific user 101 is displayed on the HMDs of the participant users 102 to 104.
[0017] Furthermore, for example, in the optimal viewpoint sharing mode, the information processing device 201 acquires an optimal viewpoint position and orientation that are preset as the viewpoint position and orientation from which the headlights 106 of the virtual object 105 are most easily visible. Then, the information processing device 201 generates an image in which the virtual object 105 seen from the viewpoint position and orientation is superimposed on the real space captured by the cameras of the HMDs of the users 102 to 104, and displays the image on the HMDs of the users 102 to 104. As a result, an image in which the headlights 106 of the virtual object 105 are viewed from the optimal viewpoint position and orientation is displayed on the HMDs of the users 102 to 104. Note that in the optimal viewpoint sharing mode, the optimal viewpoint position and orientation for the virtual object 105 do not necessarily need to be shared with the HMD of a specific user 101; however, in this embodiment, the optimal viewpoint position and orientation are also shared with the HMD of a specific user 101. In this way, the information processing device 201 enables the viewpoint position and orientation for the virtual object to be shared among the users according to the set display mode.
[0018] 4 is a diagram showing an example of the presenter viewpoint sharing mode. Images 601 to 604 represent images displayed on the HMDs of users 101 to 104. In the presenter viewpoint sharing mode, when a specific user 101 points to the headlight 106 of the virtual object 105, the information processing device 201 shares the viewpoint position and orientation of the user 101 in the virtual space with respect to the virtual object 105 with the HMDs of the other users. As a result, an image is displayed in the HMDs of each user 101 to 104, in which the virtual object 105 seen from the viewpoint position and orientation of the specific user 101 who is the presenter is superimposed on a real image seen from the viewpoint position and orientation of each user in real space.
[0019] FIG. 5 is a diagram used to explain the optimal viewpoint sharing mode. Images 701 to 704 represent images displayed on the HMDs of the corresponding users 101 to 104. A viewpoint position 705 represents the viewpoint position and orientation at which the headlight 106 of the virtual object 105 is most easily seen. In the optimal viewpoint sharing mode, when a specific user 101 points to the headlight 106 of the virtual object 105, the information processing device 201 shares the optimal viewpoint position 705 for the headlight 106 with the HMDs of the users 101 to 104 in real space. As a result, the HMDs of the users 101 to 104 display the virtual object 105 as seen from the optimal viewpoint position 705 and orientation for the headlight 106 superimposed on a real image as seen from the viewpoint position of each user in real space.
[0020] The configuration and operation of the information processing device 201 of this embodiment will be described below. 1 is a diagram showing the configuration of a personal computer (PC) as an example of the hardware configuration of an information processing device 201. The information processing device 201 generates an image of an MR space in which a virtual object generated by computer graphics (CG) technology is superimposed on an image of real space captured by an imaging device (camera) mounted on the HMD of each user 101 to 104, and transmits the image to each HMD. Note that since the MR technology is an existing technology, a detailed description thereof will be omitted here.
[0021] The CPU 202 is a system control unit that controls the entire information processing apparatus, and also executes an information processing program to implement information processing according to this embodiment. The ROM 203 is a read-only memory that stores programs and parameters that do not require modification, such as basic programs and initial data. The RAM 204 is a memory for temporarily storing input information, calculation results in information processing, image processing, and the like.
[0022] The operation unit 210 includes operation devices such as a keyboard, mouse, and controller that are capable of pointing operations and inputting various commands, and acquires operation instructions, commands, and the like from a user. In the present embodiment, information from the operation unit 210 can include instruction information when a specific user 101 points to a virtual object. When a specific user 101 inputs an instruction for a virtual object, the instruction information is sent to the input detection unit 206. The operation unit 210 may include not only a function for acquiring instruction input by a pointing device, but also a function for acquiring instruction input by voice. The input detection unit 206 receives required data input. In this embodiment, the required data includes operation instructions and commands from the operation unit 210.
[0023] The recording unit 205 is a device capable of writing and reading various information, such as a hard disk or memory card built into or external to the information processing device, or a memory card, flexible disk, or IC card detachable from the information processing device. The information processing program according to this embodiment is recorded in the recording unit 205, read from the recording unit 205, loaded into the RAM 204, and executed by the CPU 202. The information processing program may also be stored in the ROM 203. The recording unit 205 can also record information on operation instructions input from a specific user 101 via the input detection unit 206, information on the positions and orientations of the users 101-104 detected by each HMD, and information on the positions and orientations of virtual objects. Examples of instructions input from a specific user 101 include an instruction on a part of a virtual object that the specific user 101 wants the other users 102-104 to focus on, or an instruction on one of the parts of a virtual object that is divided into parts. In this embodiment, the recording unit 205 also records information on an optimal viewpoint position and orientation (direction) preset for the virtual object.
[0024] The GPU board 208 is a general-purpose graphics board that performs processes such as image generation and composition. The GPU board 208 in this embodiment can perform image generation and composition processes such as superimposing an image of a virtual object generated using CG technology on an image of real space captured by the camera of the HMD. Note that the image of the virtual object may be generated in advance and recorded in the recording unit 205.
[0025] The display driver 209 is software for controlling the display unit 211 . The display unit 211 is an electronic display device such as a liquid crystal display device mounted on each HMD of the users 101 to 104. The display unit 211 of this embodiment allows the user to give instructions on the display using a mouse cursor (not shown) or the like, and displays images of the MR world or the like generated by the CPU 202 on the GPU board 208 based on information from the recording unit 205. Note that the example in FIG. 1 is based on the assumption that processing is performed by a PC, and therefore the operation unit 210 and the display unit 211 are connected as external components; however, if the HMD has the information processing function according to this embodiment, the operation unit 210 and the display unit 211 are also included as internal components.
[0026] The communication I / F 207 is an interface unit capable of transmitting and receiving data to and from an operation device, a cloud, etc. In this embodiment, the communication I / F 207 can receive, via a network, data of real images captured by cameras provided in the HMDs of the users 101 to 104, position and orientation information detected by the HMDs of the users 101 to 104, and the like.
[0027] FIG. 2 is a functional block diagram showing an example of the functional configuration of an information processing device 201 according to this embodiment. The information processing device 201 has, as its functional configuration, an operation unit 210, an object instruction acquisition unit 305, a mode determination unit 306, a viewpoint determination unit 307, an image generation unit 308, and a user position and orientation recording unit 301 and an object position and orientation recording unit 303 as position and orientation acquisition units.
[0028] The user position and orientation recording unit 301 acquires and records the position and orientation of the user in a mixed reality space (MR space). In this embodiment, the user position and orientation recording unit 301 acquires position coordinates and line-of-sight vector information from each HMD worn by the users 101 to 104, and stores them as position and orientation information of the users 101 to 104 (referred to as user position and orientation information 302). The user position and orientation information 302 is stored in the recording unit 205, for example.
[0029] The object position and orientation recording unit 303 acquires and records the position and orientation of a virtual object in mixed reality space. In this embodiment, the object position and orientation recording unit 303 acquires the position, orientation, size, etc. of the virtual object in the virtual space and stores them as object position and orientation information 304. The object position and orientation information 304 is stored in the recording unit 205, for example.
[0030] The operation unit 210 receives various instructions and commands from the user, such as an instruction to select a virtual object and an instruction to switch the display mode. The object instruction acquisition unit 305 acquires information about an instruction given by a user to a virtual object in mixed reality space via the operation unit 210. In this embodiment, the object instruction acquisition unit 305 acquires an instruction given to a virtual object by a specific user 101, but can also acquire instructions given by other users 102 to 104.
[0031] The mode determination unit 306 determines which display mode to apply, the presenter viewpoint sharing mode or the optimal viewpoint sharing mode. If the display mode to be used is set in advance, the mode determination unit 306 determines the display mode according to the setting. If the user instructs one of the display modes via the operation unit 210, the mode determination unit 306 determines the display mode according to the instruction from the user. If the display mode is not set in advance or if there is no instruction from the user, the mode determination unit 306 determines the display mode according to the setting associated with the virtual object to be displayed. Note that when a user instructs a display mode, the instruction is made by a specific user 101 who is the presenter, but the instruction may also be made by other users 102 to 104 who are participants.
[0032] The viewpoint determination unit 307 determines a viewpoint for a virtual object in mixed reality space in accordance with the display mode determined by the mode determination unit 306. In this embodiment, the viewpoint determination unit 307 determines a viewpoint position and orientation (line of sight) as a viewpoint for a virtual object, based on the user position and orientation information 302, the object position and orientation information 304, and the display mode determined by the mode determination unit 306. For example, when the display mode is the presenter viewpoint sharing mode, the viewpoint determination unit 307 determines the viewpoint position and orientation in the virtual space for a virtual object 105 of a specific user 101 as the viewpoint position and orientation to be shared by the HMDs of the other users 102 to 104. On the other hand, when the display mode is the optimal viewpoint sharing mode, the viewpoint determination unit 307 determines the optimal viewpoint position and orientation, which are preset as the most visible viewpoint position for the headlights 106 of the virtual object 105, as the viewpoint position and orientation to be shared by the HMDs of the users 101 to 104. Furthermore, the viewpoint determination unit 307 can also arbitrarily switch between users who share the viewpoint position and direction based on an instruction from the operation unit 210, for example.
[0033] The image generation unit 308 generates an image of mixed reality space including the virtual object 105 viewed from the viewpoint determined by the viewpoint determination unit 307. In this embodiment, the image generation unit 308 generates an image in which an image of the virtual object is superimposed on an image of real space, based on the user position and orientation information 302, the object position and orientation information 304, and the viewpoint position and orientation determined by the viewpoint determination unit 307. The image generation unit 308 then sends the image in which the virtual object is superimposed on the image of real space to the HMD of each of the users 101 to 104. As a result, an image in which the virtual object is superimposed on the real space image is displayed on the HMD of each of the users 101 to 104. That is, in the presenter viewpoint sharing mode, an image in which the virtual object 105 viewed from the viewpoint position and orientation of a specific user 101 is superimposed on the real space image captured by the camera of each user's HMD is displayed on the HMD of the other users 102 to 104. In the optimal viewpoint sharing mode, the HMD of each user 101 to 104 displays an image in which a virtual object 105 viewed from a preset optimal viewpoint position and direction is superimposed on a real space image captured by the camera of the HMD of each user.
[0034] 6 is a flowchart showing the flow of a series of information processing by the information processing device 201. The processing shown in this flowchart is realized by the CPU 202 executing the information processing program according to this embodiment.
[0035] First, in step S401, the user position and orientation recording unit 301 analyzes a real-space image acquired by a camera (image capturing device) mounted on the HMD of each of the users 101 to 104 using an existing self-position estimation technology to acquire position and orientation information for each user. In this embodiment, the user's position and orientation information refers to the position and orientation of the HMD, which is information capable of representing the viewpoint position and orientation of the user. An example of a self-position estimation method is SLAM (Simultaneous Localization and Mapping). As long as the user's position and orientation in real space (the position and orientation of the HMD) can be acquired, the method is not limited to SLAM. For example, the user's position and orientation may be acquired using a GPS, a gyro sensor, a geomagnetic sensor, or the like. The user position and orientation recording unit 301 saves the information acquired in step S401 as user position and orientation information 302.
[0036] Next, in step S402, the user position and orientation recording unit 301 determines whether or not the position and orientation information of all users present in the mixed reality space has been acquired. If the user position and orientation recording unit 301 has not yet acquired the position and orientation information (viewpoint position, orientation) of all users, the process returns to step S401, and the information processing device 201 acquires the position and orientation information of any users that have not yet been acquired. On the other hand, if the acquisition has been completed, the process of the information processing device 201 proceeds to step S403.
[0037] In step S403, the object position and orientation recording unit 303 acquires position and orientation information (viewpoint position and orientation) of the virtual object. In this embodiment, since the specific user 101 is the presenter, a superimposed image in which a virtual object is superimposed on real space captured by a camera of the presenter's HMD is recorded in the recording unit 205. The object position and orientation recording unit 303 reads the superimposed image from the recording unit 205 to the RAM 204, and acquires position and orientation information in the virtual space of a virtual object selected by the presenter from the superimposed image via the operation unit 210. The object position and orientation recording unit 303 then saves the acquired information as object position and orientation information 304.
[0038] Next, in step S404, the object instruction acquisition unit 305 acquires an instruction for the virtual object from the specific user 101 who is the presenter, based on an operation input via the operation unit 210. Note that the instruction from the presenter may be an audio instruction instead of an instruction via the operation unit 210.
[0039] Next, in step S405, the mode determination unit 306 determines which display mode, the presenter viewpoint sharing mode or the optimal viewpoint sharing mode, to apply to the virtual object received from the object instruction acquisition unit 305.
[0040] Next, in step S406, the viewpoint determination unit 307 determines the viewpoint position and orientation of each user with respect to the virtual object, based on the user position and orientation information 302, the object position and orientation information 304, and the display mode determined by the mode determination unit 306. The display viewpoint determination process in the viewpoint determination unit 307 will be described in detail later using the flowchart in FIG.
[0041] Next, in step S407, the image generation unit 308 generates an image in which the virtual object is superimposed on the real space image, in accordance with the viewpoint position and direction determined by the viewpoint determination unit 307, the user position and orientation information 302, and the object position and orientation information 304. The superimposed image generation process by the image generation unit 308 will be described in detail later using the flowchart in FIG. Then, in the next step S408, the display unit 211 displays an image in which the virtual object is superimposed on the real space image on the HMD of each of the users 101 to 104.
[0042] 7 is a flowchart showing the flow of the display viewpoint determination process performed by the viewpoint determination unit 307 in step S406 and the superimposed image generation process performed by the image generation unit 308 in step S407. Steps S501 to S506 correspond to the process of step S406, and steps S507 and S508 correspond to the process of step S407.
[0043] First, in step S501, the viewpoint determination unit 307 determines whether the display mode determined in step S405 is the presenter viewpoint sharing mode or the optimal viewpoint sharing mode. If the viewpoint determination unit 307 determines in step S501 that the display mode is the presenter viewpoint sharing mode, the process proceeds to step S502. On the other hand, if the viewpoint determination unit 307 determines that the display mode is the optimal viewpoint sharing mode, the process proceeds to step S503.
[0044] If the process proceeds to step S502, the viewpoint determination unit 307 acquires the viewpoint position and direction of the specific user 101 who is the presenter, based on the user position and orientation information 302 and the object position and orientation information 304. Thereafter, the viewpoint determination unit 307 proceeds to step S504.
[0045] In step S504, the viewpoint determination unit 307 determines whether the line of sight of the specific user 101 who is the presenter has deviated from the target virtual object. If it is determined that the presenter's line of sight has not deviated from the target virtual object, the viewpoint determination unit 307 proceeds to step S506. In step S506, the viewpoint determination unit 307 shares the current position and direction of the presenter's viewpoint with the HMDs of the other users 102 to 104.
[0046] On the other hand, if it is determined in step S504 that the presenter's line of sight has deviated from the target virtual object, the viewpoint determination unit 307 proceeds to the process of step S505. For example, when a presenter is giving an explanation while pointing at the headlight 106 of the virtual object 105 with a pointer or the like, the presenter may look away from the virtual object to read the explanation. If the presenter's viewpoint position and direction at this time are shared with other users, the gazes of the other users 102 to 104 will also deviate from the virtual object. In other words, the displays on the HMDs of the other users 102 to 104 will also deviate from the virtual object. Therefore, if it is determined that the presenter's line of sight has deviated from the target virtual object, the viewpoint determination unit 307 proceeds to the process of step S505.
[0047] In step S505, the viewpoint determination unit 307 acquires the viewpoint position and orientation at the time when the presenter was looking at the virtual object, and the process proceeds to step S506. That is, in this case, in step S506, the viewpoint determination unit 307 shares the viewpoint position and orientation at the time when the presenter was looking at the virtual object with the HMDs of the other users 102 to 104. Note that in step S505, the viewpoint determination unit 307 may share the viewpoint position and orientation previously set for the virtual object with the HMDs of the other users 102 to 104.
[0048] On the other hand, if the process proceeds from step S501 to step S503, the viewpoint determination unit 307 acquires the optimum viewpoint position and orientation that have been set in advance for the virtual object, and the process proceeds to step S506. Then, in step S506, the viewpoint determination unit 307 shares the optimum viewpoint position and orientation with the HMDs of the other users 102 to 104.
[0049] After step S506, the process proceeds to step S507 by the image generation unit 308. In step S507, the image generation unit 308 generates a superimposed image based on the viewpoint position and orientation shared by each user in step S506. The superimposed image may be generated on the HMD, or a superimposed image generated by another configuration may be transmitted to the HMD.
[0050] Next, in step S508, the information processing device 201 determines whether the presentation by the presenter, who is the specific user 101, has finished. If the presentation has finished, the image generation unit 308 transmits an image, a message, or the like indicating that the presentation has finished to the display unit 211. On the other hand, if the presentation has not finished, the information processing device 201 repeats the processes from step S501 to step S507 until the presentation finishes. Note that while the processes from step S501 to S507 in FIG. 5 are being repeated during the presentation, the display mode can also be switched as desired.
[0051] <Second embodiment> In the first embodiment, a superimposed image generated by sharing a specific user's viewpoint position and direction relative to a virtual object, or a preset optimal viewpoint position and direction relative to a virtual object, is displayed on each user's HMD. Here, when multiple users surround and view a virtual object in an MR space, each user may want to view the virtual object from a different position and hold a discussion. Therefore, in the second embodiment, a different viewpoint position and direction relative to a virtual object are calculated for each user, and superimposed images generated according to the different viewpoint positions and orientations are displayed on each user's HMD.
[0052] FIG. 8 is a diagram showing an example of calculating different viewpoint positions and directions of other users 102 to 104 based on the viewpoint position and direction of a specific user who is a presenter. 8(a) is a diagram showing the viewpoint positions and orientations of users 801 to 804 in a virtual space 800. In this example, a specific user is user 802, and the viewpoint position and orientation of user 802 relative to a virtual object 105 in the virtual space 800 are used as a reference. In the virtual space 800, the users 801 to 804 are arranged in a predetermined configuration, such as side-by-side or in a fan shape, and the viewpoint positions and orientations of the other users 801, 803, and 804, who are arranged in the predetermined configuration, are calculated based on the reference viewpoint position and orientation of the specific user 802. Note that while this example illustrates an example in which the viewpoint position and orientation of the specific user 802 are used as a reference, the reference viewpoint position and orientation may be the viewpoint position and orientation of an arbitrarily selected user, or may be the viewpoint position and orientation preset for the virtual object 105.
[0053] 8(b) is a diagram showing the users 101-104 in the real space 100 and the virtual object 105 in the virtual space in the mixed reality space of the second embodiment. Images 805-808 in FIG. 8(b) represent superimposed images displayed on the HMDs of the users 101-104 in the second embodiment.
[0054] In the second embodiment, the HMDs of the users 101 to 104 display an image in which a virtual object viewed from the viewpoint positions and orientations of the users 801 to 804 arranged in the virtual space 800 shown in Fig. 8(a) is superimposed on an image of the real space captured by each camera. That is, in the second embodiment, each of the users 101 to 104 can hold a discussion while remaining in the same position in the real space 100, as if they were viewing the virtual object 105 from different positions like the users 801 to 804 in the virtual space 800. Furthermore, in this embodiment, each of the users 101 to 104 can independently perform various operations such as zooming in and out and rotating the superimposed image displayed on the HMD via the operation unit 210.
[0055] 9 is a diagram showing an example in which avatars 902, 903, and 904 of users (e.g., users 802, 803, and 804) in the virtual space 800 of FIG. 8(a) are superimposed on an image in real space in the second embodiment. Image 900 represents an image displayed on the HMD of user 104, for example.
[0056] At this time, the image generation unit 308 generates an image in which the avatars 902 to 904 of the other users are positioned and superimposed on the image displayed on the HMD of the user 104, based on the viewpoint positions and orientations of the users 801 to 804 in the virtual space 800. This allows the users 101 to 104 to feel as if the other users are right next to them in the virtual space, discussing the virtual object 105, even if they are located far away from each other in the real space.
[0057] Such avatars may be superimposed on images generated by the viewpoint determination unit 307 in the first embodiment so as to correspond to the viewpoints determined for the users 101 to 104. In this case, the avatars are generated by the image generation unit 308 in the first embodiment.
[0058] Fig. 10 is a flowchart showing the flow of the display viewpoint determination process performed by the viewpoint determination unit 307 in step S406 in Fig. 6 and the superimposed image generation process performed by the image generation unit 308 in step S407 in the second embodiment. Steps S1001 to S1003, like steps S501 to S503 in Fig. 7, are processes for determining whether or not the presenter's viewpoint position and direction are to be shared with other users, and acquiring the viewpoint position and direction according to the determination result.
[0059] In the second embodiment, in S1002, the viewpoint determination unit 307 acquires the viewpoint position and direction of the presenter (specific user 101) from the user position and orientation information 302 and the object position and orientation information 304, and then proceeds to step S1004.
[0060] When the process proceeds to step S1004, the viewpoint determination unit 307 calculates the viewpoint position and orientation of each user in the virtual space from the viewpoint position and orientation acquired in step S1002 or step S1003. In the second embodiment, the viewpoint determination unit 307 calculates the viewpoint position and orientation of each user when the viewpoint positions and orientations of each user are arranged in a predetermined manner, such as in a fan shape or side by side, with respect to the target virtual object. This allows each user to view the virtual object from a different angle.
[0061] After step S1004, the process proceeds to step S1005, in which the viewpoint determination unit 307 shares the viewpoint position and orientation relative to the virtual object in the virtual space of each user calculated in step S1004 with the HMD of each corresponding user.
[0062] Next, in step S1006, the image generation unit 308 generates an image based on the viewpoint position and orientation shared by each user, similar to step S506 in the first embodiment. The generation of an image in which a virtual object is superimposed on a real-space image may be performed on the HMD, or an image generated by another configuration may be transmitted to the HMD. Furthermore, when generating an image in each user's HMD, as described above with reference to FIG. 9, avatars may be generated from the viewpoint positions of other users in the virtual space and simultaneously placed within the field of view. This can improve the sense of realism during discussions between users.
[0063] After step S1006, the process proceeds to step S1007, where the information processing device 201 determines whether the presentation has ended, as in step S508. If the presentation has ended, the image generation unit 308 transmits a notification that the presentation has ended to the display unit 211. On the other hand, if the presentation has not ended, the information processing device 201 repeats the processes from step S1001 to step S1006 until the presentation ends. In the second embodiment, too, while the processes from step S1001 to S1006 are being repeated, the display mode can be switched arbitrarily.
[0064] As described above, in the information processing device 201 of the first and second embodiments, the viewpoint position and orientation of a virtual object of another user can be changed and shared in response to an instruction for the virtual object from a specific user such as a presenter. Furthermore, in the information processing device 201 of the first and second embodiments, a user can arbitrarily switch the display mode for a virtual object, so that a specific user can clearly share with other users a part that the specific user wants to focus on.
[0065] In the above-described embodiment, an example of a mixed reality space was presented in which multiple users gather in the same real space to view a virtual object. However, the multiple users may be located in different real spaces. In this case, the HMDs of the users or personal computers connected to the HMDs are connected via the Internet, and images and audio captured by the HMD cameras are exchanged via a server, enabling communication between the users. The HMDs of users in different real spaces display images in which virtual objects are superimposed on images captured of their respective real spaces. This allows multiple users in different locations to view virtual objects over a network. Furthermore, the viewpoint position and orientation of other users relative to the virtual object can be shared depending on the display mode (presenter viewpoint sharing mode or optimal viewpoint sharing mode) and the user's instruction to the virtual object.
[0066] In addition, in the above-described embodiment, an example was given in which each user wears an HMD and views an image in which a virtual object is superimposed on real space, but the mobile device used by each user is not limited to an HMD and may be a smartphone, tablet device, etc.
[0067] The present invention can also be realized by supplying a program that realizes one or more of the functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more of the functions. The above-described embodiments are merely examples of specific implementations of the present invention, and the technical scope of the present invention should not be construed as being limited by these embodiments. In other words, the present invention can be implemented in various forms without departing from its technical concept or main features. [Explanation of symbols]
[0068] 201: Information processing device, 301: User position and orientation recording unit, 303: Object position and orientation recording unit, 304: Object instruction acquisition unit, 306: Mode determination unit, 307: Viewpoint determination unit, 308: Image generation unit
Claims
1. an instruction acquisition means for acquiring an instruction from a first user wearing the first head mounted display; a mode determination means for determining a display mode of a second head mounted display worn by a second user in response to the instruction; a viewpoint determination means for determining a viewpoint for a virtual object in accordance with the determined display mode; an image generation means for generating an image of a mixed reality space; and the mode determination means determines whether to use a first display mode in which a viewpoint of a specific user of the first user and the second user with respect to the virtual object is determined as a viewpoint of another user other than the specific user with respect to the virtual object, or to use a second display mode in which a viewpoint set in advance with respect to the virtual object is determined as a viewpoint of the other user with respect to the virtual object; The information processing device is characterized in that the image generation means generates an image of the mixed reality space to be displayed on the second head-mounted display, by superimposing an image of the virtual object seen from the determined viewpoint different from the viewpoint of the second user on an image of the real space obtained by the second head-mounted display.
2. a position and orientation acquisition means for acquiring a position and orientation of the virtual object and a position and orientation of each of the first user and the second user in the mixed reality space; the viewpoint determination means determines a viewpoint for the virtual object according to the display mode, based on a position and orientation of the virtual object, the positions and orientations of the first user and the second user, and the instruction for the virtual object; The information processing device according to claim 1, characterized in that the image generation means generates an image of the mixed reality space by superimposing an image of the virtual object seen from the determined viewpoint on an image of the real space obtained for each of the first user and the second user.
3. 2. The information processing device according to claim 1, wherein, when the first display mode is determined by the mode determination means, the viewpoint determination means determines the position and direction of the viewpoint of the particular user with respect to the virtual object as the position and direction of the viewpoint of the other user with respect to the virtual object.
4. 4. The information processing device according to claim 3, wherein, when the line of sight of the specific user deviates from the virtual object, the viewpoint determination means determines the position and direction of the viewpoint of the specific user when the specific user was pointing at the virtual object as the position and direction of the viewpoint with respect to the virtual object.
5. 4. The information processing device according to claim 3, wherein the viewpoint determination means determines a preset position and direction of the viewpoint relative to the virtual object when the line of sight of the specific user deviates from the virtual object.
6. 2. The information processing device according to claim 1, wherein, when the second display mode is determined by the mode determination means, the viewpoint determination means determines a position and orientation of a viewpoint that is preset with respect to the virtual object as the position and orientation of a viewpoint with respect to the virtual object for at least the other user.
7. 7. The information processing apparatus according to claim 1, wherein the mode determination means determines the first display mode and the second display mode by switching between them in response to an instruction from the specific user.
8. 8. The information processing device according to claim 1, wherein the mode determination means determines whether the first display mode or the second display mode is to be used for each of the first user and the second user.
9. 2 . The information processing apparatus according to claim 1 , wherein the viewpoint determining means determines, as the viewpoint for the virtual object, a position and a direction of a viewpoint that is disposed in a predetermined position with respect to the virtual object.
10. 10. The information processing apparatus according to claim 9, wherein the viewpoint determination means determines, based on a position and a direction of a reference viewpoint, the positions and directions of a plurality of viewpoints that are arranged in a predetermined arrangement with respect to the virtual object as the viewpoints with respect to the virtual object.
11. 11. The information processing apparatus according to claim 10, wherein the reference viewpoint position and direction are the viewpoint position and direction of a user who is used as a reference, or a predetermined viewpoint position and direction.
12. The information processing device according to any one of claims 1 to 11, characterized in that the image generation means generates an image in which avatars of each of the first user and the second user are superimposed on an image of the mixed reality space based on the position and orientation of the viewpoints of each of the first user and the second user in the mixed reality space.
13. 13. The information processing apparatus according to claim 1, further comprising an operation means for each of the first user and the second user to perform an independent operation on the virtual object.
14. an instruction acquisition step of acquiring an instruction from a first user wearing the first head mounted display; a mode determination step of determining a display mode of a second head mounted display worn by a second user in response to the instruction; a viewpoint determination step of determining a viewpoint for a virtual object in accordance with the determined display mode; an image generation step of generating an image of a mixed reality space; and In the mode determination step, it is determined whether to use a first display mode in which a viewpoint of a specific user of the first user and the second user with respect to the virtual object is determined as a viewpoint of another user other than the specific user with respect to the virtual object, or to use a second display mode in which a viewpoint set in advance with respect to the virtual object is determined as a viewpoint of the other user with respect to the virtual object; An information processing method characterized in that in the image generation process, an image of the mixed reality space to be displayed on the second head-mounted display is generated by superimposing an image of the virtual object seen from the determined viewpoint different from the viewpoint of the second user onto an image of the real space obtained by the second head-mounted display.
15. A program that causes a computer to function as the information processing device according to any one of claims 1 to 13.
Citation Information
Patent Citations
Information processing device, information processing system, control method and program thereof
JP2017033299A
Information processing system, control method thereof, program, information processor, control method thereof, and program
JP2017120650A
Image display device and control method of image display device
JP2020074066A
Transmission terminal, reception terminal, transmission / reception system, and program thereof
JP2020162136A
Information processing apparatus, information processing system, information processing method, and program
JP2021099816A