Information processing apparatus, server, control method of information processing apparatus, and program

The information processing device improves visibility of focused virtual objects in VR or MR conferences by adjusting the transparency of overlapping objects based on acquired information, addressing occlusion issues while maintaining immersion.

JP2025177342APending Publication Date: 2025-12-05CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024084082
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-23
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

In VR or MR conferences, when a specific virtual object is occluded by another opaque virtual object, it becomes difficult or impossible to focus on the specific virtual object, and displaying multiple virtual objects transparently or semi-transparently can damage the sense of immersion.

Method used

An information processing device that processes spatial images with avatars, capable of generating and acquiring information about focus and obstructing virtual objects, and performs transparency processing to differentiate their transparency based on acquired information, improving visibility of the focus virtual object.

Benefits of technology

Enhances the visibility of desired virtual objects even when they overlap, maintaining immersion in VR or MR conferences by adjusting the transparency of obstructing objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025177342000001_ABST
    Figure 2025177342000001_ABST
Patent Text Reader

Abstract

To provide an information processing apparatus, a server, a control method of the information processing apparatus, and a program which allow for improving visibility of a desired virtual object out of a plurality of virtual objects even in a state where the virtual objects overlap each other.SOLUTION: An HMD 1000 processes information about a space image including at least a virtual space in which avatars of a plurality of participants can participant. The HMD 1000 comprises generation means (GPU 1009) for generating virtual objects in the space image and acquisition means (CPU 1001) for acquiring first information about a virtual object of interest out of the virtual objects and second information about a screening virtual object screening the virtual object of interest in the space image. On the basis of the first information and the second information, the generation means performs transparency processing which makes transparency levels of the virtual object of interest and the screening virtual object different from each other.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, a server, a control method for an information processing device, and a program. [Background technology]

[0002] In recent years, there has been active development of technologies related to mixed reality (MR), which combines real space and virtual space. One known example of this technology is a technology that uses a video see-through head mounted display (HMD). By using an HMD, it is possible to display a mixed reality image in which a computer graphics (CG) image generated according to the position and orientation of an imaging device is superimposed as a virtual object on an image of real space captured by an imaging device. MR is also a technology included in XR (Extended Reality / Cross Reality). Another technology included in XR, like MR, is technology related to virtual reality (VR). HMDs can also be used in virtual reality.

[0003] Also, VR conferences and MR conferences are known in which multiple users (participants) wearing HMDs share a single virtual space. In a VR conference or MR conference, for example, when one of the multiple users wants to gaze at a specific virtual object, the virtual object may be occluded by another virtual object, preventing the user from fully gazing at the object. For example, Patent Document 1 discloses a system that prevents a virtual object (CG model) that a user (participant) wants to gaze at from being occluded by the movement of an avatar of another user (co-participant). Furthermore, Patent Document 2 discloses a system that switches to enhance the visibility of a virtual object to be viewed in response to user actions such as eye movement. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Publication No. 2018-106297 [Patent Document 2] Patent No. 7333051 Summary of the Invention [Problem to be solved by the invention]

[0005] However, if a specific virtual object that a user wants to focus on is completely occluded by another opaque (non-transparent) virtual object, it becomes difficult or impossible to focus on the specific virtual object. Also, in VR or MR conferences, multiple virtual objects may be displayed, and displaying each virtual object transparently or semi-transparently may damage the sense of immersion in the conference.

[0006] The present invention has been made in consideration of the above-mentioned problems. It is an object of the present invention to provide an information processing device that can improve the visibility of a desired virtual object among a plurality of virtual objects even when the plurality of virtual objects overlap. Similarly, it is an object of the present invention to provide a server, a control method for an information processing device, and a program that can improve the visibility of a desired virtual object among a plurality of virtual objects even when the plurality of virtual objects overlap. [Means for solving the problem]

[0007] In order to achieve the above object, an information processing device of the present invention is an information processing device that processes information relating to a spatial image that includes at least a virtual space in which avatars of multiple participants can participate, and is equipped with a generation means for generating a virtual object within the spatial image, and an acquisition means for acquiring first information relating to a focus virtual object among the virtual objects that is determined to be the focus of at least some of the multiple participants, and second information relating to an obstructing virtual object that obstructs the focus virtual object in the spatial image viewed by a predetermined participant among the multiple participants, wherein the generation means performs transparency processing that makes the transparency of the focus virtual object and the obstructing virtual object different from each other based on the first information and the second information. [Effects of the Invention]

[0008] According to the present invention, even when a plurality of virtual objects overlap each other, it is possible to improve the visibility of a desired virtual object among the plurality of virtual objects. [Brief explanation of the drawings]

[0009] [Figure 1] 2 is a block diagram showing an example of the hardware configuration of an HMD and a server according to the first embodiment. FIG. [Figure 2] FIG. 1 is a diagram showing an example of a configuration in real space of a communication system having a plurality of HMDs and a server. [Figure 3] FIG. 3 is a bird's-eye view showing an example of a virtual space conference using the communication system shown in FIG. 2. [Figure 4] 10 is a flowchart showing a process executed by the server. [Figure 5] FIG. 10 is a bird's-eye view showing an example of a virtual space conference established by a server. [Figure 6] 1A and 1B are a top view and a front view seen from the avatar's point of view showing an example of a virtual space conference created using an HMD. [Figure 7] 10 is a flowchart showing processing executed by the HMD. [Figure 8] 1A and 1B are a top view and a front view seen from the avatar's point of view showing an example of a virtual space conference created using an HMD. [Figure 9] FIG. 10 is a bird's-eye view showing an example of a virtual space conference established by an HMD according to a second embodiment. [Figure 10] FIG. 11 is a bird's-eye view showing an example of a virtual space conference established by an HMD according to a third embodiment. [Figure 11] FIG. 10 is a bird's-eye view showing an example of a virtual space conference established by an HMD according to a fourth embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0010] Each embodiment of the present invention will be described in detail below with reference to the drawings. However, the configurations described in each of the following embodiments are merely examples, and the scope of the present invention is not limited to the configurations described in each embodiment. For example, each component constituting the present invention can be replaced with any configuration that can perform the same function. Also, any component may be added. Furthermore, any two or more configurations (features) of each embodiment can be combined.

[0011] First Embodiment A first embodiment will be described below with reference to FIGS. 1 to 8. FIG. 1 is a block diagram showing an example of the hardware configuration of an HMD and a server according to the first embodiment. An HMD (head mounted display) 1000 shown in FIG. 1 is a device to which an information processing device is applied, and is communicably connected to a server 1200 via the Internet 1300. The HMD 1000 includes a CPU 1001, a ROM 1002, a RAM 1003, and a mass storage unit 1004. The HMD 1000 also includes a sensor unit 1005, a communication unit 1006, a gaze measurement unit (detection means) 1007, an imaging unit 1008, a GPU (generation means) 1009, and a display unit (display means) 1010. These pieces of hardware included in the HMD 1000 are communicably connected to each other via an internal bus 1011.

[0012] The CPU 1001 is a computer that executes programs stored in the ROM 1002 and programs deployed in the RAM 1003. These programs include, for example, programs for causing the CPU 1001 to execute each unit and each means (control method of an information processing device) of the HMD 1000. The ROM 1002 is a non-volatile memory configured, for example, by a flash memory. The ROM 1002 holds a program for controlling the HMD 1000. When the HMD 1000 is powered on, the CPU 1001 reads the program from the ROM 1002 and starts controlling the HMD 1000. The RAM 1003 is a rewritable memory configured, for example, by a volatile memory (DRAM) using a semiconductor element. The RAM 1003 is also used as a work area for the CPU 1001. In addition, programs stored in the ROM 1002 or the mass storage unit 1004 are deployed in the RAM 1003 and executed by the CPU 1001. The RAM 1003 also stores data acquired by, for example, the sensor unit 1005, the communication unit 1006, the gaze measurement unit 1007, and the imaging unit 1008. The mass storage unit 1004 is a rewritable and non-volatile storage unit configured by, for example, a semiconductor storage device such as an eMMC (embedded multi media card) or an SSD (solid state disk). The mass storage unit 1004 stores programs for controlling the HMD 1000, various data used by the programs, and data acquired by the sensor unit 1005, the communication unit 1006, the gaze measurement unit 1007, and the imaging unit 1008.

[0013] The communication unit 1006 includes an antenna and a communication controller that controls the antenna. For example, an antenna for realizing short-range wireless communication such as Bluetooth may be used. Other examples of the antenna include antennas for realizing wireless communication conforming to IEEE802.11a, IEEE802.11b, IEEE802.11g, IEEE802.11n, IEEE802.11ac, and IEEE802.11ax. The communication unit 1006 may also be a unit for realizing wired communication conforming to the USB standard. The communication unit 1006 is controlled by the CPU 1001 and communicates with the server 1200 via the Internet 1300. This allows the communication unit 1006 to transmit data to the server 1200 and receive data from the server 1200. Thus, the communication unit 1006 functions as a transmitting unit that transmits data and a receiving unit that receives data. The gaze measurement unit 1007 is configured with an image sensor positioned so as to capture both eyes of a user wearing the HMD 1000. The gaze measurement unit 1007 may also be configured with a myoelectricity sensor that measures the electrical potential of the muscles that move both eyes of the user. The gaze measurement unit 1007 configured in this manner is controlled by the CPU 1001 and measures (detects) the user's gaze direction, convergence angle, etc. The imaging unit 1008 is configured with a stereo camera or the like positioned so as to capture the outside world of the HMD 1000. The imaging unit 1008 is controlled by the CPU 1001 and captures the outside world of the HMD 1000 with the stereo camera to acquire still images or moving images. The GPU 1009 is configured with an LSI or the like designed to perform various calculations for rendering three-dimensional model data at high speed and in parallel. The GPU 1009 performs parallel data processing, rendering of the virtual space, and superimposing an image captured by the imaging unit 1008 onto a virtual object created by CG, to generate a display image that is displayed on the display unit 1010. The virtual object is generated by, for example, the GPU 1009 (generation process). The display image is stored in the RAM 1003 or the mass storage unit 1004 via the VRAM in the GPU 1009 or the internal bus 1011.The display unit 1010 is controlled by the CPU 1001 or the GPU 1009, and has a display device for the user's left eye and a display device for the user's right eye. This allows different images to be displayed taking into account parallax, thereby enabling stereoscopic viewing. The display device is not particularly limited, and may be, for example, a liquid crystal panel or an organic EL panel.

[0014] As shown in FIG. 1, the server 1200 includes a CPU 1201, a ROM 1202, a RAM 1203, a mass storage unit 1204, and a communication unit 1205, which are communicatively connected to one another via an internal bus 1206. The CPU 1201 is a computer that executes programs stored in the ROM 1202 and programs deployed in the RAM 1203. The ROM 1202 is a non-volatile memory formed, for example, of a flash memory. The ROM 1202 holds a program for controlling the server 1200. When the server 1200 is powered on, the CPU 1201 reads the program from the ROM 1202 and begins controlling the server 1200. The RAM 1203 is a rewritable memory formed, for example, of a volatile memory (DRAM) using a semiconductor element. The RAM 1203 is also used as a work area for the CPU 1201. In addition, programs stored in the ROM 1202 or the mass storage unit 1204 are deployed in the RAM 1203 and executed by the CPU 1201. The RAM 1203 also stores data acquired by the communication unit 1205. The mass storage unit 1204 is a rewritable, non-volatile storage unit configured, for example, by a semiconductor storage device such as an eMMC or an SSD. The mass storage unit 1204 stores a program for controlling the server 1200, various data used by the program, and data acquired by the communication unit 1006. The communication unit 1205 is a communication unit for realizing wired communication. The communication unit 1205 is composed of hardware for wired communication and a communication controller for processing wired signals, and realizes wired communication in accordance with the IEEE 802.3 standard. The communication unit 1205 is controlled by the CPU 1201 and communicates with the HMD 1000 via the Internet 1300. This allows the communication unit 1205 to transmit data to the HMD 1000 and receive data from the HMD 1000. In this way, the communication unit 1205 functions as a transmitting means for transmitting data and as a receiving means for receiving data.

[0015] FIG. 2 is a diagram showing an example of the configuration of a communication system in real space having a plurality of HMDs and a server. The communication system 2000 shown in FIG. 2 has six HMDs 1000 and a server 1200, which are connected to each other via the Internet 1300 so as to be able to communicate with each other. Note that the number of HMDs 1000 arranged in the communication system 2000 is six in this embodiment, but is not limited to this and may be, for example, two to five, or seven or more. Each HMD 1000 is used by a user 2400 to a user 2405. The communication system 2000 is used in a virtual space conference (see FIG. 3). Each user 2400 to a user 2405 can participate in this virtual space conference as a participant. Note that in this embodiment, the users 2400 to a user 2405 are located at different locations in real space, but are not limited to this.

[0016] FIG. 3 is a bird's-eye view showing an example of a virtual space conference using the communication system shown in FIG. 2. A user 2400 can participate in a virtual space conference (virtual space) 300 shown in FIG. 3 as an avatar 3000. Similarly, a user 2401 can participate as an avatar 3001, a user 2402 can participate as an avatar 3002, and a user 2403 can participate as an avatar 3003. A user 2404 can participate as an avatar 3004, and a user 2405 can participate as an avatar 3005. The avatars 3000 to 3005 are each a virtual object. The HMD 1000 can process information related to an image (space image) within the virtual space conference 300. The virtual space conference 300 only needs to include at least a virtual space. For example, the virtual space conference 300 may include a composite space of a virtual space and a real space. A virtual whiteboard 3006 is arranged in the virtual space conference 300 as a virtual object. Avatars 3003 to 3005 are arranged in a line in the front row, facing virtual whiteboard 3006. Of avatars 3003 to 3005, avatar 3004 is arranged in the center, avatar 3003 is arranged to the left of avatar 3004 when facing the front, and avatar 3005 is arranged to the right of avatar 3004 when facing the front. Avatars 3000 to 3002 are arranged in a line in the back row, facing virtual whiteboard 3006. Of avatars 3000 to 3002, avatar 3000 is arranged in the center, avatar 3001 is arranged to the left of avatar 3000 when facing the front, and avatar 3002 is arranged to the right of avatar 3000 when facing the front. Users 2400 to 2405 can view virtual whiteboard 3006 from the eye level of their respective avatars.

[0017] The initial coordinates and initial orientations of the avatars 3000 to 3005 and the virtual whiteboard 3006 in the virtual space conference 300 are stored in the mass storage unit 1204 of the server 1200. At the start of the virtual space conference 300, the CPU 1201 of the server 1200 controls the communication unit 1205 to transmit the initial coordinates and initial orientations stored in the mass storage unit 1204 to each HMD 1000. The CPU 1001 and GPU 1009 of each HMD 1000 perform rendering processing using the initial coordinates and initial orientations received by the communication unit 1006. In this way, the virtual space conference 300 shown in FIG. 3 is constructed. Furthermore, by synchronizing the coordinates and orientation information of the virtual objects in each HMD 1000, it is possible to prevent relatively large positional deviations between virtual objects in the virtual space conference 300.

[0018] Fig. 4 is a flowchart showing processing executed by the server. In the flowchart shown in Fig. 4, steps S4001 to S4006 are periodically repeated and executed in order. In step S4001, the CPU 1201 of the server 1200 controls the communication unit 1205 to receive information about the movement amount and orientation of the HMD 1000 transmitted from each HMD 1000 in step S7003 (see Fig. 7), which will be described later. This information is stored in the RAM 1203 of the server 1200.

[0019] In step S4002, the CPU 1201 controls the communication unit 1205 to receive the user's line of sight information transmitted from each HMD 1000 in step S7004 (see FIG. 7), which will be described later. This line of sight information is stored in the RAM 1203. Note that the order of execution of steps S4001 and S4002 may be reversed.

[0020] In step S4003, based on the movement amount and orientation information of each HMD 1000 stored in step S4001, the CPU 1201 calculates the movement amount, coordinates, and orientation of the avatar of the user using each HMD 1000 within the virtual space conference 300. The calculation results are stored in the RAM 1203.

[0021] In step S4004, the CPU 1201 identifies at least one gaze target candidate object (candidate virtual object) from among the virtual objects in the virtual space conference 300, based on the gaze information of each user stored in step S4002. Information about this gaze target candidate object is stored in the RAM 1203.

[0022] In step S4005, the CPU 1201 controls the communication unit 1205 to transmit the calculation results stored in the RAM 1203 in step S4003 to each HMD 1000.

[0023] In step S4006, the CPU 1201 controls the communication unit 1205 to transmit the information about the gaze target candidate object stored in the RAM 1203 in step S4004 to each HMD 1000. Note that the order of execution of steps S4005 and S4006 may be reversed.

[0024] Fig. 5 is a bird's-eye view showing an example of a virtual space conference established by a server. As shown in Fig. 5, avatar 3000 has a line of sight α5000 directed toward virtual whiteboard 3006 and is gazing at point of gaze P5000 on virtual whiteboard 3006. Similarly, avatar 3001 has a line of sight α5001 directed toward virtual whiteboard 3006 and is gazing at point of gaze P5001 on virtual whiteboard 3006. Furthermore, avatar 3002 has a line of sight α5002 directed toward virtual whiteboard 3006 and is gazing at point of gaze P5002 on virtual whiteboard 3006. Avatar 3003 has a line of sight α5003 directed toward virtual whiteboard 3006 and is gazing at point of gaze P5003 on virtual whiteboard 3006. Avatar 3004 has a line of sight α5004 directed toward virtual whiteboard 3006, and is gazing at point of gaze P5004 on virtual whiteboard 3006. Avatar 3005 has a line of sight α5005 directed toward virtual whiteboard 3006, and is gazing at point of gaze P5005 on virtual whiteboard 3006. The line of sight direction of each avatar is calculated by CPU 1201 of server 1200 based on the line of sight information of each user stored in RAM 1203 in step S4002 and the calculation result stored in RAM 1203 in step S4003.

[0025] The CPU 1201 identifies a virtual object of interest based on the line of sight of each avatar. The "virtual object of interest" identified as a result of this identification refers to a virtual object within the virtual space conference 300 that at least some of the users among the multiple users are focusing on through their avatars, i.e., that the users are concentrating on and viewing. Note that at least one virtual object of interest is sufficient. The method for identifying the virtual object of interest is not particularly limited. For example, first, virtual objects present in the line of sight of each avatar are identified. Next, virtual objects that have been identified a threshold or more times are identified as virtual objects on which users are concentrating their gazes. The threshold is a ratio of the total number of participants in the virtual space conference 300, and can be set to, for example, 1 / 2, 1 / 3, 1 / 4, or more. Note that the threshold may be determined arbitrarily by the organizer of the virtual space conference 300, or may be determined by each HMD 1000 or the server 1200 of the communication system 2000. In this embodiment, the threshold is set to ⅓ or more of the total number of participants, and the virtual whiteboard 3006 on which all of the avatars' line of sight are focused is identified as the virtual object of interest. In this way, in this embodiment, the CPU 1201 also functions as an identifying unit that identifies the virtual object of interest. Note that a part that functions as the identifying unit may be provided separately from the CPU 1201. Furthermore, although the identifying unit is provided in the server 1200 in this embodiment, the present invention is not limited to this and may be provided in the HMD 1000, for example.

[0026] FIG. 6 shows an overhead view and a front view from the avatar's point of view, illustrating an example of a virtual space conference created using an HMD. The virtual space conference 6000 shown in FIGS. 6(a) and 6(b) is created using an HMD 1000 worn by a user 2400. FIG. 6(a) is an overhead view of the virtual space conference 6000, focusing on the avatar 3000 among the avatars 3000 to 3005. The arrangement of each virtual object, such as the avatar 3000, in this virtual space conference 6000 is the same as the arrangement of each virtual object, such as the avatar 3000, in the virtual space conference 300. As described above, the avatar 3000's line of sight α5000 is directed toward the virtual whiteboard 3006, and the avatar 3000 is gazing at a gaze point P5000 on the virtual whiteboard 3006. FIG. 6(b) is a view of the virtual space conference 6000 viewed from the front of the avatar 3000 of the user 2400. The virtual space conference 6000 shown in FIG. 6(b) is an image displayed on the display unit 1010 of the HMD 1000. As shown in FIG. 6(a), avatars 3003 to 3005 are present in front of the avatar 3000, i.e., on the virtual whiteboard 3006 side. In this case, as shown in FIG. 6(b), even if the user 2400 attempts to view the entire virtual whiteboard 3006 from the line of sight of the avatar 3000, the virtual whiteboard 3006 is blocked by the avatars 3003 to 3005. As a result, it becomes difficult for the user 2400 to fully view the entire virtual whiteboard 3006 through the avatar 3000; that is, the visibility of the virtual whiteboard 3006 decreases. This may make it difficult for the user 2400 to grasp the content displayed on the virtual whiteboard 3006.

[0027] Therefore, the HMD 1000 is configured to reduce such a decrease in visibility. The configuration and operation will be described below. FIG. 7 is a flowchart showing processing executed by the HMD. In the flowchart shown in FIG. 7, steps S7001 to S7017 are periodically repeated and executed in order. FIG. 8 is a bird's-eye view and a front view seen from the avatar's eye level showing an example of a virtual space conference created by an HMD. The virtual space conference 8000 shown in FIGS. 8(a) and 8(b) is created by the HMD 1000 worn by the user 2400. FIG. 8(a) is a bird's-eye view of the virtual space conference 8000 when focusing on the avatar 3000 among the avatars 3000 to 3005. The arrangement of each virtual object, such as the avatar 3000, in this virtual space conference 8000 is the same as the arrangement of each virtual object, such as the avatar 3000, in the virtual space conference 6000. In Fig. 8(a), as in the above, the line of sight α5000 of the avatar 3000 is directed toward the virtual whiteboard 3006, and the avatar 3000 is gazing at the gaze point P5000 on the virtual whiteboard 3006. Fig. 8(b) is a diagram of the virtual space conference 8000 as seen from the front of the avatar 3000 of the user 2400. The virtual space conference 8000 shown in Fig. 8(b) is an image displayed on the display unit 1010 of the HMD 1000.

[0028] 7, in step S7001, the CPU 1001 of the HMD 1000 controls the sensor unit 1005 to acquire information related to the amount of movement and the posture of the HMD 1000. This information is stored in the RAM 1003 of the HMD 1000. After step S7001 is executed, the process proceeds to step S7002.

[0029] In step S7002, the CPU 1001 controls the gaze measurement unit 1007 to acquire gaze information such as the gaze direction and convergence angle of the user (predetermined participant) wearing the HMD 1000, i.e., the detection results of the gaze measurement unit 1007. This gaze information is stored in the RAM 1003. After step S7002 is executed, the process proceeds to step S7003. Note that the order of execution of steps S7001 and S7002 may be reversed.

[0030] In step S7003, the CPU 1001 controls the communication unit 1006 to transmit the information stored in the RAM 1003 in step S7001 to the server 1200. After step S7003 is executed, the process proceeds to step S7004.

[0031] In step S7004, the CPU 1001 controls the communication unit 1006 to transmit the gaze information stored in the RAM 1003 in step S7002 to the server 1200. After step S7004 is executed, the process proceeds to step S7005. Note that the order in which steps S7003 and S7004 are executed may be reversed. Alternatively, steps S7001, S7003, S7002, and S7004 may be executed in this order.

[0032] In step S7005, the CPU 1001 controls the communication unit 1006 to receive the calculation result transmitted from the server 1200 in step S4005 (see FIG. 4). This calculation result is stored in the RAM 1003. After step S7005 is executed, the process proceeds to step S7006.

[0033] In step S7006, the CPU 1001 controls the communication unit 1006 to receive the gaze target candidate object information transmitted from the server 1200 in step S4006 (see FIG. 4). The gaze target candidate object information is stored in the RAM 1003. After step S7006 is executed, the process proceeds to step S7007. Note that the order of execution of steps S7005 and S7006 may be reversed.

[0034] In step S7007, the CPU 1001 reflects the calculation result stored in the RAM 1003 in step S7005 in the movement amount, coordinates, and posture of each avatar in the virtual space conference 8000 (see FIG. 8) created by the HMD 1000. As described above, the communication system 2000 executes steps S7001, S7003, S4001, S4003, S4005, S7005, and S7007. This synchronizes the movement, coordinates, and posture of virtual objects in the virtual space conference 8000 created by each HMD 1000. After executing step S7007, the process proceeds to step S7008.

[0035] In step S7008, CPU 1001 determines, based on the line-of-sight information stored in step S7002, whether the gaze target object identified in step S7010 is present in the line-of-sight direction α5000 of avatar 3000 (user 2400). As described above, in the flowchart shown in FIG. 7, steps S7001 to S7017 are periodically repeated and executed in order. If the determination in step S7008 indicates that the gaze target object identified in step S7010 in the previous cycle is not present in the line-of-sight direction α5000, the process proceeds to step S7009. On the other hand, if the determination in step S7008 indicates that the gaze target object identified in step S7010 in the previous cycle is present in the line-of-sight direction α5000, the process proceeds to step S7015. If step S7008 is executed for the first time, step S7010 has not been executed and there is no gaze target object, so the process proceeds to step S7009.

[0036] In step S7009, CPU 1001 determines whether or not a gaze target candidate object exists in the gaze direction α5000, based on the gaze information stored in step S7002 and the gaze target candidate object information stored in step S7006. If the determination in step S7009 determines that a gaze target candidate object exists in the gaze direction α5000, the process proceeds to step S7010. On the other hand, if the determination in step S7009 determines that a gaze target candidate object does not exist in the gaze direction α5000, the process proceeds to step S7014. In the state shown in FIG. 8(a), the virtual whiteboard 3006 identified as the gaze target candidate object in step S4004 (see FIG. 4) exists in the gaze direction α5000. In this case, the process proceeds to step S7010.

[0037] In step S7010, CPU 1001 identifies the gaze target candidate object determined in step S7009 to exist in line of sight α5000 as the gaze target object (noted virtual object). If there are multiple gaze target candidate objects, the gaze target candidate object that is located farthest from avatar 3000 is identified as the gaze target object. In the state shown in FIG. 8(a), virtual whiteboard 3006 is identified as the gaze target object. After step S7010 is executed, the process proceeds to step S7011.

[0038] In step S7011, CPU 1001 determines whether one or more occluding virtual objects exist. An "occluding virtual object" refers to a virtual object that occludes virtual whiteboard 3006, which is the gaze target object, in the field of view of user 2400, i.e., on virtual space conference 8000 displayed on display unit 1010. The determination method in step S7011 involves first calculating a virtual space visible in the field of view of avatar 3000 based on the user's line of sight information, virtual object coordinates, and posture information stored in RAM 1003. Next, the CPU 1001 calculates an overlap state in the calculated virtual space between the gaze target object and a virtual object located in front of the gaze target object. If the calculation result is equal to or greater than a threshold, it is determined that an occluding virtual object exists. If the calculation result is less than the threshold, it is determined that an occluding virtual object does not exist. In FIG. 8(b), in a virtual space conference 8000, virtual whiteboard 3006, which is the object of gaze, overlaps avatars 3003 to 3005 located in front of virtual whiteboard 3006 by a threshold value or more. In this case, avatars 3003 to 3005 are determined (identified) as occluding virtual objects. Alternatively, as another determination method in step S7011, for example, a determination method based on the relative coordinates of avatar 3000, the object of gaze, and a virtual object located between avatar 3000 and the object of gaze may be used. Alternatively, as another determination method, a determination method based on the overlapping state of a virtual object rendered so as to overwrite the object of gaze in rendering processing may be used. If it is determined in step S7011 that an occluding virtual object exists, the process proceeds to step S7012. On the other hand, if it is determined in step S7011 that an occluding virtual object does not exist, the process proceeds to step S7014.

[0039] In step S7012, the CPU 1001 performs transparency processing to change the transparency of the occluding virtual object. The transparency processing is processing to improve the visibility of the gaze target object while ensuring a degree of transparency that allows the presence of the occluding virtual object to be visible. The transparency of the occluding virtual object may be fully transparent or semi-transparent. The transparency may be changed by increasing or decreasing the amount of change from the current transparency, but it is preferable to overwrite it with a predetermined transparency value. In FIG. 8(b), the transparency of avatars 3003 to 3005 determined to be occluding virtual objects is changed. After step S7012 is executed, the processing proceeds to step S7013.

[0040] In step S7013, the CPU 1001 sets the transparency of virtual objects other than the occluding virtual object to a default value. This step S7013 is executed to return the transparency of a virtual object that has been released from the determination of being an occluding virtual object to the default value due to, for example, a movement of the user 2400's line of sight or a movement of the occluding object being gazed at. After step S7013 is executed, the process proceeds to step S7014. Note that the order of execution of steps S7012 and S7013 may be reversed.

[0041] In step S7014, the GPU 1009 of the HMD 1000 performs rendering processing using the virtual space information stored in the RAM 1003. As a result of this rendering processing, the display unit 1010 of the HMD 1000 displays a virtual space conference 8000 in the state shown in FIG. 8(b). The user 2400 can view this virtual space conference 8000. As shown in FIG. 8(b), the virtual space conference 8000 includes a virtual whiteboard 3006, which is the gaze target object, and avatars 3003 to 3005, which are located in front of the virtual whiteboard 3006 and have been determined to be occluding virtual objects. The avatars 3003 to 3005 have a transparency that allows the user 2400 to fully view the virtual whiteboard 3006.

[0042] As described above, there are cases where avatars 3003 to 3005 overlap with virtual whiteboard 3006, which is a gaze target object (desired virtual object) located behind avatars 3003 to 3005. In this case, avatars 3003 to 3005 become occluding virtual objects that occlude virtual whiteboard 3006. In HMD 1000, under the control of CPU 1001 (acquisition means), i.e., through steps S7001 to S7006, information about virtual whiteboard 3006 (hereinafter referred to as "first information") is acquired (acquisition process). In addition to the first information, information about avatars 3003 to 3005 (hereinafter referred to as "second information") is also acquired (acquisition process). Specifically, position information and posture information about virtual whiteboard 3006 within virtual space conference 8000 are acquired as the first information. In addition, the first information also includes information about the gaze target candidate object. As the second information, position information and posture information of the avatars 3003 to 3005 within the virtual space conference 8000 is acquired.

[0043] Then, the HMD 1000 performs transparency processing based on this information. The transparency processing makes the transparency of the avatars 3003 to 3005 different from that of the virtual whiteboard 3006. That is, the transparency of the avatars 3003 to 3005 can be made higher than that of the virtual whiteboard 3006. This improves the visibility of the virtual whiteboard 3006 for the user 2400. This allows the user 2400 to understand the content displayed on the virtual whiteboard 3006 (for example, "ABCDEFG" in FIG. 8B). Note that when increasing the transparency higher than that of the gaze target object, it is preferable to be able to change the degree of increase. In this case, for example, the transparency of the obscuring virtual object can be changed depending on the degree of darkness of the gaze target object. Specifically, when the gaze target object is a light color, the transparency of the obscuring virtual object can be made relatively high, and when the gaze target object is a dark color, the transparency of the obscuring virtual object can be made relatively low. Furthermore, a rendering process for highlighting the contour of the gaze target candidate object may be set between steps S7006 to S7014. This allows the user to easily recognize the gaze target candidate object. Furthermore, instead of highlighting, a marker or the like pointing to the gaze target candidate object may be superimposed above the gaze target candidate object.

[0044] 7, step S7015 is executed after step S7008 is executed. In step S7015, CPU 1001 determines whether a virtual object other than the gaze target object is being gazed at, based on the gaze information (particularly, convergence angle information of the eyeballs of user 2400) stored in RAM 1003 in step S7002. If it is determined in step S7015 that a virtual object other than the gaze target object is being gazed at, the process proceeds to step S7016. On the other hand, if it is determined in step S7015 that a virtual object other than the gaze target object is not being gazed at, the process proceeds to step S7011.

[0045] In step S7016, CPU 1001 determines whether the transparency of the virtual object determined to be gazed upon in step S7015 was changed in step S7012, which was executed in the previous iteration of the process. This determination is made, for example, by referencing transparency information of the virtual object stored in RAM 1003 and determining whether the transparency differs from the default transparency. If it is determined in step S7016 that the transparency has been changed, the process proceeds to step S7017. On the other hand, if it is determined in step S7016 that the transparency has not been changed, the process proceeds to step S7014.

[0046] In step S7017, CPU 1001 changes the transparency of the virtual object whose transparency was determined to have been changed in step S7016 to the default transparency. After step S7017 is executed, the process proceeds to step S7014.

[0047] Second Embodiment The second embodiment will be described below with reference to FIG. 9. Differences from the previous embodiment will be mainly described, and similar points will not be described again. This embodiment is similar to the first embodiment except that there are multiple gaze target candidate objects that can become gaze target objects. FIG. 9 is an overhead view showing an example of a virtual space conference established with an HMD according to the second embodiment. FIG. 9 shows states that change over time in the order of FIG. 9(a), FIG. 9(b), and FIG. 9(c). As shown in FIGS. 9(a) to 9(c), a virtual space conference 9000 includes avatars 3000 to 3005, a virtual whiteboard 3006, and a virtual whiteboard 9007. The arrangement of the avatars 3000 to 3005 is the same as in the first embodiment. The virtual whiteboard 3006 is located on the left side as viewed from the front, and the virtual whiteboard 9007 is located on the right side as viewed from the front. The virtual whiteboard 3006 and the virtual whiteboard 9007 are each gaze target candidate objects.

[0048] In the state shown in FIG. 9(a), avatars 3000, 3001, and 3004 are gazing at virtual whiteboard 3006, and avatars 3002, 3003, and 3005 are gazing at virtual whiteboard 9007. In the state shown in FIG. 9(a), avatars 3003 and 3004 are present in front of avatar 3000, i.e., between avatar 3000 and virtual whiteboard 3006. Avatars 3003 and 3004 are transparent due to transparency processing. This allows avatar 3000 to fully view virtual whiteboard 3006. Then, avatar 3000 turns toward virtual whiteboard 9007, resulting in the state shown in FIG. 9(b).

[0049] In the state shown in FIG. 9(b), avatars 3000, 3002, 3003, and 3005 are gazing at virtual whiteboard 9007, while avatars 3001 and 3004 are gazing at virtual whiteboard 3006. In the state shown in FIG. 9(b), avatars 3004 and 3005 are located in front of avatar 3000, i.e., between avatar 3000 and virtual whiteboard 9007. Avatar 3004 remains transparent, but avatar 3005 is opaque because transparency processing has not been performed. In this case, avatar 3005 blocks the virtual whiteboard 9007 from avatar 3000, making it difficult for avatar 3000 to fully view the virtual whiteboard 9007. Therefore, in the HMD 1000 of this embodiment, by performing transparency processing on the avatar 3005, it is possible to make the avatar 3005 transparent as shown in Fig. 9(c). This allows the avatar 3000 to fully view the virtual whiteboard 9007. As described above, even when there are multiple gaze target candidate objects, it is possible to improve the visibility of the gaze target object.

[0050] Third Embodiment The third embodiment will be described below with reference to FIG. 10. Differences from the previous embodiment will be mainly described, and similar points will not be described again. This embodiment is similar to the first embodiment except that the avatars' line of sight differ from each other. FIG. 10 is a bird's-eye view showing an example of a virtual space conference established using an HMD according to the third embodiment. FIG. 10 shows a state that changes over time from FIG. 10(a) to FIG. 10(b). As shown in FIGS. 10(a) and 10(b), the virtual space conference 10000 includes avatars 3000 to 3005 and a virtual whiteboard 3006. Avatars 3003 and 3004 are positioned in front of avatar 3000, i.e., between avatar 3000 and virtual whiteboard 3006. Avatars 3001, 3002, and 3005 are each positioned to the right of avatar 3000 as viewed from the front (virtual whiteboard 3006).

[0051] In the state shown in Fig. 10(a), avatars 3000 to 3005 are gazing at virtual whiteboard 3006. In the state shown in Fig. 10(a), avatars 3003 and 3004, which are in front of avatar 3000, i.e., between avatar 3000 and virtual whiteboard 3006, are made transparent by transparency processing. This allows avatar 3000 to fully view virtual whiteboard 3006. Then, avatars 3002 to 3005 turn toward avatar 3001, resulting in the state shown in Fig. 10(b).

[0052] In the state shown in FIG. 10(b), avatar 3000 is still gazing at virtual whiteboard 3006, but avatars 3002 to 3005 are gazing at avatar 3001. In the state shown in FIG. 10(b), avatars 3003 and 3004, which are in front of avatar 3000, remain transparent. This allows avatar 3000 to continue to have a good view of virtual whiteboard 3006. As described above, even when the line of sight of avatar 3000 and the line of sight of avatars 3002 to 3005 differ from each other, the visibility of the gaze target object for avatar 3000 can be improved.

[0053] <Fourth embodiment> The fourth embodiment will be described below with reference to FIG. 11. Differences from the previous embodiments will be mainly described, and similar points will not be described again. This embodiment is similar to the first embodiment except that the line of sight of the avatars changes. FIG. 11 is a bird's-eye view showing an example of a virtual space conference established with an HMD according to the fourth embodiment. FIG. 11 shows states that change over time in the order of FIG. 11(a), FIG. 11(b), and FIG. 11(c). As shown in FIGS. 11(a) to 11(c), the virtual space conference 11000 includes avatars 3000 to 3003 and a virtual whiteboard 3006. Avatars 3001 to 3003 are arranged in a line in front of avatar 3000, i.e., between avatar 3000 and virtual whiteboard 3006.

[0054] In the state shown in Fig. 11(a), avatars 3000 to 3003 are gazing at virtual whiteboard 3006. In the state shown in Fig. 11(a), avatars 3001 to 3003 that are in front of avatar 3000, i.e., between avatar 3000 and virtual whiteboard 3006, are made transparent by transparency processing. This allows avatar 3000 to fully view virtual whiteboard 3006. Then, avatar 3000 changes its line of sight to avatar 3002 in front of avatar 3000, resulting in the state shown in Fig. 11(b).

[0055] In the state shown in FIG. 11(b), avatars 3001 to 3003 are still gazing at virtual whiteboard 3006, but avatar 3000 is gazing at avatar 3002. In the state shown in FIG. 11(b), the gaze target object for avatar 3000 is changed from virtual whiteboard 3006 to avatar 3002. As a result, for avatar 3000, the overlapping state between avatars 3001 to 3003 and virtual whiteboard 3006 is released. Note that this release is determined based on the convergence angle θ11 of avatar 3000 measured from the rotation of the left and right eyeballs of user 2400. Then, by stopping the transparency processing for avatars 3001 to 3003, the state shown in FIG. 11(c) is achieved. As shown in FIG. 11(c), avatars 3001 to 3003 each return to the state before the transparency processing was performed. As described above, when the line of sight of the avatar 3000 changes and it is acceptable to stop the transparency processing, the processing can be stopped. This makes it possible to prevent unnecessary execution of transparency processing.

[0056] Although preferred embodiments of the present invention have been described above, the present invention is not limited to the above-described embodiments and various modifications and variations are possible within the spirit and scope of the present invention. The present invention can also be realized by providing a program that implements one or more functions of the present invention to a system or device via a network or a storage medium, and having one or more general-purpose processors (ASICs) in the computer of the system or device read and execute the program. The present invention can also be realized by a dedicated processor (e.g., an ASIC or FPGA) that implements one or more functions. Furthermore, the present invention can also be realized by a combination of a general-purpose processor and a dedicated processor. Note that the term "processor" here refers to a processor in a broad sense and includes both general-purpose and dedicated processors. Furthermore, the processing that implements the present invention may be performed by a single processor alone, or may be performed by multiple processors located in physically separate locations in cooperation with each other. Furthermore, although the HMD 1000 is used in the above embodiments as a device to which the information processing device can be applied, this is not limited thereto and can also be, for example, a desktop or notebook personal computer, a tablet terminal, a smartphone, or the like. In this case, the personal computer or the like is separately connected to the HMD so as to be able to communicate with it.

[0057] Furthermore, the communication system 2000 may include a case where, for example, the server 1200 (hereinafter simply referred to as the "server") is located outside Japan and the terminal device, the HMD 1000, is located within Japan. Even in this case, files and data can be transmitted from the server to the terminal device, and the terminal device can receive the files and data. Even if the server is located outside Japan, the transmission and reception (transmission and reception) of files and data in this system are performed as a single unit, i.e., without any separate operation by the user of the terminal device. Since the system functions when a terminal device located within Japan receives the files and data, the transmission and reception can be considered to have occurred within Japan. Furthermore, in this system, even if, for example, the server is located outside Japan and the terminal device, the HMD 1000, is located within Japan, the terminal device can still perform the main function of the system (transparency processing). Furthermore, the effect of this function (improved visibility of desired virtual objects) can be realized within Japan. Furthermore, even if the server is located outside Japan, as long as the terminal device constituting the system is located within Japan, the system can be used within Japan using the terminal device. And the use of the system can affect the economic interests of, for example, patent holders.

[0058] The disclosure of this embodiment includes the following configuration, method, and program. (Configuration 1) An information processing device that processes information relating to a space image including at least a virtual space in which avatars of multiple participants can participate, a generating means for generating a virtual object within the spatial image; an acquisition means for acquiring first information on a target virtual object determined to be the virtual object that at least some of the participants are paying attention to, and second information on a blocking virtual object that blocks the target virtual object on the spatial image viewed by a predetermined participant of the plurality of participants; The information processing device is characterized in that the generation means performs transparency processing that makes the transparency of the target virtual object and the obscuring virtual object different from each other based on the first information and the second information. (Configuration 2) The information processing device according to Configuration 1, wherein the acquisition means acquires, as the first information, position information and orientation information of the target virtual object within the spatial image, and acquires, as the second information, position information and orientation information of the obscuring virtual object within the spatial image. (Configuration 3) The information processing device according to Configuration 2, wherein when it is determined based on the first information and the second information that the target virtual object and the occluding virtual object are in an overlapping state, the generation means performs the transparency processing. (Configuration 4) In the information processing device according to Configuration 3, when it is determined that the target virtual object is located further back than the occluding virtual object, the generation means performs the transparency processing by increasing the transparency of the occluding virtual object to be greater than the transparency of the target virtual object. (Configuration 5) The information processing device according to configuration 4, wherein when the transparency of the target virtual object is increased, the degree of increase can be changed. (Configuration 6) The information processing device according to configuration 3, wherein the generating means stops the transparency processing when it is determined that the overlapping state has been released. (Configuration 7) The information processing device according to any one of configurations 1 to 6, wherein the virtual object of interest is determined to be a virtual object that is being watched by a predetermined percentage of the participants among the plurality of participants. (Configuration 8) The information processing device described in any one of configurations 1 to 7, characterized in that the first information includes information about a plurality of candidate virtual objects that can become the target virtual object, and the target virtual object is determined from among the plurality of candidate virtual objects. (Configuration 9) The information processing device according to any one of configurations 1 to 8, wherein the occluding virtual object is an avatar of a participant other than the predetermined participant among the plurality of participants. (Configuration 10) A detection means for detecting the line of sight of the predetermined participant is provided, The information processing device according to any one of configurations 1 to 9, wherein the target virtual object and the occluding virtual object on the spatial image visually recognized by the predetermined participant are determined based on the detection results of the detection means. (Configuration 11) The information processing device according to any one of configurations 1 to 10, characterized in that it is a head-mounted display including display means for displaying the spatial image. (Configuration 12) A server communicably connected to a plurality of information processing devices according to any one of configurations 1 to 11, Each of the information processing devices has a detection means for detecting the line of sight of the participant using the information processing device; a transmitting means for transmitting a detection result by the detecting means to the server, The server a receiving means for receiving the detection result of the detecting means from the transmitting means; an identification means for identifying the target virtual object and the obscuring virtual object based on the detection result of the detection means received by the reception means; a transmitting unit configured to transmit the results of the identification by the identifying unit to the information processing device as the first information and the second information. (Method 1) A method for controlling an information processing device that processes information relating to a space image including at least a virtual space in which avatars of multiple participants can participate, comprising: a generation step of generating a virtual object within the spatial image; an acquisition step of acquiring first information on a target virtual object determined to be the virtual object that at least some of the participants are paying attention to, and second information on a blocking virtual object that blocks the target virtual object on the spatial image viewed by a predetermined participant of the plurality of participants, a control method for an information processing device, characterized in that in the generation step, transparency processing is performed to make the transparency of the target virtual object and the obscuring virtual object different from each other based on the first information and the second information. (Program 1) A program for causing a computer to execute each means of the information processing device according to any one of configurations 1 to 11. [Explanation of symbols]

[0059] 1000 HMD 1001 CPU 1006 Communications Department 1009 GPU 1200 Server 1201 CPU 1205 Communications Department 3000~3005 Avatar 3006 Virtual Whiteboard 8000 Virtual Space Conference

Claims

1. An information processing device that processes information relating to a space image including at least a virtual space in which avatars of multiple participants can participate, a generating means for generating a virtual object within the spatial image; an acquisition means for acquiring first information on a target virtual object determined to be the virtual object that at least some of the participants are paying attention to, and second information on a blocking virtual object that blocks the target virtual object on the spatial image viewed by a predetermined participant of the plurality of participants; The information processing apparatus is characterized in that the generation means performs transparency processing that makes the transparency of the target virtual object and the transparency of the obstructing virtual object different from each other based on the first information and the second information.

2. 2. The information processing device according to claim 1, wherein the acquisition means acquires, as the first information, position information and orientation information of the target virtual object within the spatial image, and acquires, as the second information, position information and orientation information of the obscuring virtual object within the spatial image.

3. 3. The information processing device according to claim 2, wherein when it is determined based on the first information and the second information that the target virtual object and the obscuring virtual object are in an overlapping state, the generation means performs the transparency processing.

4. 4. The information processing device according to claim 3, wherein, when it is determined that the target virtual object is located behind the occluding virtual object, the generation means performs the transparency processing by increasing the transparency of the occluding virtual object more than the transparency of the target virtual object.

5. 5. The information processing apparatus according to claim 4, wherein when the transparency of the target virtual object is increased, the degree of increase can be changed.

6. 4. The information processing apparatus according to claim 3, wherein said generating means stops said transparency processing when it is determined that said overlapping state has been released.

7. 2 . The information processing apparatus according to claim 1 , wherein the virtual object of interest is determined to be a virtual object that is being watched by a predetermined percentage of the participants among the plurality of participants.

8. 2 . The information processing apparatus according to claim 1 , wherein the first information includes information about a plurality of candidate virtual objects that can become the target virtual object, and the target virtual object is determined from among the plurality of candidate virtual objects.

9. The information processing apparatus according to claim 1 , wherein the occluding virtual object is an avatar of a participant other than the predetermined participant among the plurality of participants.

10. a detection means for detecting the line of sight of the predetermined participant; 2 . The information processing device according to claim 1 , wherein the target virtual object and the occluding virtual object on the spatial image visually recognized by the predetermined participant are each determined based on a detection result by the detection means.

11. 2. The information processing apparatus according to claim 1, wherein the information processing apparatus is a head-mounted display having a display means for displaying the spatial image.

12. A server communicably connected to a plurality of information processing devices according to claim 1, Each of the information processing devices has a detection means for detecting the line of sight of the participant using the information processing device; a transmitting means for transmitting a detection result by the detecting means to the server, The server a receiving means for receiving the detection result of the detecting means from the transmitting means; an identification means for identifying the target virtual object and the obscuring virtual object based on the detection result of the detection means received by the reception means; a transmitting unit configured to transmit the result of the identification by the identifying unit to the information processing device as the first information and the second information.

13. A method for controlling an information processing device that processes information relating to a space image including at least a virtual space in which avatars of multiple participants can participate, comprising: a generation step of generating a virtual object within the spatial image; an acquisition step of acquiring first information on a target virtual object determined to be the virtual object that at least some of the participants are paying attention to, and second information on a blocking virtual object that blocks the target virtual object on the spatial image viewed by a predetermined participant of the plurality of participants, a control method for an information processing device, characterized in that, in the generation step, transparency processing is performed to make the transparency of the target virtual object and the obscuring virtual object different from each other based on the first information and the second information.

14. 2. A program for causing a computer to execute each means of the information processing apparatus according to claim 1.

Citation Information

Patent Citations

  • Mixed reality presentation system, information processing apparatus and control method thereof, and program

    JP2018106297A

  • Display system, display method, controller, and computer program

    JP7333051B2