Information processing apparatus
The information processing apparatus addresses the discomfort in MR communication by synthesizing visual field and person images, shifting the display position to simulate face-to-face interaction, thereby enhancing the communication experience.
Patent Information
- Application Number
- JP2023208220
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-11
- Publication Date
- 2025-06-23
AI Technical Summary
Existing MR technology makes communication uncomfortable for remote participants, as they cannot replicate the feeling of facing each other, leading to discomfort during video communication using HMDs.
An information processing apparatus that acquires visual field images, person images, and information on the positional difference between imaging devices and display positions, then controls the display to synthesize these images, shifting the visual field image display position to match the person image's position, thereby creating a composite reality image that simulates face-to-face interaction.
Enables comfortable communication between individuals at separate locations by creating a composite reality image that simulates face-to-face interaction, reducing discomfort and enhancing the sense of immersion.
Smart Images

Figure 2025092848000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, and particularly to composite reality image display.
Background Art
[0002] As a technology for seamlessly integrating the real world and the virtual world in real time, a composite reality, so-called MR (Mixed Reality) technology, is known. The MR technology may be used in a video see-through type head-mounted display (HMD). In a video see-through type HMD, for example, a range corresponding to the field of view of a user wearing the HMD in the real space is imaged by a video camera. Then, CG (Computer Graphic) is synthesized with the image (video) of the real space obtained by the video camera, and the synthesized image (an image obtained by synthesizing CG with the image of the real space) is displayed on a display panel inside the HMD. By looking at the synthesized image displayed on the display panel, the user can obtain a feeling as if the virtual object represented by the CG exists in the real space.
[0003] Also, in a video see-through type HMD, by synthesizing an image of a person imaged at a remote location with the image of the real space, an experience (feeling) as if the user wearing the HMD is in the same space as the person at the remote location can be provided to the user.
[0004] Patent Document 1 discloses a technique for changing the orientation of a virtual screen on which an image of a person at a remote location is projected according to the position of the person in chroma key synthesis. According to the technique disclosed in Patent Document 1, a synthesized image without a sense of incongruity can be obtained for a user wearing an HMD.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] Consider a case where a user wearing an HMD communicates with a person in a remote location using their respective images (videos). In this case, with the technology disclosed in Patent Document 1, the communication may become uncomfortable (unnatural). For example, a user wearing an HMD can communicate with the person in a remote location without discomfort as if they were facing each other face to face. However, the person in the remote location cannot obtain the feeling of facing the user wearing the HMD and may feel discomfort in communicating with the user.
[0007] An object of the present invention is to provide a technology that enables two persons in separate locations to communicate with each other using their respective images (videos) without discomfort.
Means for Solving the Problems
[0008] A first aspect of the present invention includes a first acquisition means for acquiring a visual field image representing a range corresponding to the visual field of the user in the real space where the user is located, a second acquisition means for acquiring a person image which is an image of a person imaged at a remote location, and a third acquisition means for acquiring information on the difference between the position of an imaging device that transmits an image of the user imaged in the real space to the remote location and the position where the person image is displayed. The control means controls a display means to display a composite image obtained by synthesizing the person image with the visual field image, and the control means shifts the position where the visual field image is displayed according to the difference. The information processing apparatus is characterized by this. obtains, and control means for controlling a display means to display a composite image obtained by synthesizing the person image with the visual field image, wherein the control means shifts the position where the visual field image is displayed according to the difference, and is an information processing apparatus.
[0009] A second aspect of the present invention includes a first acquisition step of acquiring a field-of-view image representing a range corresponding to the field of view of a user in the real space where the user is located; a second acquisition step of acquiring a person image which is an image of a person located at a remote location; a third acquisition step of acquiring information on the difference between the position of an imaging device that transmits an image of the user in the real space to the remote location and the position where the person image is displayed; and a control step of controlling a display means to display a composite image obtained by compositing the person image with the field-of-view image. In the control step, the position where the field-of-view image is displayed is shifted according to the difference. This is a control method for an information processing apparatus.
[0010] A third aspect of the present invention is a program for causing a computer to function as each means of the information processing apparatus described above. A fourth aspect of the present invention is a computer-readable storage medium storing a program for causing a computer to function as each means of the information processing apparatus described above.
Advantages of the Invention
[0011] According to the present invention, when two persons located at mutually distant locations communicate with each other using their respective images (videos), each person can communicate without a sense of discomfort.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Mode for Carrying Out the Invention
[0013] <First Embodiment> Hereinafter, a first embodiment of the present invention will be described. FIG. 1(A) is a schematic diagram showing the configuration of the information processing system 100 according to the first embodiment, and FIG. 1(B) is a block diagram showing the configuration of the information processing system 100. The information processing system 100 includes head-mounted displays (HMDs) 110 and 120 and cameras 130 and 140. The HMD 110, HMD 120, camera 130, and camera 140 are connected to the network 190 and can communicate with each other via the network 190. The space 170 (for example, the user's own room of the person 150) where the person 150 who is the user of the HMD 110 is located is different from the space 180 (for example, the user's own room of the person 160) where the person 160 who is the user of the HMD 120 is located. The person 150 wears the HMD 110 on the head, and the person 160 wears the HMD 120 on the head. The camera 130 is installed in the space 170 where the person 150 is located, and the camera 140 is installed in the space 180 where the person 160 is located.
[0014] The HMD 110 includes a CPU 111, a ROM 112, a RAM 113, an attitude sensor 114, a communication unit 115, an imaging unit 116, and a display unit 117. The CPU 111 is an arithmetic unit that controls the entire HMD 110. For example, it executes various programs stored in the ROM 112 to perform various processes. The ROM 112 is a read-only non-volatile memory that stores various information (e.g., various programs and various parameters). The RAM 113 is a memory that temporarily stores various information and is also used as the work memory of the CPU 111. The attitude sensor 114 is a sensor that detects the attitude of the HMD 110 and includes, for example, at least one of an acceleration sensor and a geomagnetic sensor. A gyro sensor is a type of acceleration sensor. The communication unit 115 communicates with external devices by wire or wirelessly. The imaging unit 116 acquires a visual field image representing a range (in front of the HMD 110) corresponding to the visual field of the person 150 wearing the HMD 110 in the real space by imaging the range. The imaging unit 116 includes, for example, an imaging sensor such as a CCD sensor or a CMOS sensor. The display unit 117 can display various images and various information. The display unit 117 includes, for example, a display panel such as a liquid crystal panel or an organic EL panel. The person 150 can view the image (video) displayed on the display unit 117 by wearing the HMD 110. The HMD 110 is a video see-through type HMD, and the CPU 111 can display the visual field image obtained by the imaging unit 116 on the display unit 117.
[0015] The HMD 120 includes a CPU 121, a ROM 122, a RAM 123, an attitude sensor 124, a communication unit 125, an imaging unit 126, and a display unit 127. The configuration of the HMD 120 is the same as that of the HMD 110.
[0016] The camera 130 is an imaging device that obtains a spatial image representing at least a part of the space 170 by imaging the person 150 (the space 170 where the person 150 is located). The camera 130 transmits the spatial image of the space 170 to the HMD 120 via the network 190. In the HMD 120, at least a part of the spatial image of the space 170 is displayed on the display unit 127 under the control of the CPU 121. Among the spatial images of the space 170, the part (display range) displayed on the display unit 127 can be changed according to the posture of the HMD 120 (the detection result of the posture sensor 124). By looking at the image displayed on the display unit 127, the person 160 can obtain a feeling of being in the same space 170 as the person 150.
[0017] The camera 140 is an imaging device that obtains a spatial image representing at least a part of the space 180 by imaging the person 160 (the space 180 where the person 160 is located). The camera 140 transmits the spatial image of the space 180 to the HMD 110 via the network 190. In the HMD 110, the area of the person 160 is extracted from the spatial image of the space 180 under the control of the CPU 111. As a result, a person image that is an image of the person 160 (an image representing only the person 160) is obtained. Then, the person image is synthesized with the view image (the view image of the person 150) obtained by the imaging unit 116, and the synthesized image (an image obtained by synthesizing the person image with the view image) is displayed on the display unit 117. By looking at the image displayed on the display unit 117, the person 150 can obtain a feeling of the person 160 being in the same space 170 as himself / herself (the person 150).
[0018] In this way, in the information processing system 100, the person 150 in the space 170 and the person 160 in the space 180 can obtain a feeling of being in the same space 170.
[0019] According to the above-described operation, the person 160 is invited to the space 170 where the person 150 is located. Therefore, the space 170 where the person 150 is located can be regarded as the Host side, and the space 180 where the person 160 is located can be regarded as the Guest side.
[0020] The cameras 130 and 140 may be monocular cameras, but are preferably stereo cameras. If the camera 130 is a stereo camera, the person 150 can wear the HMD 110 and stereoscopically view two images with a parallax, so that a higher sense of immersion can be obtained. Similarly, if the camera 140 is a stereo camera, the person 160 can wear the HMD 120 and stereoscopically view two images with a parallax, so that a higher sense of immersion can be obtained.
[0021] Also, the imaging range of the Host-side camera 130 is preferably wide. For example, the viewing angle (field angle) may be in the range of 180 degrees, or may be in the range of 360 degrees. The image obtained by the Host-side camera 130 may be a hemispherical image (VR180 image), a full-spherical image (omnidirectional image), or a panoramic image. Since the imaging range of the Host-side camera 130 is wide, the Guest-side person 160 can see the image of the Host-side space 170 even if the face direction is changed greatly, and can obtain a feeling of being in the Host-side space 170.
[0022] In the first embodiment, it is assumed that the Guest-side person 160 is facing the direction of the Guest-side camera 140. Therefore, the Host-side person 150 can communicate with the Guest-side person 160 (a person at a remote location) without a sense of discomfort as if they were facing each other face to face.
[0023] However, in the prior art, the Guest-side person 160 cannot obtain a feeling as if facing the Host-side person 150 (a person in a remote location), and may feel uncomfortable in communicating with the Host-side person 150. In the information processing system 100, the position of the Host-side camera 130 becomes the viewpoint of the Guest-side person 160. Here, assume that a person image representing the Guest-side person 160 is displayed at a position different from the position of the Host-side camera 130. In this case, when the Host-side person 150 faces the direction of the person image, the Guest-side person 160 cannot obtain a feeling as if facing the Host-side person 150.
[0024] Therefore, in the first embodiment, a person image is displayed at the position of the Host-side camera 130. By doing so, the Guest-side person 160 can also obtain a feeling as if facing the Host-side person 150.
[0025] FIG. 2 is a flowchart showing the overall processing performed by the information processing system 100. For example, when the information processing system 100 (Host-side HMD 110, Guest-side HMD 120, Host-side camera 130, and Guest-side camera 140) is activated, the overall processing in FIG. 2 starts.
[0026] In step S201, the CPU 111 of the Host-side HMD 110 determines whether the Host-side HMD 110, Guest-side HMD 120, Host-side camera 130, and Guest-side camera 140 are connected to each other (via the network 190). The CPU 111 waits until it determines that these four devices are connected to each other, and when it determines that the four devices are connected to each other, it performs the processing in step S202.
[0027] In step S201, the CPU 121 of the Guest-side HMD 120 also determines whether the above four devices are connected to each other. The CPU 121 waits until it determines that the four devices are connected to each other, and when it determines that the four devices are connected to each other, it performs the process of step S203. In FIG. 2, step S203 is shown as the next step after step S202, but the process of step S203 may be performed before the process of step S202, or the processes of step S202 and step S203 may be performed in parallel. Yes.
[0028] In step S202, the CPU 111 of the Host-side HMD 110 performs Host-side display processing. The details of the Host-side display processing will be described later with reference to FIG. 3.
[0029] In step S203, the CPU 121 of the Guest-side HMD 120 performs Guest-side display processing. The details of the Guest-side display processing will be described later with reference to FIG. 5.
[0030] In step S204, the CPU 111 of the Host-side HMD 110 determines whether to stop the information processing system 100. The CPU 111 repeatedly performs the Host-side display processing (step S202) for each frame until it determines to stop the information processing system 100, and when it determines to stop the information processing system 100, it ends the overall processing of FIG. 2.
[0031] In step S204, the CPU 121 of the Guest-side HMD 120 also determines whether to stop the information processing system 100. The CPU 121 repeatedly performs the Guest-side display processing (step S203) for each frame until it determines to stop the information processing system 100, and when it determines to stop the information processing system 100, it ends the overall processing of FIG. 2.
[0032] FIG. 3 is a flowchart showing the Host-side display processing (step S202 in FIG. 2). The Host-side display processing is performed by the CPU 111 of the Host-side HMD 110.
[0033] In step S301, the CPU 111 uses the communication unit 115 to receive (acquire) a spatial image of the Guest-side space 180 from the Guest-side camera 140. If a part of the Guest-side person 160 is not shown in the spatial image of the Guest-side space 180, an uncomfortable composite image (an unnatural composite image in which a part of the Guest-side person 160 is not drawn) may be obtained. Therefore, it is preferable that the entire body of the Guest-side person 160 is shown in the spatial image of the Guest-side space 180.
[0034] In step S302, the CPU 111 performs geometric transformation processing on the spatial image acquired in step S301 as necessary. For example, when the Guest-side camera 140 performs imaging using an ultra-wide-angle lens such as a fish-eye lens, the CPU 111 needs to perform geometric transformation processing such as orthographic cylindrical transformation in order to obtain a spatial image suitable for synthesis.
[0035] In step S303, the CPU 111 extracts the area of the Guest-side person 160, which is the main subject, from the spatial image acquired in step S301 (the spatial image after the geometric transformation processing in step S302). As a result, a person image, which is an image of the Guest-side person 160 (an image representing only the Guest-side person 160), is obtained. The method for extracting the area of the Guest-side person 160 is not particularly limited. For example, the Guest-side person 160 may be imaged against a green screen, and the area of the Guest-side person 160 may be extracted by chroma key processing. The area of the Guest-side person 160 may be extracted using an arithmetic unit (trained model) trained by machine learning such as deep learning. The extraction method is selected, for example, in consideration of system resources and extraction accuracy.
[0036] In step S304, the CPU 111 determines the composite position (display position) of the person image acquired in step S303 (composite position determination processing). The details of the composite position determination processing will be described later with reference to FIG. 4.
[0037] In step S305, the CPU 111 synthesizes the person image acquired in step S303 with the field-of-view image (the field-of-view image of the host-side person 150) obtained by the imaging unit 116 of the host-side HMD 110. The person image is synthesized at the synthesis position determined in step S304. As a result, a synthesized image in which the guest-side person 160 is present in the host-side space 170 is obtained. Although details will be described later, in step S304, the synthesis size (display size) of the person image is also determined. In step S305, the person image is synthesized with the synthesis size determined in step S304. In step S306, the CPU 111 displays the synthesized image acquired in step S305 on the display unit 117.
[0038] Figure 4 is a flowchart showing the synthesis position determination process (step S304 in FIG. 3). As described above, in the information processing system 100, the position of the host-side camera 130 becomes the viewpoint of the guest-side person 160. Therefore, it is preferable to display the person image at the position of the host-side camera 130, and it is more preferable to display the face area of the guest-side person 160 in the person image at the position of the host-side camera 130. And it is particularly preferable to display the eye area (the area of the guest-side HMD 120) of the guest-side person 160 in the person image at the position of the lens of the host-side camera 130. In the synthesis position determination process of FIG. 4, the CPU 111 determines the position of the host-side camera 130 as the synthesis position (display position) of the person image.
[0039]
[0040] In step S401, the CPU 111 detects the host-side camera 130 from the field-of-view image (the field-of-view image of the host-side person 150) obtained by the imaging unit 116 of the host-side HMD 110. For example, the host-side camera 130 is detected by pattern matching.
[0041] In addition, if the Host-side camera 130 can be detected from the Host-side space 170, the method is not particularly limited. For example, a marker (known marker) pre-arranged in the Host-side space 170 may be detected, and the Host-side camera 130 may be detected according to the detection result of the marker. The marker may be arranged on the Host-side camera 130 or at a position away from the Host-side camera 130. If the marker is arranged on the Host-side camera 130, even if the position of the Host-side camera 130 is changed, the Host-side camera 130 can be detected according to the detection result of the marker. For example, the position of the marker can be detected as the position of the Host-side camera 130. The Host-side camera 130 may be detected according to an instruction from the Host-side person 150 (user of the Host-side HMD 110). For example, the position designated by the Host-side person 150 may be detected as the position of the Host-side camera 130. The Host-side camera may be detected without using the field-of-view image.
[0042] The Host-side person 150 can issue various instructions, for example, using a controller (not shown) of the Host-side HMD 110. The field-of-view image may be transmitted to an information processing device such as a personal computer or a smartphone via the network 190, and the field-of-view image may be displayed on the information processing device side. Then, the Host-side person 150 may specify the position of the Host-side camera 130 using an operation member provided on the information processing device or an operation device (such as a keyboard and a mouse) connected to the information processing device.
[0043] In step S402, the CPU 111 determines the composite size (display size) of the person image based on the positional relationship between the Host-side person 150 (Host-side HMD 110) and the Host-side camera 130. For example, the CPU 111 determines the resize ratio of the person image as a value corresponding to the composite size (display size) of the person image.
[0044] In a predetermined direction such as the vertical direction, the actual size of the Host-side camera 130 is 15 Assume that it is 15 cm, and the size of the Host-side camera 130 in the field-of-view image is 150 pix. Also assume that the actual size of the face of the Guest-side person 160 is 30 cm, and the size of the face of the Guest-side person 160 in the person image is 500 pix. In this case, if the person image is synthesized into the field-of-view image such that 1 cm becomes 10 pix (= 150 pix / 15 cm), the size of the Guest-side person 160 and the sizes of other objects will be consistent in the synthesized image. Therefore, the CPU 111 determines the reduction ratio of 300 pix / 500 pix as the reduction ratio of the person image so that a face of 30 cm is synthesized at 300 pix.
[0045] Let the actual size of the Host-side camera 130 be A cm, the size of the Host-side camera 130 in the field-of-view image be X pix, the actual size of the face of the Guest-side person 160 be B cm, and the size of the face of the Guest-side person 160 in the person image be Y pix. In this case, the resizing ratio α of the person image can be calculated using the following Equation 1. α = (B × X) / (A × Y) ···(Equation 1)
[0046] Incidentally, the actual size of the host-side camera 130 may be registered in advance in the host-side HMD 110. The actual size of the host-side camera 130 may be calculated based on the focal length of the host-side camera 130, the shooting distance of the host-side camera 130, and the size of the host-side camera 130 in the field-of-view image. The size of the host-side camera 130 in the field-of-view image may be the size (number of pixels) of the area of the host-side camera 130 detected by pattern matching, or may be a size corresponding to the size of the marker in the field-of-view image. The size of the host-side camera 130 in the field-of-view image may be specified by the host-side person 150. The actual size of the host-side camera 130 may be calculated according to the size of the marker in the field-of-view image. Similarly, the actual size of the face of the guest-side person 160 may be registered in advance in the host-side HMD 110. The actual size of the face of the guest-side person 160 may be calculated based on the focal length of the guest-side camera 140, the shooting distance of the guest-side camera 140, and the size of the face of the guest-side person 160 in the person image.
[0047] In step S403, the CPU 111 determines the position of the host-side camera 130 detected in step S401 as the composite position (display position) of the person image. For example, the composite position of the person image is determined so that the area of the face of the guest-side person 160 in the person image is arranged at the position of the host-side camera 130. The composite position of the person image may be determined so that the area of the eyes of the guest-side person 160 (the area of the guest-side HMD 120) in the person image is arranged at the position of the lens of the host-side camera 130.
[0048] FIG. 5 is a flowchart showing the guest-side display process (step S203 in FIG. 2). The guest-side display process is performed by the CPU 121 of the guest-side HMD 120.
[0049] In step S501, the CPU 121 receives (acquires) the space image of the host-side space 170 from the host-side camera 130 using the communication unit 125.
[0050] In step S502, the CPU 121 performs geometric transformation processing on the spatial image acquired in step S501 as necessary. For example, when the host-side camera 130 is a wide-angle stereo camera, the CPU 121 needs to perform geometric transformation processing such as orthographic cylindrical transformation and perspective projection transformation in order to obtain a spatial image suitable for stereoscopic vision.
[0051]
[0051] In step S503, the CPU 121 displays the spatial image (the spatial image after the geometric transformation processing in step S502) acquired in step S501 on the display unit 127.
[0052] As described above, according to the first embodiment, on the host side, a person image that is an image of a person located remotely is acquired. Then, the acquired person image is displayed at the position of the imaging device that transmits an image of the user in the real space where the user is located to a remote location. By doing so, when two people located at separate locations communicate with each other using their respective images (videos), each person can communicate without a sense of discomfort.
[0053] <Second Embodiment> The following describes the second embodiment of the present invention. In the following, descriptions of the same points as those in the first embodiment (for example, the same configuration and processing as those in the first embodiment) are omitted, and differences from the first embodiment are described.
[0054] In the first embodiment, it was assumed that the face of the guest-side person 160 was arranged at the position of the host-side camera 130, or the eyes of the guest-side person 160 (guest-side HMD 120) were arranged at the position of the lens of the host-side camera 130. However, depending on the height of the guest-side person 160 and the installation position of the host-side camera 130 (for example, the height of the tripod supporting the host-side camera 130), an unnatural composite image may be obtained. Therefore, in the second embodiment, the composite position (display position) of the person image is adjusted.
[0055] Figs. 6(A) to 6(E) are diagrams showing various images according to the second embodiment. Fig. 6(A) shows a field-of-view image (field-of-view image of the host-side person 150) obtained by the imaging unit 116 of the host-side HMD 110. The host-side camera 130 is shown in the center of the field-of-view image. The host-side camera 130 is supported by a tripod. Fig. 6(B) shows a person image representing the guest-side person 160. Note that the guest-side HMD 120 is omitted in the person image of Fig. 6(B).
[0056] Fig. 6(C) shows a composite image obtained by compositing the person image of Fig. 6(B) with the field-of-view image of Fig. 6(A). The eyes (guest-side HMD 120) of the guest-side person 160 are arranged at the position of the lens of the host-side camera 130 so that the guest-side person 160 can be felt as facing the host-side person 150 face to face. However, since the height of the tripod is too low, the composite image is an unnatural image in which the feet of the guest-side person 160 are buried under the floor. In addition, when the height of the tripod is too high, the composite image is an unnatural image in which the guest-side person 160 floats in the air. Also, in the case of these unnatural composite images, since the viewpoint of the guest-side person 160 is at an unnatural position (too high or too low), the guest-side person 160 feels discomfort.
[0057] Therefore, on the host side, when the composite image shown in Fig. 6(C) is obtained, the composite position (display position) of the person image is adjusted. Fig. 7 is a flowchart showing the adjustment operation of the composite position.
[0058] In step S701, the CPU 111 in the host-side HMD 110 moves the person image according to an instruction from the host-side person 150 (user of the host-side HMD 110).
[0059] In step S702, the CPU 111 increases the transparency of the person image. An instruction to change the transparency of the person image may be given by the host-side person 150, or the transparency of the person image may be automatically changed when an instruction to move the person image is given. The transparency of the person image By raising it, the positional relationship between the person image and the Host-side camera 130 becomes easier to grasp. FIG. 6(D) shows the composite image after step S702.
[0060] The Host-side person 150 can give various instructions, for example, using a controller (not shown) of the Host-side HMD 110. The composite image may be transmitted to an information processing device such as a personal computer or a smartphone via the network 190, and the view image may be displayed on the information processing device side. Then, the Host-side person 150 may give an instruction to move the person image using an operation member provided on the information processing device or an operation device (for example, a keyboard and a mouse) connected to the information processing device.
[0061] In step S703, the Host-side person 150 adjusts the height of the tripod. For example, the height of the tripod is adjusted so that the face area of the Guest-side person 160 in the person image is arranged at the position of the Host-side camera 130. The height of the tripod may be adjusted so that the eye area (the area of the Guest-side HMD 120) of the Guest-side person 160 in the person image is arranged at the position of the lens of the Host-side camera 130. FIG. 6(E) shows the composite image after step S703. After that, for example, according to an instruction from the user, the transparency of the person image is restored.
[0062] Note that when the position of the Host-side camera 130 changes, the CPU 111 may update the composite position (display position) of the person image so that the person image continues to be displayed at the position of the Host-side camera 130 (the person image follows the Host-side camera 130). By doing so, the instruction to move the person image becomes unnecessary. Also in this case, when the Host-side person 150 adjusts the height of the tripod, the transparency of the person image may be changed (increased).
[0063] As described above, according to the second embodiment, the composite position (display position) of the person image is adjusted to a suitable position. At that time, the transparency of the person image is changed (increased) to assist the adjustment of the composite position. By doing so, when two people who are in separate locations communicate with each other using their respective images (videos), each person can communicate with less discomfort.
[0064] <Third Embodiment> Hereinafter, a third embodiment of the present invention will be described. In the following, the description of the same points as those in the first embodiment and the second embodiment (for example, the same configuration and processing as those in the first embodiment) will be omitted, and the points different from those in the first embodiment and the second embodiment will be described.
[0065] In the first and second embodiments, as shown in FIG. 6(E), the eyes of the guest-side person 160 (guest-side HMD 120) are arranged at the position of the host-side camera 130. For this reason, the host-side camera 130 must be installed at the position where the guest-side person 160 (the person image of the guest-side person 160) is to be displayed (superimposed). However, it may not be possible to install the host-side camera 130 at the position where the guest-side person 160 is to be displayed. In such a case, a marker may be arranged at the position where the guest-side person 160 is to be displayed, and the guest-side person 160 may be displayed at the position of the marker. However, if the position of the marker (the position where the guest-side person 160 is to be displayed) is different from the position of the host-side camera 130, the line of sight between the host-side person 150 and the guest-side person 160 will not match, making it difficult to have a natural conversation within the same space and significantly reducing the sense of immersion.
[0066] FIG. 8 is a diagram showing the state of the host-side space 170, which is a bird's-eye view of the host-side space 170 from directly above. In FIG. 8, the host-side person 150 is sitting on the sofa 801 actually placed in the host-side space 170, and the sofa actually placed in the host-side space 170 It is desired to display the Guest-side person 160 on the sofa 802. However, since the top of the sofa 802 is unstable, it is not practical to place the Host-side camera 130 on the sofa 802. Therefore, a marker 803 is placed on the sofa 802. By displaying the Guest-side person 160 at the position of the marker 803, the Guest-side person 160 can be displayed on the sofa 802. As a result, a more realistic situation where the Host-side person 150 and the Guest-side person 160 sit on the sofa and talk in the same space can be realized.
[0067] If the Host-side camera 130 can be installed at the position of the marker 803 (on the sofa 802), the Host-side person 150 and the Guest-side person 160 can face each other and talk. However, as described above, it is not practical to place the Host-side camera 130 on the sofa 802. Therefore, in FIG. 8, the Host-side camera 130 is installed near the sofa 802. If the position of the marker (the position where the Guest-side person 160 is displayed) and the position of the Host-side camera 130 are different, the Guest-side person 160 may feel that the Host-side person 150 is not facing them.
[0068] FIG. 9(A) shows a composite image (a composite image in which the person image of the Guest-side person 160 is synthesized with the visual field image of the Host-side person 150) displayed on the display unit 117 of the Host-side HMD 110 in the state of FIG. 8. FIG. 9(B) shows a space image (a space image of the Host-side space 170) displayed on the display unit 127 of the Guest-side HMD 120 in the state of FIG. 8. Assume that the Host-side person 150 is wearing the Host-side HMD 110 and facing the sofa 802. Assume that the Guest-side person 160 is wearing the Guest-side HMD 120 and facing the Guest-side camera 140 directly. Since only the image of the Guest-side person 160 needs to be obtained as the image of the Guest-side space 180, there are almost no restrictions on the installation of the Guest-side camera 140, and the Guest-side person 160 can easily face the Guest-side camera 140.
[0069] In Fig. 9(A), the Guest-side HMD 120 is omitted to clearly show the face direction of the Guest-side person 160. Similarly, in Fig. 9(B), the Host-side HMD 110 is omitted to clearly show the face direction of the Host-side person 150. In the case of an optical see-through type HMD, the eye movements of the person wearing the HMD can be visually recognized from the outside of the HMD. Also, as a technology related to a video see-through type HMD, a technology that enables visual recognition of the eye movements of the person wearing the HMD from the outside of the HMD has been proposed. Therefore, regardless of whether the Guest-side HMD 120 is of the optical see-through type or the video see-through type, the Host-side person 150 can recognize the face direction and line of sight of the Guest-side person 160 wearing the Guest-side HMD 120 from the composite image. Similarly, the Guest-side person 160 can recognize the face direction and line of sight of the Host-side person 150 wearing the Host-side HMD 110 from the spatial image of the Host-side space 170.
[0070] In the composite image of Fig. 9(A), the Guest-side person 160 is shown on the sofa 802, and the Host-side camera 130 is shown beside it. Since the Guest-side person 160 is facing directly the Guest-side camera 140, in Fig. 9(A), the Guest-side person 160 is facing straight ahead. Therefore, the Host-side person 150 can obtain the feeling of facing the Guest-side person 160 face to face by looking straight ahead.
[0071] The spatial image of Fig. 9(B) is an image captured by the Host-side camera 130. The Host-side person 150 is facing the direction of the Guest-side person 160 displayed at the position of the marker 803 (on the sofa 802) and is not facing the Host-side camera 130. Therefore, in Fig. 9(B), the Host-side person 150 is not facing straight ahead, and the Guest-side person 16 0 makes the host-side person 150 feel uncomfortable as if they are not facing themselves. Also, since the host-side person 150 can get the feeling of facing the guest-side person 160 face to face, they cannot even perceive that the guest-side person 160 is feeling uncomfortable.
[0072] Therefore, in the third embodiment, the position of the field-of-view image (the field-of-view image of the host-side person 150) displayed on the host-side HMD 110 is adjusted. By doing so, the host-side person 150 can unconsciously turn their face in the direction of the host-side camera 130, and the guest-side person 160 can also get the feeling of facing the host-side person 150 face to face.
[0073] FIG. 10 is a flowchart showing the host-side display process (step S202 in FIG. 2) according to the third embodiment. The host-side display process is performed by the CPU 111 of the host-side HMD 110. Steps S1001 to S1003 are the same as steps S301 to S303 in FIG. 3.
[0074] In step S1004, the CPU 111 acquires information on the difference between the position of the host-side camera 130 in the host-side space 170 and the position where the guest-side person 160 (person image) is displayed. For example, the CPU 111 acquires the angle (line-of-sight deviation angle) by which the line of sight of the guest-side person 160 and the line of sight of the host-side person 150 deviate.
[0075] The method of acquiring the line-of-sight deviation angle will be described in detail with reference to FIG. 11. FIG. 11 is an overhead view similar to FIG. 8. The CPU 111 acquires the angle θ between the straight line passing through the position of the host-side person 150 and the position of the marker 803 and the straight line passing through the position of the host-side person 150 and the position of the host-side camera 130 as the line-of-sight deviation angle. For example, the CPU 111 acquires (calculates) the line-of-sight deviation angle θ from the distance d1 from the position of the host-side person 150 to the position of the marker 803 and the distance d2 from the position of the host-side person 150 to the position of the host-side camera 130 according to the following formula 2. θ = arccos(d1 / d2) ···(Equation 2)
[0076] The method for obtaining the distances d1 and d2 is not limited. However, when the Host-side HMD 110 has a function of obtaining a depth map corresponding to the field-of-view image, the CPU 111 can obtain the distance d1 from the depth map. For example, the CPU 111 obtains, as the distance d1, the depth corresponding to the position of the marker 803 among the plurality of depths (plural distances) indicated by the depth map. Also, usually, the shooting distance (focus distance) of the Host-side camera 130 is adjusted so as to be in focus on the Host-side person 150. Therefore, the CPU 111 can also obtain the information on the shooting distance of the Host-side camera 130 from the Host-side camera 130 via the network 190 as the information on the distance d2.
[0077] Note that the information on the difference between the position of the Host-side camera 130 and the position where the Guest-side person 160 is displayed in the Host-side space 170 may be obtained according to an instruction from the Host-side person 150. For example, the distances d1 and d2 may be specified by the Host-side person 150, and the CPU 111 may calculate the line-of-sight deviation angle θ from the specified distances d1 and d2. The line-of-sight deviation angle θ may be specified by the Host-side person 150. The position where the Guest-side person 160 is displayed and the position of the Host-side camera 130 may be specified by the Host-side person 150, and the CPU 111 may obtain, as the distances d1 and d2, the two depths corresponding to the two specified positions from the depth map.
[0078] Also, the line-of-sight deviation angle θ may be an angle in the horizontal direction (left-right direction), may be an angle in the vertical direction (up-down direction), or may be an angle considering both the horizontal direction and the vertical direction. By considering both the horizontal direction and the vertical direction, a more suitable line-of-sight deviation angle θ can be obtained.
[0079] Return to the description of FIG. 10. In step S1005, the CPU 111 obtains a shift amount (display shift amount) of the position where the visual field image of the Host-side person 150 is displayed according to the information (difference in position, line-of-sight deviation angle θ) obtained in step S1004.
[0080] The method for obtaining the display shift amount will be described in detail with reference to FIG. 12. FIG. 12 shows the composite image displayed on the display unit 117 of the Host-side HMD 110 in the state of FIG. 8. In FIG. 12, subjects other than the sofa 802 are omitted. The CPU 111 obtains a shift amount d (d pixels) for rotating the line of sight of the Host-side person 150 by the line-of-sight deviation angle θ as the display shift amount. When the line-of-sight deviation angle θ is minute, the CPU 111 can obtain (calculate) the display shift amount d from the horizontal display size H (horizontal display resolution, number of horizontal display pixels) and the display angle φ of the display unit 117 according to the following formula 3. d = H × (θ / φ) ···(Formula 2)
[0081] Return to the description of FIG. 10. In step S1006, the CPU 111 shifts the position where the visual field image of the Host-side person 150 is displayed by the display shift amount d obtained in step S1005. As a result, as shown in FIG. 12, the sofa 802 is shifted by the display shift amount d. As a result, the Host-side person 150 unconsciously rotates his / her face by the line-of-sight deviation angle θ and turns it toward the direction of the Host-side camera 130.
[0082] Note that the viewing angle of the field-of-view image is preferably wider than the display viewing angle φ of the display unit 117, and the size (resolution, number of pixels) of the field-of-view image is preferably larger than the display size (display resolution, number of display pixels) of the display unit 117. By doing so, it is possible to suppress the display viewing angle from becoming narrow (angle-of-view loss) due to the shift of the field-of-view image. For example, when the viewing angle of the field-of-view image is equal to the display viewing angle φ of the display unit 117, if the field-of-view image is shifted to the left as shown in FIG. 12, the image displayed at the right end disappears, and the display viewing angle becomes narrow. If the viewing angle of the field-of-view image is wider than the display viewing angle φ of the display unit 117, it is possible to suppress the disappearance of the image displayed at the right end, and thus suppress the narrowing of the display viewing angle.
[0083] In step S1007, the CPU 111 synthesizes the person image acquired in step S1003 with the field-of-view image after the shift in step S1006. The person image is synthesized so as to face the direction of the host-side person 150. However, if the shift in step S1006 is performed after the synthesis of the person image, in the synthesized image, the guest-side person 160 will not face the direction of the host-side person 150. Therefore, in the third embodiment, the person image is synthesized after the shift in step S1006. By doing so, both the host-side person 150 and the guest-side person 160 can obtain a feeling as if the host-side person 150 and the guest-side person 160 are facing each other. Note that the size of the person image may be changed, or the transparency of the person image may be changed. For example, the CPU 111 may determine the size of the person image based on the positional relationship between the host-side person 150 and the position where the person image is to be displayed in the same manner as in the first embodiment.
[0084] As described above, according to the third embodiment, by shifting the field-of-view image of the host-side person 150, natural communication is possible even when the host-side camera 130 cannot be installed at the position where the guest-side person 160 is to be displayed.
[0085] Note that the above embodiments (including modified examples) are merely examples, and configurations obtained by appropriately modifying or changing the configurations of the above embodiments within the scope of the gist of the present invention are also included in the present invention. Configurations obtained by appropriately combining the configurations of the above embodiments are also included in the present invention.
[0086] For example, although an example using a video see-through type HMD has been described, in the first embodiment or the second embodiment, an optical see-through type HMD may display a person image. The optical see-through type HMD may or may not have an imaging unit for acquiring a field-of-view image. Also, the Guest-side HMD 120 may perform the same processing as the Host-side display processing (Figs. 3, 10). By doing so, the Guest-side person 160 can obtain a feeling as if the Host-side person 150 is in the same Guest-side space 180 as itself (Guest-side person 160). Also, at least a part of the above processing assumed to be performed by the HMD may be performed by an information processing device different from the HMD, and the device to which the present invention is applied may not be an HMD. For example, the present invention may be applied to a display device that is not an HMD, such as a tablet terminal. The present invention may be applied to an information processing device that does not have a display unit, such as a personal computer connected to the HMD. The present invention may be applied to a cloud server provided in the network 190.
[0087] The information processing device may be selectively capable of executing the processing of the first embodiment, the processing of the second embodiment, and the processing of the third embodiment. For example, when it is not practical for the user to install a camera at the position where the person image is to be displayed, the user sets a mode for executing the processing of the third embodiment, and when a camera can be installed at the position where the person image is to be displayed, the user sets a mode for executing the processing of the first embodiment.
[0088] <Other Embodiments> The present invention can also be implemented by supplying a program that realizes one or more functions of the above-described embodiments to a system or apparatus via a network or a storage medium, and causing one or more processors in a computer of the system or apparatus to read and execute the program. Further, it can also be implemented by a circuit (for example, ASIC) that realizes one or more functions.
[0089] The disclosure of the present embodiment includes the following configurations, methods, programs, and media. (Configuration 1) A first acquisition means for acquiring a field-of-view image representing a range corresponding to the field of view of the user in the real space where the user is located; A second acquisition means for acquiring a person image that is an image of a person located at a remote location; A third acquisition means for acquiring information on the difference between the position of an imaging device that transmits an image of the user in the real space to the remote location and the position where the person image is displayed; A control means for controlling a display means to display a composite image obtained by synthesizing the person image with the field-of-view image; and the control means shifts the position where the field-of-view image is displayed according to the difference. An information processing apparatus characterized by the above. (Configuration 2) After shifting the position where the field-of-view image is displayed, the control means synthesizes the person image with the field-of-view image. The information processing apparatus according to Configuration 1, characterized by the above. (Configuration 3) The third acquisition means acquires the information on the difference based on the distance from the position of the user to the position of the imaging device and the distance from the position of the user to the position where the person image is displayed. acquires The information processing apparatus according to Configuration 1 or 2, characterized by the above. (Configuration 4) The third acquisition means uses the shooting distance of the imaging device as the distance from the position of the user to the position of the imaging device. The information processing apparatus according to Configuration 3, characterized by the above. (Configuration 5) The third acquisition means acquires, from the depth map corresponding to the visual field image, the distance from the position of the user to the position where the person image is displayed. The information processing apparatus according to Configuration 3 or 4, characterized in that. (Configuration 6) The third acquisition means acquires the difference information according to an instruction from the user. The information processing apparatus according to Configuration 1 or 2, characterized in that. (Configuration 7) Further comprising the display means The information processing apparatus according to any one of Configurations 1 to 6, characterized in that. (Configuration 8) The information processing apparatus is a head-mounted display (HMD). The information processing apparatus according to any one of Configurations 1 to 7, characterized in that. (Configuration 9) The control means controls the display means so as to display the person image in a size based on the positional relationship between the user and the position where the person image is displayed. The information processing apparatus according to any one of Configurations 1 to 8, characterized in that. (Configuration 10) Further comprising a changing means for changing the transparency of the person image The information processing apparatus according to any one of Configurations 1 to 9, characterized in that. (Method) A first acquisition step of acquiring a visual field image representing a range corresponding to the visual field of the user in the real space where the user is located; A second acquisition step of acquiring a person image which is an image of a person located at a remote location; A third acquisition step of acquiring information on the difference between the position of the imaging device that transmits the image of the user imaged in the real space to the remote location and the position where the person image is displayed; A control step of controlling the display means to display a composite image obtained by compositing the person image on the visual field image; and having In the control step, the position where the view image is displayed is shifted according to the difference. A control method for an information processing apparatus, characterized by the above. (Program) A program for causing a computer to function as each means of the information processing apparatus according to any one of Configurations 1 to 10. (Medium) A computer-readable storage medium storing a program for causing a computer to function as each means of the information processing apparatus according to any one of Configurations 1 to 10.
Explanation of Signs
[0090] 110: Head-mounted display (Host-side HMD) 111: CPU
Claims
1. First acquisition means for acquiring a field-of-view image representing a range corresponding to the field of view of the user in the real space where the user is located; Second acquisition means for acquiring a person image which is an image of a person located at a remote location; Third acquisition means for acquiring information on the difference between the position of an imaging device that transmits an image of the user taken in the real space to the remote location and the position where the person image is displayed; Control means for controlling a display means to display a composite image obtained by compositing the person image on the field-of-view image; comprising: The control means shifts the position where the field-of-view image is displayed according to the difference. An information processing apparatus characterized by this.
2. After shifting the position where the field-of-view image is displayed, the control means composites the person image on the field-of-view image. The information processing apparatus according to claim 1, characterized by this.
3. The third acquisition means acquires the information on the difference based on the distance from the position of the user to the position of the imaging device and the distance from the position of the user to the position where the person image is displayed. The information processing apparatus according to claim 1, characterized by this.
4. The third acquisition means uses the shooting distance of the imaging device as the distance from the position of the user to the position of the imaging device. The information processing apparatus according to claim 3, characterized by this.
5. The third acquisition means acquires the distance from the position of the user to the position where the person image is displayed from a depth map corresponding to the field-of-view image. The information processing apparatus according to claim 3, characterized by this.
6. The third acquisition means acquires the information on the difference according to an instruction from the user. The information processing apparatus according to claim 1, characterized by this.
7. further comprising the display means The information processing apparatus according to claim 1, characterized in that.
8. The information processing apparatus is a head-mounted display (HMD) The information processing apparatus according to claim 1, characterized in that.
9. The control means controls the display means to display the person image in a size based on the positional relationship between the user and the position where the person image is displayed The information processing apparatus according to claim 1, characterized in that.
10. further comprising a changing means for changing the transparency of the person image The information processing apparatus according to claim 1, characterized in that.
11. A first acquisition step of acquiring a visual field image representing a range corresponding to the visual field of the user in the real space where the user is located; A second acquisition step of acquiring a person image which is an image of a person located at a remote location; In the real space, the position of the imaging device that transmits an image of the user taken to the remote location A third acquisition step of acquiring information on the difference between the position where the person image is displayed; A control step of controlling the display means to display a composite image obtained by synthesizing the person image with the visual field image having In the control step, the position where the visual field image is displayed is shifted according to the difference A control method for an information processing apparatus, characterized in that.
12. A program for causing a computer to function as each means of the information processing apparatus according to any one of claims 1 to 10.
13. A computer-readable storage medium storing a program for causing a computer to function as each means of the information processing apparatus according to any one of claims 1 to 10.
Citation Information
Patent Citations
Information processing device and image generation method
WO2019097639A1