Information processing device, information processing method, and program
The information processing device adjusts external sound volumes based on real and virtual image associations to allow hearing external sounds during virtual or mixed reality, improving immersion and situational awareness.
Patent Information
- Application Number
- JP2022078133
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-05-11
- Publication Date
- 2025-11-17
- Estimated Expiration
- 2042-05-11
AI Technical Summary
Existing technologies fail to allow viewers to hear external sounds when it is desirable during virtual or mixed reality experiences, particularly when talking actions are not detected.
An information processing device that associates external sounds with real and virtual images, adjusting the relative volume between them based on transparency information and head movement, enabling the combination and output of composite images and sounds through a head-mounted device.
Enables viewers to hear external sounds when necessary during virtual or mixed reality experiences, enhancing immersion and situational awareness without manual switching.
Smart Images

Figure 0007770990000001 
Figure 0007770990000002 
Figure 0007770990000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing technology for processing sound when experiencing virtual reality or mixed reality. [Background technology]
[0002] One type of display device is a head-mounted display (HMD), which is worn on the head to allow viewers to enjoy images with a sense of presence. HMDs can be used in VR (Virtual Reality), which displays only images from a virtual reality world, or in MR (Mixed Reality), which combines images from a virtual reality world with images from the real world around the viewer.
[0003] There are also sound output devices that can adjust the intensity of external sounds, such as headphones and earphones with a noise cancellation function. Hereinafter, noise cancellation will be referred to as "NC," and headphones and earphones with an NC function will be referred to as "NC earphones." NC earphones can also block out most external sounds. For example, when NC earphones are used in VR applications, blocking out external sounds from the surroundings makes it easier for viewers to achieve a sense of immersion. On the other hand, in MR applications, depending on the content displayed on the HMD, it may be desirable to be able to hear sounds in the surrounding real space.
[0004] Patent Document 1 describes a method that enables a viewer whose visual and auditory information from the surroundings is limited by an HMD and NC earphones to detect speech from people around them, and adjusts the sound so that external sounds can be heard when a speech is detected.By adjusting the sound so that external sounds can be heard, the viewer can easily respond to speech from people around them. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Application Laid-Open No. 2017-69687 Summary of the Invention [Problem to be solved by the invention]
[0006] However, with the technology described in Patent Document 1, if a talking action is not detected during viewing, external sounds cannot be heard.
[0007] Therefore, an object of the present invention is to enable a viewer to hear external sounds when it is preferable to do so, even when experiencing virtual reality or mixed reality. [Means for solving the problem]
[0008] The information processing device of the present invention is corresponds to Real and virtual images of Get image Acquisition means; The external sound in the real space and the virtual image are associated with each other. Virtual Sound and Get sound Acquisition means; The external sound and the virtual image are detected based on at least one of the real image and the virtual image. The virtual sound of an adjusting means for adjusting the relative volume between the and an output means for outputting a composite image obtained by combining the real image and the virtual image based on mixing information to a display device provided in the head-mounted device, wherein the adjustment means generates a composite sound by combining the virtual sound and the external sound based on at least one of the real image and the virtual image, the real image being an image obtained by capturing an image of the real space using an imaging device mounted on the head-mounted device, the mixing information including transparency information that determines a transparency of an image based on the virtual image, the adjustment means performing the adjustment according to the transparency information, the mixing information including the transparency information such that the transparency decreases as the amount of movement of the head-mounted device in the real space decreases, and the adjustment means performs control such that the amount of reduction of the external sound increases as the transparency of the transparency information decreases. It is characterized by: [Effects of the Invention]
[0009] According to the present invention, even when a viewer is experiencing virtual reality or mixed reality, the viewer can hear external sounds when it is desirable to do so. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 2 is a diagram illustrating an example of a hardware configuration of an information processing device. [Figure 2] FIG. 1 is a diagram showing an example of how an HMD is used for MR purposes. [Figure 3] FIG. 2 is a diagram illustrating a functional configuration of an information processing device according to the first embodiment. [Figure 4] 4 is a flowchart of information processing in the first embodiment. [Figure 5] 10 is a flowchart of information processing in the second embodiment. [Figure 6] 10 is a detailed flowchart from spatial information acquisition to external sound synthesis. [Figure 7] 10 is a flowchart of information processing in the third embodiment. [Figure 8] 10 is a flowchart of information processing in the fourth embodiment. [Figure 9] 13 is a flowchart of information processing in the fifth embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, embodiments of the present invention will be described with reference to the drawings. The following embodiments do not limit the present invention, and not all of the combinations of features described in the present embodiments are necessarily essential to the solution of the present invention. The configurations of the embodiments may be modified or changed as appropriate depending on the specifications of the device to which the present invention is applied and various conditions (such as usage conditions and usage environment). Furthermore, the present invention may be configured by appropriately combining parts of each of the embodiments described below. In the following embodiments, the same configurations and processes are described with the same reference symbols.
[0012] <Hardware configuration of information processing device according to first embodiment> The hardware configuration of the information processing apparatus of the first embodiment will be described below with reference to FIG. The CPU 101 uses the RAM 102 as a work memory to execute programs stored in the ROM 103 or hard disk drive (HDD) 105, and controls the operation of each block (described later) via a system bus 114. The programs executed by the CPU 101 include an information processing program according to this embodiment (described later). The HDD interface (hereinafter, interface will be referred to as "I / F") 104 connects a secondary storage device such as the HDD 105 or an optical disk drive. The HDD I / F 104 is, for example, an I / F such as a serial ATA (SATA). The CPU 101 can read data from the HDD 105 and write data to the HDD 105 via the HDD I / F 104. Furthermore, the CPU 101 can load data stored in the HDD 105 into the RAM 102, and conversely, can save data loaded in the RAM 102 to the HDD 105. The CPU 101 can then execute programs loaded into the RAM 102.
[0013] The input I / F 106 connects to an input device 107 such as a keyboard, a mouse, a digital camera, a scanner, or an acceleration sensor. The input I / F 106 can also connect a stereo camera provided in an HMD (head-mounted display), which is a head-mounted device, as the input device 107. The input I / F 106 is, for example, a serial bus I / F such as USB or IEEE 1394. The CPU 101 can read data from the input device 107 via the input I / F 106. The output I / F 108 connects the information processing device 100 to a display device 109 serving as an HMD. The output I / F 108 is, for example, an image output I / F such as DVI or HDMI (registered trademark). The CPU 101 can display images of the virtual reality world, images of the mixed reality world, etc. on the HMD by sending image data of the virtual reality world or image data of the mixed reality world (described later) to the HMD via the output I / F 108.
[0014] The sound input I / F 110 connects a sound input device 111 capable of collecting sound, such as a microphone or a directional microphone. The sound input I / F 110 is a serial bus I / F such as USB or IEEE1394. The sound output I / F 112 connects a sound output device 113 that outputs sound, such as a headphone device or a speaker. The connection between the sound output I / F 112 and the sound output device 113 may be a wired connection or a wireless connection. The sound output device 113 may be configured separately from the information processing device 100, and the information processing device 100 can control the sound output device 113 by controlling it or by transmitting a control signal to it. This allows the CPU 101 of the information processing device to integrally control the HMD and the sound output device 113 (such as a headphone device). The sound output I / F 110 is a serial bus I / F such as USB or IEEE1394. The CPU 101 sends sound data from the virtual reality world (described later) and synthesized sound data that combines virtual sound data from the virtual reality world with external sound data from the real world around the viewer to a headphone device or the like via the sound output I / F 112, and outputs these sounds from the headphone device or the like.
[0015] The information processing device 100 may not include the HDD 105 or the display device 109. The information processing device 100 may or may not include an HMD. If the information processing device 100 does not include an HMD, the information processing device 100 receives data from the HMD and sends data to the HMD by connecting to an external HMD via the input I / F 106. The components of the information processing device 100 may include configurations other than the components described above, but illustration and description of those components will be omitted here.
[0016] Before describing the detailed operation and processing of the information processing device 100, an example will be described using FIG. 2 in which an image of a mixed reality (MR) world in which an image of a virtual reality world is synthesized with an image of the surrounding real space is provided to a viewer.
[0017] As shown in FIG. 2(A), it is assumed that a viewer wears an HMD 1, which is a head-mounted device, and a headphone device 2. The HMD 1 is equipped with a camera (imaging device) (not shown), which captures real image data of the real space around the viewer. The HMD 1 is also equipped with a microphone, which is capable of capturing external sound data, which is sound in the real space around the viewer. Note that the microphone does not have to be equipped in the HMD 1; in that case, the microphone is connected to the HMD or the information processing device 100 via a wired or wireless connection.
[0018] Image 3 in FIG. 2(B) shows a real image (referred to as real image 3) captured by the HMD camera of a real space. The HMD 1 also acquires a virtual image 4. The virtual image 4 is a virtual reality image, hereafter referred to as a VR image 4. The HMD 1 then displays a mixed reality image 5 (referred to as MR image 5) that is a combination of the real image 3 and the VR image 4 on its built-in display as a display image. This allows a viewer wearing the HMD 1 to view the MR image 5 that is a combination of the real image 3 and the VR image 4, that is, to experience mixed reality. Note that FIG. 2 shows an example of an MR application in which the MR image 5 is displayed on the HMD 1. However, when the HMD 1 is used for virtual reality (VR) applications, only a VR image of the virtual reality world is displayed on the HMD 1 as a display image.
[0019] The headphone device 2 is a sound output device that the viewer wears on their ears. The sound output device may be an earphone device. When the HMD 1 is used to display only images of a virtual reality world (VR images), the viewer can be more easily immersed in the virtual reality world by outputting virtual sound data, that is, sound of the virtual reality world (hereinafter referred to as VR sound), from the headphone device 2. In particular, when the HMD 1 is used to display only VR images, performing information processing to block external sounds using, for example, a noise cancellation function (NC function) can easily make the viewer feel immersed in the virtual reality world. On the other hand, when the HMD 1 is used to combine images of the real world (actual images) around the viewer with images of the virtual reality world (VR images), it is often desirable to be able to hear external sounds from the surrounding real space, depending on the content of the display image displayed on the HMD 1.
[0020] Below, we will explain the first embodiment of an information processing device that allows viewers to easily feel immersed in a virtual reality world when displaying only VR images, and that allows viewers to hear external sounds from the surrounding real space in accordance with the display of the HMD1 when displaying VR images and real images.
[0021] <Functional configuration of the information processing device according to the first embodiment> FIG. 3 is a functional block diagram showing the functional configuration of an information processing device 100 according to the first embodiment. As shown in FIG. 3, the information processing device 100 includes a VR image acquisition unit 11, a mixed information acquisition unit 12, a real image acquisition unit 13, an output image acquisition unit 14, a mixed ratio calculation unit 15, a VR sound acquisition unit 16, an external sound acquisition unit 17, a sound synthesis unit 18, and an output unit 19. These functional units are implemented by the CPU 101 executing an information processing program according to this embodiment. However, some or all of the functional units may be implemented by hardware configurations such as circuits. In the following description, unless otherwise specified, image data handled by the information processing device 100 will be referred to simply as "image," and sound data handled by the information processing device 100 will be referred to simply as "sound" (or "external sound" in the case of external sound data). The information processing device 100 will also be referred to simply as the information processing device.
[0022] The VR image acquisition unit 11 is a virtual image acquisition unit that acquires a VR image, which is virtual image data generated by rendering, or a VR image that has been prepared in advance. Note that the VR image acquisition unit 11 is a functional unit configured by the CPU 101 executing the information processing program of this embodiment, and therefore rendering of the VR image is executed by the CPU 101, but may also be performed by, for example, a GPU (not shown). The VR image acquisition unit 11 outputs the acquired VR image to the output image acquisition unit 14. Details of the operation of the VR image acquisition unit 11 will be described later.
[0023] The real image acquisition unit 13 acquires a real image captured by a camera (image capture device) of a real space. Then, the real image acquisition unit 13 outputs the real image to the output image acquisition unit 14. In this embodiment, the camera is included in the input device 107.
[0024] The mixture information acquisition unit 12 outputs the mixture information to the output image acquisition unit 14 and the mixture ratio calculation unit 15. In this embodiment, mask image data (hereinafter referred to as a mask image) is used as the mixture information. The mask image is used to specify the range of the VR image to be superimposed on the actual image captured by a camera of the real space. The mixture information and the mask image in this embodiment will be described in detail later.
[0025] The mixture ratio calculation unit 15 is a ratio acquisition unit that calculates the mixture ratio of the real image and the VR image based on the mixture information acquired by the mixture information acquisition unit 12. The mixture ratio calculation unit 15 outputs information on the calculated mixture ratio to the sound synthesis unit 18. Examples of the mixture ratio and other details will be described later.
[0026] The output image acquisition unit 14 combines the VR image and the real image based on the mixed information acquired by the mixed information acquisition unit 12, and acquires the combined image data (referred to as a combined image) as an output image. Note that since the output image acquisition unit 14 is a functional unit configured by the CPU 101 executing an information processing program, generation of the combined image is executed by the CPU 101, but may also be performed by, for example, a GPU (not shown). Then, the output image acquisition unit 14 outputs the output image to the output unit 19.
[0027] The VR sound acquisition unit 16 is a virtual sound acquisition unit that generates VR sound, which is virtual reality sound data, or acquires VR sound that is prepared in advance and associated with a VR image. Note that since the VR sound acquisition unit 16 is a functional unit configured by the CPU 101 executing an information processing program, the generation of VR sound is executed by the CPU 101, but may also be executed by, for example, a GPU (not shown). The VR sound acquisition unit 16 then outputs the VR sound to the sound synthesis unit 18. Details of the operation of the VR sound acquisition unit 16 will be described later.
[0028] The external sound acquisition unit 17 acquires external sounds in the surrounding real space using a sound input device 111 such as a microphone. Then, the external sound acquisition unit 17 outputs the acquired external sounds to the sound synthesis unit 18. The sound synthesis unit 18 adjusts the relative volume between the external sound and the VR sound based on the information about the mixing ratio, and outputs the adjusted sound to the output unit 19 at the downstream side. As will be described in detail later, the sound synthesis unit 18 performs a process of adjusting the external sound so that the viewer can hear it, or a process of adjusting the external sound to reduce it, based on the information about the mixing ratio. Sound adjustment based on the mixing ratio includes generating a synthesized sound by directly synthesizing the external sound with the VR sound, generating a synthesized sound by synthesizing the external sound with the VR sound after adjusting the external sound to reduce it, and adjusting the volume of the VR sound side (lowering the volume) so that the external sound in the real space can be heard. That is, the sound synthesis unit 18 adjusts the relative volume between the external sound and the VR sound (hereinafter referred to as sound adjustment) based on the information about the mixing ratio, and outputs the adjusted sound to the output unit 19 at the downstream side. Adjustment to reduce external sound includes adjustment to almost completely block out external sound by so-called noise canceling. Furthermore, when an adjustment is made based on the mixing ratio, for example, to lower the volume of the VR sound, the volume of the external sound becomes relatively louder compared to the VR sound, making it easier for the user to hear the external sound.
[0029] The output unit 19 outputs the output image from the output image acquisition unit 14 to the display device 109 (HMD 1 in this embodiment). The output unit 19 also outputs the synthesized sound from the sound synthesis unit 18 to the sound output device 113 (headphone device 2 in this embodiment).
[0030] The data such as the VR image, mask image (mixing information), mixing ratio, output image, and VR sound may be acquired in advance and stored in the HDD 105 or the like, and the information processing device may appropriately read them from the HDD 105. The information processing device may also appropriately acquire the data of the VR image, mask image, output image, mixing ratio, and VR sound from a cloud (not shown) via a communication device (not shown) or the like.
[0031] In this embodiment, the VR image, mask image, actual image, and output image are all assumed to have the same resolution, but they do not necessarily have to be the same resolution. For example, each of these images may be held at a different resolution, and the information processing device may adjust the resolution by scaling during calculation. Furthermore, the VR sound and external sound are assumed to be sound data with the same sample rate, but they do not necessarily have to be the same sample rate; they may be synthesized after being adjusted to the same sample rate by interpolation or resampling.
[0032] Furthermore, the VR image, real image, and output image are assumed to be three-channel image data of the three primary colors RGB, but they do not necessarily have to be three channels; for example, they may be one-channel monochrome image data, or five channels including color differences. Furthermore, although the external sound, VR sound, and synthesized sound are assumed to be one-channel monaural data, they do not necessarily have to be one-channel, and may be two-channel stereo data or five-channel or other three-dimensional sound data.
[0033] <Information Processing According to the First Embodiment> 4 is a flowchart showing the flow of information processing executed by the information processing apparatus of the first embodiment. Hereinafter, each processing step (process) will be indicated by adding "S" before the reference numeral, and the description of the processing step will be omitted. In S101, the VR image acquisition unit 11 acquires a VR image. As described above, a VR image is virtual reality image data, which is an image that represents something different from the real world around the viewer. The VR image is, for example, an image created by rendering three-dimensional model data from a virtual viewpoint using a known rendering technology, or an image created by capturing a real world, such as a place or time different from the real world around the viewer, with a camera or the like.
[0034] Next, in S102, the VR sound acquisition unit 16 acquires VR sound. As described above, VR sound is sound data of the virtual reality world. VR sound is generated as sound heard at a listening point (virtual viewpoint) based on the positional relationship, including the distance and direction, between the position of a virtual sound source set in the three-dimensional model data and a virtual listening point (virtual viewpoint in this embodiment). Furthermore, when an image captured in a situation different from the real world around the viewer is used as a VR image, the VR sound is, for example, sound acquired by collecting ambient external sounds in that situation using a microphone or the like. Alternatively, sounds such as background music may be used directly as VR sound.
[0035] Next, in S103, the mixed information acquisition unit 12 acquires a mask image as mixed information. The mask image is an image in which the pixel values of pixel positions to be superimposed on the actual image as a VR image are set to 1, and the pixel values of other pixel positions are set to 0. In other words, in the mask image, an image region with a pixel value of 1 is set to a region corresponding to the VR image, and an image region with a pixel value of 0 is set to a region corresponding to an image other than the VR image.
[0036] Next, in S104, the mixing ratio calculation unit 15 calculates a mixing ratio based on the mixing information acquired by the mixing information acquisition unit 12 in S103. Then, the mixing ratio calculation unit 15 sets the value of a predetermined flag based on the mixing ratio. For example, if the mixing ratio has a ratio corresponding to the real image, that is, if the mask image has an image area corresponding to the real image, the mixing ratio calculation unit 15 sets a predetermined flag indicating whether or not to synthesize the real image to 1. On the other hand, if the mixing ratio does not have a ratio corresponding to the real image, that is, if the mask image has no image area corresponding to the real image, the mixing ratio calculation unit 15 sets the predetermined flag to 0. In other words, if the mask image has an image area other than the VR image, the mixing ratio calculation unit 15 sets the predetermined flag to 1, and if all image areas are areas corresponding to the VR image, the mixing ratio calculation unit 15 sets the predetermined flag to 0. In the following description, the predetermined flag will be referred to as an MR flag.
[0037] Next, in S105, the information processing device determines whether the MR flag is 1 or 0. If the MR flag is 1, the processing of the information processing device proceeds to S106 and subsequent steps, whereas if the MR flag is 0, the processing of the information processing device proceeds to S108 and subsequent steps.
[0038] When it is determined in S105 that the MR flag is 0 and the process proceeds to S108, the external sound acquisition unit 17 acquires external sounds in the surrounding real space from the sound input device 111 such as a microphone. Next, in S109, the sound synthesis unit 18 synthesizes the external sound acquired by the external sound acquisition unit 17 and the VR sound acquired by the VR sound acquisition unit 16, and performs sound adjustment processing to reduce the external sound. In this embodiment, a known noise canceling method is used as a method for reducing the external sound. That is, in the external sound synthesis processing in S109, a sound adjustment processing is performed to substantially block the external sound by adding a sound that is in opposite phase to the external sound acquired in S108, and then the sound is synthesized with the VR sound. As a result, the external sound and the opposite phase sound cancel each other out, and the external sound is substantially blocked. Then, the sound synthesis unit 18 outputs the sound after the sound adjustment processing, i.e., the synthesized sound in which the external sound is blocked and substantially only the VR sound remains, to the output unit 19.
[0039] After processing S109, the process proceeds to S110, where the output unit 19 outputs the output image to the display device 109, i.e., the HMD, and outputs the synthesized sound to the sound output device 113, i.e., the headphone device. In other words, if the MR flag is determined to be 0 in S105 and processing from S108 onwards is performed, and then the process proceeds to S110, an output image consisting of only the VR image is displayed on the HMD, and synthesized sound consisting of only the VR sound after external sounds have been blocked out is output from the headphone device. Note that it is not necessary for only the VR sound to be output, as long as the VR sound is relatively stronger than other sounds.
[0040] On the other hand, if it is determined in S105 that the MR flag is 1 and the process proceeds to S106, the real image acquisition unit 13 acquires a real image of the real space captured by a camera, and the external sound acquisition unit 17 acquires external sound of the real space collected by a sound input device 111 such as a microphone.
[0041] Next, in S107, the output image acquisition unit 14 generates a composite image as an output image by combining the real image acquired by the real image acquisition unit 13 and the VR image acquired by the VR image acquisition unit 11. That is, the output image at this time is an MR image in which the real image and the VR image are combined. Here, when combining the real image and the VR image, the output image acquisition unit 14 acquires a mask image from the mixture information acquisition unit 12 and multiplies the VR image by the mask image for each pixel position. Furthermore, the output image acquisition unit 14 adds the multiplication result to each pixel position of the real image. As a result, the VR image within the range specified by the mask image to be superimposed is superimposed on the real image. At this time, the output image acquisition unit 14 may perform a known blending process to smooth the transition at the boundary between the real image and the VR image.
[0042] Also, in S107, the sound synthesis unit 18 generates, as output sound, a synthesized sound obtained by synthesizing the external sound acquired by the external sound acquisition unit 17 and the VR sound acquired by the VR sound acquisition unit 16. That is, the output sound at this time is a sound (hereinafter referred to as MR sound) obtained by synthesizing the collected external sound as it is with the VR sound. In this way, if the MR flag is 1 in S105, the sound adjustment process of S109 to reduce the external sound is not performed, and the MR sound obtained by synthesizing the external sound with the VR sound is output. Alternatively, in S107, the sound synthesis unit 18 may perform sound adjustment, such as amplifying the collected external sound, to synthesize it with the VR sound. That is, as long as it is possible to make both the external sound and the VR sound audible, the form of sound synthesis is not limited.
[0043] After processing S107, when the process proceeds to S110, the output unit 19 outputs an output image (MR image) in which the VR image and the real image are combined to the display device 109 of the HMD. Furthermore, when the process proceeds from S107 to S110, the output unit 19 outputs an output sound (MR sound) in which the VR sound and the external sound are combined to the headphone device worn by the user. That is, when the MR flag is determined to be 1 in S105 and the process from S106 onwards is performed, and then the process proceeds to S110, a combined sound including not only the VR sound but also the external sound is output from the headphone device. This allows the viewer to hear the external sound of the surrounding real space while listening to the VR sound output from the headphone device.
[0044] Thereafter, in S111, the information processing device determines whether or not to terminate the processing of this flowchart, and if not, returns the processing to S101. For example, if the viewer is watching a video, processing corresponding to the next frame is performed in S101. On the other hand, if it is determined that the processing should be terminated, for example, because the viewer issues an instruction to terminate, the information processing device terminates the processing of this flowchart.
[0045] As described above, according to the information processing device of the first embodiment, when only VR images are displayed on the HMD, external sounds can be largely blocked, allowing the viewer to easily immerse themselves in the virtual reality world. On the other hand, when an image obtained by combining VR images and real images is displayed on the MMD, external sounds can be prevented from being reduced, allowing the viewer to hear surrounding external sounds. Furthermore, the information processing device of the first embodiment automatically switches between blocking external sounds when only VR images are displayed on the HMD and when an image obtained by combining VR images with real images is displayed. Therefore, according to the first embodiment, the viewer does not need to manually switch between blocking external sounds and not blocking them, which saves the viewer time and effort.
[0046] 4 illustrates an example in which, when a composite image of a VR image and a real image is displayed, the viewer can hear both VR sound and external sound, whereas, when only a VR image is displayed, external sound is mostly blocked by noise cancellation, allowing the viewer to mostly hear only the VR sound. This embodiment is not limited to this example. For example, if an earphone device or the like is used in which external sound reaches the viewer's ears when the noise cancellation function is off, the external sound acquisition process of S106 and the external sound synthesis process of S107 in FIG. 4 do not need to be performed. In this example, when a composite image of a VR image and a real image is displayed, the viewer can hear external sounds in the surrounding real space while listening to the VR sound output from the earphone device.
[0047] In the first embodiment, the mixture information acquisition unit 12 acquires a mask image as the mixture information in S103. However, this embodiment is not limited to this example. For example, information other than a mask image may be used as long as it is possible to determine whether the HMD is for VR use, in which only VR images are displayed on the HMD, or for MR use, in which a composite image of a VR image and a real image is displayed. For example, if viewing application programs operated on the HMD are divided into those for VR and those for MR and the uses can be identified, the mixture information acquisition unit 12 may acquire, as the mixture information, information indicating which type of application program the application program is for. In this case, the mixture information acquisition unit 12 may acquire, as the mixture information, information indicating which type of application program the application program is for. In this case, in S105, the mixture ratio calculation unit 15 sets the MR flag to 0 if the application program is for VR use, and sets the MR flag to 1 if the application program is for MR use.
[0048] In the first embodiment, an example was described in which external sound was almost completely blocked by noise canceling, which adds sound in antiphase to the external sound. However, external sound may be blocked by a method other than noise canceling. For example, the viewer's ears may be covered with an object made of a material that does not easily transmit external sound. Then, when the MR flag is 1, the external sound may be collected by a microphone or the like in S106, and the external sound may be synthesized with VR sound in S107. In this example, the processes of S108 and S109 in the flowchart of FIG. 4 are unnecessary.
[0049] In the first embodiment, an example was described in which external sound was reduced when only a VR image was displayed. However, it is not necessary to reduce external sound in this case; sound adjustment may be performed so that the viewer can hear the external sound. This is done in consideration of the fact that, for example, a viewer may want to understand the surrounding situation when they move. In this case, for example, in S103, the mixed information acquisition unit 12 acquires the mixed information, including the value of an acceleration sensor mounted on the HMD. Then, in S104, the mixed ratio calculation unit 15 calculates the mixed ratio taking the acceleration sensor value into account. This enables, for example, in S109, the sound synthesis unit 18 to synthesize external sound into the VR sound based on the acceleration sensor value. That is, for example, when the sound synthesis unit 18 determines that the viewer has moved based on the acceleration sensor value, it generates synthesized sound without making adjustments to reduce the external sound. This allows the viewer to understand the surrounding situation from the external sound when they move.
[0050] In addition, determining whether the viewer is moving can be made possible by, for example, previously learning a classifier that identifies whether the viewer is moving based on the acceleration sensor value using a known machine learning process. In this case, the mixing ratio calculation unit 15 uses the classifier to calculate a mixing ratio by including a flag indicating whether the viewer is moving. Then, when the sound synthesis unit 18 determines from the flag included in the mixing ratio that the viewer is moving, it does not synthesize a sound that is out of phase with the external sound into the VR sound. On the other hand, when the sound synthesis unit 18 determines from the flag included in the mixing ratio that the viewer is not moving, it synthesizes a sound that is out of phase with the external sound into the VR sound, thereby reducing the external sound. Note that using the acceleration sensor value as a method for determining whether the viewer is moving may not be essential, and other methods may be used as long as they can detect the viewer's movement. For example, a real-world area that the viewer will be viewing may be set in advance, and exiting that area may be detected using known sensing technology. Then, when the viewer leaves the set real-world area, it may not reduce the external sound, or conversely, it may reduce the external sound.
[0051] The information processing device of this embodiment handles image data and sound data, but the sampling rate of the image data is generally different from the sampling rate of the sound data, with the latter often being higher. In this case, the processing of the sound data in S108 and S109 and the processing of the sound data in S106 and S107 may be looped until the next image data is sampled. Furthermore, the processing of the sound data and the processing of the image data may be performed in different threads. In this case, the sound synthesis unit 18, which operates in a thread separate from the processing related to the image data, acquires the MR flag, and if the MR flag is 1, does not reduce the external sound, but if the MR flag is 0, reduces the external sound.
[0052] <Second embodiment> In the first embodiment, an example has been described in which the MR flag is determined to be 1 or 0 based on a mask image that is mixture information, and whether or not to adjust the external sound is switched depending on the value of the MR flag. In the following second embodiment, an example in which spatial information is included in the mixed information will be described. In the second embodiment, the spatial information is distance information from the viewer. The distance information in the second embodiment is information that indicates the distance from the viewer to a sound source in real space. There are cases in which the viewer wants to hear external sounds from a nearby sound source in real space. An example of such a situation is when the viewer is working while viewing a VR image. In such an assumed example, it is often desirable for the viewer to hear external sounds from a sound source that is close to the viewer.
[0053] In order to be able to handle such situations, the information processing device of the second embodiment includes distance information as spatial information in the mixed information, and performs sound adjustment based on the distance information. Note that the hardware configuration and functional configuration of the information processing device in the second embodiment are the same as those in the first embodiment. In the second embodiment, the same functional configurations and processing steps as those in the first embodiment are assigned the same reference numerals and their explanations are omitted, and the following description will mainly focus on the parts that are different from the first embodiment. The functional configuration of the information processing device according to the second embodiment is the same as that shown in FIG. 3 described above. However, in the second embodiment, the mixture information acquisition unit 12 outputs mixture information including spatial information to the output image acquisition unit 14 and the mixture ratio calculation unit 15.
[0054] <Information Processing of the Second Embodiment> 5 is a flowchart showing the flow of information processing executed by an information processing device according to the second embodiment. In the case of the second embodiment, if the MR flag is 0 in S105, the information processing device proceeds to processing from S108 onwards, whereas if the MR flag is 1, the information processing device proceeds to processing from S201 onwards. After processing from S201 to S206, the information processing device proceeds to processing of S110.
[0055] In the second embodiment, when the MR flag is 1 in S105 and the process proceeds to S201, the real image acquisition unit 13 acquires a real image of the real space captured by the camera. Next, in S202, the output image acquisition unit 14 generates a composite image by combining the real image acquired by the real image acquisition unit 13 and the VR image acquired by the VR image acquisition unit 11. The process of combining the real image and the VR image is the same as the image combination process in S107 described above.
[0056] Next, in S203, the mixed information acquisition unit 12 acquires distance information as spatial information. FIG. 6 is a flowchart showing details of the process of S203 when acquiring distance information, the subsequent process of S204, and the processes of S206. The detailed process of S203 will be described with reference to FIG.
[0057] In S2001, the mixed information acquisition unit 12 acquires distance setting information. The distance setting is information set by, for example, a viewer or a system setter as a distance at which it is desirable to hear external sounds. In this embodiment, an example is given in which the viewer sets the distance. In this example, the information processing device displays a user interface (UI) screen on, for example, the HMD, and the mixed information acquisition unit 12 acquires, as distance setting information, a distance threshold Dth arbitrarily set by the viewer via the input device 107. The distance threshold Dth set by the viewer as the distance setting corresponds to the distance from the HMD to an object in the real world. The UI screen used for distance setting displays objects whose distance from the HMD is less than the distance threshold (less than Dth), and does not display objects whose distance is equal to or greater than the distance threshold (Dth or greater). The viewer sets the distance at which it is desirable to hear external sounds while looking at the UI screen. Note that the method of setting the distance and the definition of the distance are not limited to setting using the UI. In the case where the sound source is an object existing in real space, such as an object on a desk, the distance from the HMD to the desk may be measured using a known technique such as stereo matching, and the setting may be based on the measured distance. Furthermore, the definition of the distance does not necessarily have to be the distance from the HMD, and may be, for example, the distance from the center of gravity of the viewer.
[0058] Next, in S2002, the mixed information acquisition unit 12 acquires depth information from the real image. The depth information is information that indicates the depth to an object for each pixel in an object captured in the real image, and can be obtained using a known technique such as stereo matching. The mixed information acquisition unit 12 may also obtain the depth by using a range finder such as Lidar (Light Detection and Ranging or Laser Imaging Detection and Ranging).
[0059] Next, the process proceeds to S204, where the mixture ratio calculation unit 15 calculates the mixture ratio. The mixture ratio in the second embodiment is different from that in the first embodiment. The detailed process of S204 will be described with reference to FIG. In S2003, the mixture ratio calculation unit 15 calculates depth threshold information. The depth threshold information is information that stores, for each pixel, whether the depth to an object corresponding to each pixel of the real image is less than the distance threshold (less than Dth) or greater than or equal to the distance threshold (greater than Dth). The mixing ratio calculation unit 15 records the depth of each pixel of the object in the actual image acquired in S2002 as 1 if it is less than the distance threshold Dth, or 0 if it is greater than or equal to the distance threshold Dth, at each pixel position corresponding to the depth information.
[0060] Next, in S2004, the mixing ratio calculation unit 15 calculates the connected components of the depth threshold information. The mixing ratio calculation unit 15 treats adjacent positions in the depth threshold information as one connected component and stores the connected components in a list. As will be described later, the list of connected components represents a mixing ratio where external sounds in the area indicated by the connected component are synthesized and external sounds in other areas are not synthesized (external sounds are reduced). Next, in S205, the external sound acquisition unit 17 acquires external sound, and then the processing of the information processing device proceeds to S206. Note that the external sound acquisition processing in the external sound acquisition unit 17 is the same as in S108 described above.
[0061] In S206, the sound synthesis unit 18 synthesizes an external sound. The sound synthesis process in S206 differs from the external sound synthesis process in the first embodiment described above. The detailed process of S206 will be described with reference to FIG. In S2005, the sound synthesis unit 18 refers to the list calculated in S2004 and determines whether there are any connected components for which external sound has not yet been synthesized. If there are any connected components for which external sound has not yet been acquired, the sound synthesis unit 18 selects one from the list and performs the subsequent process of S2006 on that connected component. In other words, the process of S206 is a loop process, and if there are no connected components in the list for which external sound has not yet been synthesized, the process of the information processing device proceeds to S110.
[0062] Proceeding to S2006, the sound synthesis unit 18 acquires the external sound at the selected connected component position from the external sound acquisition unit 17. Note that, as a method for acquiring the external sound at the connected component position, for example, a method can be used in which a directional microphone is used as the microphone and the external sound acquisition unit 17 acquires the external sound from the directional microphone. Alternatively, for example, a method can be used in which a plurality of microphones are installed, the sound generation position is estimated by a known sound generation source estimation method that estimates the sound generation position (sound source) from the phase shift between the microphones, and if there is sound at the position indicated by the connected component, the sound at that position is acquired.
[0063] In the second embodiment, when external sounds are synthesized in S206, it is not necessary to synthesize the external sounds at their original intensity, and for example, in S204, the intensity may be adjusted using a gain based on distance instead of depth threshold information, and the external sounds may be synthesized at the intensity adjusted based on that gain. In this case, the connected components may be found by calculating the connected components of pixels whose values are not 0.
[0064] In the second embodiment, in an application in which only VR images are displayed on an HMD, external sounds are almost completely blocked by noise canceling, as in the example of the first embodiment described above, but it is also possible to physically cover the ears of the viewer so that the viewer cannot hear external sounds. In this case, the external sound acquisition process of S108 and the sound adjustment and external sound synthesis process of S109 in the flowchart of FIG. 5 are not necessary.
[0065] As described above, according to the information processing device of the second embodiment, distance information as spatial information is included in the mixed information, and sound adjustment can be performed based on the distance information. That is, in the case of the second embodiment, the mixed ratio calculation unit 15 calculates a mixed ratio that reduces external sounds from sound sources that are equal to or greater than the distance threshold in real space, and on the other hand, calculates a mixed ratio that does not reduce external sounds from sound sources that are less than the distance threshold. This allows the viewer to hear nearby sounds in real space. Furthermore, in the second embodiment, as in the first embodiment described above, the viewer can be spared the trouble of adjusting external sounds themselves.
[0066] <Third embodiment> In the second embodiment, an example has been described in which distance information is included as spatial information in the mixed information. In the third embodiment, an example will be described in which transparency information of a VR image is included in the mixed information. For example, even when a VR image is being displayed, a viewer may want to temporarily check information about their surroundings. As an example, when the viewer moves, they may come into contact with another person approaching. In such cases, it is preferable for the viewer to be able to check the situation around them. Therefore, the information processing device of the third embodiment includes transparency information of a VR image in the mixed information and adjusts external sounds based on the transparency information. The hardware configuration and functional configuration of the information processing device in the third embodiment are the same as those in the first embodiment. In the third embodiment, the same functional configurations and processing steps as those in the first and second embodiments are denoted by the same reference numerals and will not be described again, and differences will be mainly described below.
[0067] <Functional configuration of information processing device according to third embodiment> The functional configuration of the information processing device according to the third embodiment is the same as that shown in Fig. 3. In the third embodiment, the mixture information acquisition unit 12 also outputs the mixture information to the output image acquisition unit 14 and the mixture ratio calculation unit 15. In the case of the third embodiment, the mixture information includes transparency information of the VR image. <Information Processing of the Third Embodiment> 7 is a flowchart showing the flow of information processing executed by the information processing apparatus of the third embodiment. In the information processing apparatus of the third embodiment, if the MR flag is 1 in S105, the processing proceeds to S301 and subsequent steps, whereas if the MR flag is 0, the processing proceeds to S108 and subsequent steps. After the processing from S301 to S306, the information processing apparatus proceeds to the processing of S110.
[0068] When it is determined that the MR flag is 1 and the process proceeds to S301, the mixture information acquisition unit 12 acquires transparency information included in the mixture information. The transparency information is, for example, a transparency value set by a viewer or a system setter. In the present embodiment, the mixture information acquisition unit 12 acquires, as the transparency information, the transparency value set by the viewer via, for example, the input device 107. Note that in this embodiment, transparency Tv is used as the transparency information, and transparency Tv is a value between 0 and 1.
[0069] Next, proceeding to S302, the mixing ratio calculation unit 15 calculates the mixing ratio based on the transparency information (transparency Tv). The mixing ratio in the third embodiment uses the transparency Tv as it is. Note that the transparency Tv does not necessarily have to be the mixing ratio. For example, the mixing ratio may be the transparency Tv raised to a power. In this case, the information processing device performs the composition described below by treating the transparency Tv raised to a power as if it were substituted for the transparency Tv.
[0070] Next, in S303, the real image acquisition unit 13 acquires a real image obtained by capturing the real space around the viewer. Next, in S304, the output image acquisition unit 14 combines the VR image acquired by the VR image acquisition unit 11 and the real image acquired by the real image acquisition unit 13 according to the mixture ratio calculated in S302. At this time, the output image acquisition unit 14 multiplies the RGB values of the pixels of the real image by the transparency Tv. The output image acquisition unit 14 also multiplies the RGB values of the pixels of the VR image by (1-Tv). Thereafter, the output image acquisition unit 14 combines the real image and VR image after the multiplication by adding them together to generate a combined image.
[0071] Next, in S305, the external sound acquisition unit 17 acquires external sound. Next, in S306, the sound synthesis unit 18 synthesizes the external sound acquired by the external sound acquisition unit 17 and the VR sound acquired by the VR sound acquisition unit 16 to generate a synthesized sound. At this time, the sound synthesis unit 18 performs sound adjustment such that the intensity of the external sound is multiplied by the transparency Tv, and synthesizes the adjusted external sound with the VR sound. Note that when adjusting the intensity of the external sound, it is not necessary to multiply by Tv, and another predetermined constant may be multiplied. Furthermore, the sound synthesis unit 18 may, for example, determine the maximum value of the sound adjustment based on the transparency Tv, and apply a gain so that this maximum value becomes the maximum volume of the external sound.
[0072] As described above, according to the information processing device of the third embodiment, external sound can be adjusted based on transparency information included in the mixing information. That is, in the case of the third embodiment, the mixing ratio calculation unit 15 acquires a mixing ratio according to the transparency information. As a result, in the case of the third embodiment, the viewer can check the surrounding external sound. Furthermore, in this embodiment, the external sound is adjusted according to the transparency of the VR image, which reduces the effort required of the viewer to adjust the external sound.
[0073] In the third embodiment, an example has been given in which the viewer or a system setter sets the transparency information, but this is not limiting. For example, the mixing information acquisition unit 12 may acquire the amount of movement of the viewer while he or she is moving, i.e., the amount of movement of the HMD, and automatically set the transparency level according to the acquired amount of movement. Here, the amount of movement of the viewer while he or she is moving (the amount of movement of the HMD) can be obtained, for example, from the output of an acceleration sensor. In this example, the mixing information acquisition unit 12 acquires transparency information in which the transparency level increases as the amount of movement increases. Then, the mixing ratio calculation unit 15 calculates a mixing ratio such that the amount of external sound reduction decreases as the transparency level of the transparency information increases, in other words, the amount of external sound reduction increases as the transparency level of the transparency information increases. Alternatively, the mixing information acquisition unit 12 may acquire transparency information in which the transparency level decreases as the amount of movement decreases. In this case, the mixing ratio calculation unit 15 calculates a mixing ratio such that the amount of external sound reduction increases as the transparency level of the transparency information decreases, in other words, the amount of external sound reduction decreases as the transparency level of the transparency information decreases.
[0074] In the third embodiment, for example, a designated area may be predefined in a virtual reality space or real space, and when the viewer leaves the designated area, the mixing information acquisition unit 12 may acquire predefined transparency information. For example, when a viewer is experiencing a VR game or the like using an HMD, a system has been put into practical use in which a real-world area in which the VR game is played is predefined, and the VR game is interrupted when the viewer leaves the area. When the viewer leaves the area defined in the real world, for example, a real image is displayed through the VR game image, allowing the viewer to check both the situation in the real space and the situation in the virtual reality world. Furthermore, in situations where the viewer checks their surroundings, it is desirable to simultaneously hear external sounds. To address this assumed example, the mixing information acquisition unit 12 may acquire transparency information that allows the VR image to be transparent when the viewer of the HMD leaves a designated area in real space, and may acquire transparency information that blocks the VR image when the viewer is within the designated area. In this case, the mixing ratio calculation unit 15 calculates a mixing ratio that reduces external sounds when it acquires transparency information that blocks the VR image.
[0075] In the third embodiment, an example has been shown in which the transparency of a VR image is set and then composited with a real image. However, as another example, the transparency of a real image may be set and then composited with a VR image. Furthermore, the transparency Tv does not necessarily have to be a value between 0 and 1. For example, it may be a value between 0 and 100, and after the process of S302, scaling may be performed by dividing the transparency Tv by 100. Furthermore, the transparency information does not necessarily have to be a numerical value representing the transparency of the entire VR image, but may be, for example, a transparent image with the same resolution as the actual image. In this case, for example, each pixel value represents the transparency of each transparent image. Then, the statistical value of these transparent images may be used as the transparency Tv. For example, the average value, median value, maximum value, minimum value, etc. of the transparent image may be used as the statistical value.
[0076] In the third embodiment, similarly to the example of the first embodiment described above, when the MR flag is 0 and only a VR image is displayed without displaying a real image, external sounds are substantially blocked by noise canceling. In contrast, as described in the second embodiment, for example, the viewer's ears may be physically covered to prevent the viewer from hearing external sounds. Also, in the third embodiment, for example, if a headphone device is used that allows external sounds from the surrounding real space to reach the viewer's ears when noise cancellation processing that reduces external sounds is not performed, the external sound acquisition processing of S106 and the external sound synthesis processing of S107 do not need to be performed.
[0077] <Fourth embodiment> In the second embodiment described above, distance information is used as spatial information, but in the fourth embodiment, an example will be described in which direction information is used as spatial information. A viewer may want to check the situation in a real space in a specific direction while also listening to external sounds in that direction. For example, by displaying a real image in the direction in which a target to be watched, such as a baby or a pet, is present and displaying a VR image in other directions, the viewer can check the situation of the target to be watched in real space. In this case, it is considered more desirable to be able to hear not only the real image of the baby or the like, but also the crying of the baby. Therefore, in the fourth embodiment, information about the direction in which the real image is captured is used as spatial information.
[0078] The hardware configuration and functional configuration of the information processing device of the fourth embodiment are the same as those of the first embodiment described above, and therefore a description thereof will be omitted. The functional configuration of the information processing device is generally the same as that of FIG. 3 described above, but in the fourth embodiment, the mixture information acquisition unit 12 acquires directional information by including it in the mixture information, and outputs this information to the output image acquisition unit 14 and the mixture proportion calculation unit 15. Thus, in the fourth embodiment, the mixture information includes directional information. Hereinafter, in the fourth embodiment, the same functional configurations and processing steps as those of the examples of the above-described embodiments are denoted by the same reference numerals, and descriptions thereof will be omitted, and differences will be mainly described below.
[0079] <Information Processing of the Fourth Embodiment> 8 is a flowchart showing the flow of information processing executed by an information processing device of the fourth embodiment. In the case of the fourth embodiment, if the MR flag is 0 in S105, the information processing device proceeds to processing from S108 onwards, whereas if the MR flag is 1, the information processing device proceeds to processing from S401 onwards. After processing from S401 to S406, the information processing device proceeds to processing of S110.
[0080] When it is determined that the MR flag is 1 and the process proceeds to S401, the real image acquisition unit 13 acquires a real image obtained by capturing an image of the real space around the viewer. Next, in S402 , the output image acquisition unit 14 combines the VR image acquired by the VR image acquisition unit 11 and the real image acquired by the real image acquisition unit 13 .
[0081] Next, proceeding to S403, the mixed information acquisition unit 12 acquires directional information as spatial information included in the mixed information. Here, the mixed information acquisition unit 12 acquires, as directional information, which of the left and right sides of the HMD screen has a larger proportion of the real image when the screen is divided into left and right halves. At this time, the mixed information acquisition unit 12 checks the number of pixels occupied by the VR image and the real image from the mask image acquired in S103 for each of the left and right sides of the HMD screen. If the number of pixels of the real image is larger, the mixed information acquisition unit 12 determines that the direction is the direction in which the real image has a larger proportion. Note that the HMD screen may be divided into left and right halves at any position other than the center. Furthermore, the HMD screen may be divided into any shape and number of halves, not just left and right. For example, if multiple directional microphones for collecting external sounds in the real space are installed facing different directions, the HMD screen may be divided according to the directions of the directional microphones. Furthermore, when an input to place the virtual display on the left half is received, the direction may be acquired based on the viewer's input, so that the right side is set as the direction with a larger proportion.
[0082] Next, in S404, the mixing ratio calculation unit 15 calculates a mixing ratio. For example, the mixing ratio calculation unit 15 calculates a mixing ratio such that external sounds in the direction determined in S403 to have a large proportion occupied by the real image are synthesized with the same intensity, while external sounds in directions other than the direction determined to have a large proportion occupied by the real image are not synthesized (external sounds are reduced).
[0083] Next, in S405, the external sound acquisition unit 17 acquires external sound. Here, in a case where multiple directional microphones are installed facing different directions as described above, the external sound acquisition unit 17 acquires external sound collected by each of the directional microphones.
[0084] Next, in S406, the sound synthesis unit 18 synthesizes the external sound based on the mixture ratio calculated in S404. For example, the sound synthesis unit 18 synthesizes the external sound acquired from a directional microphone installed facing the direction in which the external sound is to be synthesized at the same intensity, into the VR sound.
[0085] As described above, according to the information processing device of the fourth embodiment, it is possible to adjust external sounds based on directional information as spatial information. That is, in the case of the fourth embodiment, the mixing ratio calculation unit 15 calculates a mixing ratio that reduces external sounds from sound sources in directions other than the direction corresponding to the directional information in real space, and also calculates a mixing ratio that does not reduce external sounds from sound sources in the direction corresponding to the directional information. This allows the viewer to hear sounds from specific directions and also saves the viewer the trouble of adjusting external sounds.
[0086] The mixing ratio calculated in S404 does not necessarily have to be a mixing ratio that synthesizes the external sound with its original intensity. For example, an image showing the mixing ratio with the same resolution as the real image may be prepared, and a mixing ratio with a gain for the external sound intensity may be calculated by corresponding each pixel value. In this example, the mixing information acquisition unit 12 acquires directional information indicating the direction in which the image based on the real image occupies a larger proportion in a synthesized image obtained by synthesizing the real image and the VR image based on the mixing information. In addition, in this case, the mixing ratio calculation unit 15 calculates a normal distribution centered on the direction based on the proportion of the image based on the real image, that is, a normal distribution centered on the center of the area in the direction in which the proportion of the real image is considered to be large. Then, the mixing ratio calculation unit 15 uses the normal distribution value corresponding to each pixel position of the synthesized image as a gain for adjusting the intensity of the external sound. In other words, the mixing ratio calculation unit 15 acquires a mixing ratio that adjusts the gain of the external sound using the normal distribution value as a gain. As a result, in S406, the sound synthesis unit 18 multiplies the external sound by the gain corresponding to each direction and then synthesizes it into VR sound.
[0087] For example, when an object emitting sound moves between a direction in which external sound is not reduced and a direction in which external sound is reduced, the sound emitted by the moving object will change significantly at the boundary. A similar situation can also occur when the viewer is moving. Therefore, to prevent such sudden changes in the volume of external sound, the external sound may be adjusted gradually over time. For example, suppose an HDM screen is divided into two, left and right, and external sound is being reduced on the left side when a sound source enters the left side from the right. In this case, the movement of the object is detected using a known object detection method, and when detected, only that sound source is separated using a known sound separation technology. Then, sound data for that sound source is synthesized so that the sound is gradually reduced over a predetermined period. This is done, for example, by applying a gain that gradually decreases from 1 to 0 to the intensity of the separated external sound.
[0088] In addition, in S403, the mixture information acquisition unit 12 may recognize the area of a specific object using a known object recognition technology and use the area as direction information. In this case, in S404, the mixture ratio calculation unit 15 calculates a mixture ratio that synthesizes external sounds from sound sources in that area.
[0089] Also in the fourth embodiment, as in the above, for example, the viewer's ears may be physically covered to prevent the viewer from hearing external sounds. Also in the fourth embodiment, for example, if a headphone device is used in which external sounds from the surrounding real space reach the viewer's ears when noise cancellation processing that reduces external sounds is not performed, the external sound acquisition processing in S106 and the external sound synthesis processing in S107 do not need to be performed.
[0090] <Fifth embodiment> In the fourth embodiment, directional information was used as spatial information, but in the fifth embodiment, an example will be described in which area information is used as spatial information. Here, for example, if the proportion of the area occupied by a VR image in the image displayed on the HMD is large, it is considered that the content is centered on a virtual reality space, and in this case, it is considered desirable to increase the amount of reduction in external sound. On the other hand, if the proportion of the area occupied by a VR image is small, it is considered that the content is not centered on a virtual reality space, and it is desirable to be able to hear external sound. To deal with these situations, in the fifth embodiment, area information occupied by a VR image is used.
[0091] The hardware configuration and functional configuration of the information processing device in the fifth embodiment are the same as those in the first embodiment, and therefore will not be described further. The functional configuration of the information processing device is generally the same as that in FIG. 3 , but in the fifth embodiment, the mixture information acquisition unit 12 acquires mixture information including area information, and outputs this information to the output image acquisition unit 14 and the mixture proportion calculation unit 15. Thus, in the fifth embodiment, the mixture information includes area information. Hereinafter, in the fifth embodiment, the same functional configurations and processing steps as those in the examples of the above-described embodiments are denoted by the same reference numerals, and their description will be omitted, and differences will be mainly described below.
[0092] <Information Processing of the Fifth Embodiment> 9 is a flowchart showing the flow of information processing executed by an information processing device of the fifth embodiment. In the case of the fifth embodiment, if the MR flag is 0 in S105, the information processing device proceeds to processing from S108 onwards, whereas if the MR flag is 1, the information processing device proceeds to processing from S501 onwards. After processing from S501 to S506, the information processing device proceeds to processing of S110.
[0093] When it is determined that the MR flag is 1 and the process proceeds to S501, the real image acquisition unit 13 acquires a real image obtained by capturing the real space around the viewer. Next, in S502, the output image acquisition unit 14 combines the VR image acquired by the VR image acquisition unit 11 and the real image acquired by the real image acquisition unit 13.
[0094] Next, proceeding to S503, the mixed information acquisition unit 12 acquires area information of the VR image and the real image. For example, the mixed information acquisition unit 12 counts the number of pixels whose value is 1 in the mask image acquired in S104, and determines the ratio AR obtained by dividing the count of the number of pixels whose value is 1 by the total number of pixels as area information.
[0095] Next, in S504, the mixing ratio calculation unit 15 calculates the mixing ratio. In the fifth embodiment, as the mixing ratio, a gain of the ratio AR is applied to the intensity of the VR sound, and a gain of (1-AR) is applied to the intensity of the external sound.
[0096] Next, in S505, the external sound acquisition unit 17 acquires external sound. Then, in S506, the sound synthesis unit 18 synthesizes the external sound based on the mixture ratio calculated in S504.
[0097] As described above, according to the information processing device of the fifth embodiment, it is possible to adjust external sound based on area information as spatial information. That is, in the case of the fifth embodiment, the mixing ratio calculation unit 15 calculates a mixing ratio that reduces external sound when the area ratio of the VR image is larger than that of the real image in a composite image obtained by combining the real image and the VR image displayed on the HMD. On the other hand, the mixing ratio calculation unit 15 calculates a mixing ratio that does not reduce external sound when the area ratio of the real image is larger than that of the VR image. This allows the viewer to hear sound according to the content displayed on the HMD and also saves the trouble of adjusting external sound.
[0098] The area information may be calculated based on the angle of view corresponding to the real image and the VR image. For example, if the angle of view of the real image and the VR image is predetermined, the ratio between them is used as the area information. As described above, according to the information processing device of the fifth embodiment, it is possible to adjust the external sound based on the area information included in the spatial information. This makes it possible to adjust the external sound according to the area occupied by the VR image, thereby saving the viewer the trouble of adjusting the sound.
[0099] Also in the fifth embodiment, as in the above, for example, the viewer's ears may be physically covered to prevent the viewer from hearing external sounds. Also in the fifth embodiment, for example, if a headphone device is used in which external sounds from the surrounding real space reach the viewer's ears when noise cancellation processing that reduces external sounds is not performed, the external sound acquisition processing in S106 and the external sound synthesis processing in S107 do not need to be performed.
[0100] In the first to fifth embodiments described above, examples have been given in which a mask image, transparency information, and application program information are used as mixed information, and distance information, direction information, and area information are used individually as spatial information, but one or more of these pieces of information may be appropriately combined. That is, image synthesis, external sound synthesis, sound adjustment, etc. may be performed by appropriately combining one or more of these pieces of information.
[0101] The present invention can also be realized by supplying a program that realizes one or more functions of each of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions. The above-described embodiments are merely examples of specific implementations of the present invention, and the technical scope of the present invention should not be construed as being limited by these embodiments. In other words, the present invention can be implemented in various forms without departing from its technical concept or main features.
[0102] The disclosure of this embodiment includes the following configuration, method, and program. (Configuration 1) an information acquisition means for acquiring mixed information of a real image of a real space and a virtual image; a virtual sound acquisition means for acquiring a virtual sound; a ratio acquisition means for acquiring a mixing ratio between the virtual sound and an external sound in a real space based on the mixing information; an adjustment unit that adjusts the relative volume between the virtual sound and the external sound based on the mixing ratio; An information processing device comprising: (Configuration 2) an external sound acquisition means for acquiring a sound in a real space as the external sound; 2. The information processing device according to configuration 1, wherein the adjustment means generates a synthetic sound by synthesizing the virtual sound and the external sound based on the mixing ratio. (Configuration 3) a real image acquisition means for acquiring the real image obtained by capturing an image of the real space using an imaging device mounted on a head-mounted device; a virtual image acquisition means for acquiring the virtual image; an output means for outputting a composite image obtained by combining the real image and the virtual image based on the mixed information to a display device included in the head-mounted device; 3. The information processing device according to configuration 2, further comprising: (Configuration 4) the mixed information is information representing synthesis of the real image with the virtual image, The information processing device according to configuration 3, wherein the ratio acquisition means acquires the mixing ratio for reducing the external sound to be synthesized into the virtual sound when the mixing information does not include information indicating that the real image is synthesized into the virtual image. (Configuration 5) the mixed information is information representing synthesis of the real image with the virtual image, The information processing device according to configuration 3 or 4, wherein the ratio acquisition means acquires the mixing ratio that does not reduce the external sound to be synthesized into the virtual sound when the mixing information includes information indicating that the real image is to be synthesized into the virtual image. (Configuration 6) the mixed information is a mask image indicating an image region corresponding to the virtual image, The information processing device according to configuration 4 or 5, wherein the ratio acquisition means determines, based on the mask image, whether or not the mixed information includes information indicating that the real image is to be synthesized with the virtual image. (Configuration 7) the mixed information is information indicating an application program for displaying a composite image obtained by combining the real image of the real space and the virtual image on a display device included in the head-mounted device, The information processing device according to configuration 3, wherein the ratio acquisition means acquires the mixing ratio for reducing the external sound to be mixed with the virtual sound when the mixing information indicates that the application program is not one that displays a composite image obtained by combining the actual image of the real space with the virtual image. (Configuration 8) the mixed information is information indicating an application program for displaying a composite image obtained by combining the real image of the real space and the virtual image on a display device included in the head-mounted device, The information processing device according to configuration 3 or 7, wherein the ratio acquisition means acquires the mixing ratio that does not reduce the external sound that is mixed with the virtual sound when the mixing information indicates that the application program displays a composite image that combines the actual image of the real space with the virtual image. (Configuration 9) the mixed information includes spatial information corresponding to the real space in which the head-mounted device exists, 9. The information processing apparatus according to any one of configurations 3 to 8, wherein the ratio acquisition means acquires the mixture ratio according to the spatial information. (Configuration 10) the spatial information includes distance information from the head-mounted device in the real space, 10. The information processing device according to claim 9, wherein the ratio acquisition means acquires the mixing ratio for reducing external sounds from a sound source that is equal to or greater than a distance threshold represented by the distance information in the real space. (Configuration 11) the spatial information includes distance information from the head-mounted device in the real space, 11. The information processing device according to configuration 9 or 10, wherein the ratio acquisition means acquires the mixing ratio that does not reduce external sounds from a sound source in the real space that is less than a distance threshold represented by the distance information. (Configuration 12) the spatial information includes distance information from the head-mounted device in the real space, 10. The information processing device according to configuration 9, wherein the ratio acquisition means acquires the mixing ratio for adjusting the gain of the external sound in accordance with the distance information. (Configuration 13) the spatial information includes directional information from the head-mounted device in the real space; 10. The information processing device according to configuration 9, wherein the ratio acquisition means acquires the mixing ratio for reducing external sounds from sound sources in directions other than the direction corresponding to the direction information in the real space. (Configuration 14) the spatial information includes directional information from the head-mounted device in the real space; 14. The information processing device according to configuration 9 or 13, wherein the ratio acquisition means acquires the mixing ratio that does not reduce external sounds from a sound source in a direction in the real space that corresponds to the direction information. (Configuration 15) the spatial information includes directional information from the head-mounted device in the real space; the information acquiring means acquires the direction information indicating a direction in which an image based on the real image occupies a larger proportion in the composite image obtained by combining the real image and the virtual image based on the mixed information; The information processing device described in configuration 9, wherein the ratio acquisition means calculates a normal distribution centered on a direction based on the ratio occupied by an image based on the actual image, and acquires the mixing ratio for adjusting the gain of the external sound by using the value of the normal distribution corresponding to each pixel position of the synthetic image as a gain. (Configuration 16) 16. The information processing device according to any one of configurations 13 to 15, wherein the information acquisition means acquires the direction information indicating a direction in which an image based on the real image occupies a larger proportion in the composite image obtained by combining the real image and the virtual image based on the mixed information. (Configuration 17) the spatial information includes area information of an image based on the real image and an image based on the virtual image, The information processing device described in configuration 9, characterized in that the ratio acquisition means acquires the mixing ratio that reduces the external sound when, in a composite image obtained by combining the real image and the virtual image displayed on a display device provided in the head-mounted device, the ratio of area occupied by an image based on the virtual image is larger than that of the real image. (Configuration 18) the spatial information includes area information of an image based on the real image and an image based on the virtual image, The information processing device described in configuration 9 or 17, characterized in that the ratio acquisition means acquires the mixing ratio that does not reduce the external sound when, in a composite image obtained by combining the real image and the virtual image displayed on a display device provided in the head-mounted device, the ratio of area occupied by an image based on the real image is larger than that of the virtual image. (Configuration 19) 20. The information processing device according to any one of configurations 1 to 19, wherein the ratio acquisition means calculates the mixing ratio so that, when the mixing ratio changes, the change becomes gradual over time. (Configuration 20) the blending information includes transparency information that determines the transparency of an image based on the virtual image; 4. The information processing device according to configuration 3, wherein the ratio acquisition means acquires the mixing ratio according to the transmission information. (Configuration 21) the information acquisition means acquires the transparency information in which the transparency decreases as the amount of movement of the head-mounted device in the real space decreases, and 21. The information processing device according to configuration 20, wherein the ratio acquisition means calculates the mixing ratio such that the amount of reduction of the external sound increases as the transparency of the transmission information decreases. (Configuration 22) the information acquisition means acquires the transparency information in which the transparency increases as the amount of movement of the head-mounted device in the real space increases, and 22. The information processing device according to configuration 20 or 21, wherein the ratio acquisition means acquires the mixing ratio such that the external sound is louder as the transparency of the transparency information increases. (Configuration 23) the information acquisition means acquires the transparency information that transmits an image based on the virtual image when the head-mounted device is outside the designated area of the real space, and acquires the transparency information that does not transmit an image based on the virtual image when the head-mounted device is within the designated area of the real space; The information processing device according to any one of configurations 20 to 22, characterized in that the ratio acquisition means acquires the mixing ratio that reduces the external sound when the transmission information that does not allow an image based on the virtual image to be transmitted is acquired. (Configuration 24) An information processing device characterized by having an adjustment means for adjusting the relative volume between a sound in a real space and a sound associated with a displayed image based on at least one of an actual image corresponding to the real space and the displayed image. (Method 1) An information processing method executed by an information processing device, an information acquisition step of acquiring mixed information of a real image of a real space and a virtual image; a virtual sound acquisition step of acquiring a virtual sound; a ratio acquisition step of acquiring a mixing ratio between the virtual sound and an external sound in a real space based on the mixing information; an adjustment step of adjusting the relative volume between the virtual sound and the external sound based on the mixing ratio; An information processing method comprising: (Program 1) A program that causes a computer to function as the information processing device according to any one of configurations 1 to 24. [Explanation of symbols]
[0103] 11: VR image acquisition unit, 12: Mixed information acquisition unit, 13: Real image acquisition unit, 14: Output image acquisition unit, 15: Mixed ratio calculation unit, 16: VR sound acquisition unit, 17: External sound acquisition unit, 18: Sound synthesis unit, 19: Output unit
Claims
1. an image acquisition means for acquiring a real image and a virtual image corresponding to a real space; a sound acquisition means for acquiring an external sound in the real space and a virtual sound associated with the virtual image; an adjustment unit that adjusts the relative volume between the external sound and the virtual sound based on at least one of the real image and the virtual image; an output means for outputting a composite image obtained by combining the real image and the virtual image based on the mixed information to a display device provided in a head-mounted device; and the adjustment means generates a synthetic sound by combining the virtual sound and the external sound based on at least one of the real image and the virtual image; the real image is an image obtained by capturing an image of the real space using an imaging device mounted on the head-mounted device, the blending information includes transparency information that determines a transparency of an image based on the virtual image; the adjusting means performs the adjustment in accordance with the transmission information, the mixed information includes the transparency information in which the transparency decreases as the amount of movement of the head-mounted device in the real space decreases, and The information processing device is characterized in that the adjustment means performs control such that the amount of reduction of the external sound increases as the transparency of the transmission information decreases.
2. The image processing device further includes an information acquisition means for acquiring mixed information representing the synthesis of the real image onto the virtual image, 2. The information processing apparatus according to claim 1, wherein, when the mixing information includes information indicating that the real image is to be combined with the virtual image, the adjustment means performs control not to reduce the external sound.
3. The mixed information is a mask image indicating an image area corresponding to the virtual image, 3. The information processing apparatus according to claim 2, wherein the adjustment means determines, based on the mask image, whether or not the mixed information includes information representing that the real image is to be combined with the virtual image.
4. The mixed information is information indicating an application program for displaying a composite image obtained by combining the real image of the real space and the virtual image on a display device provided in the head-mounted device, The information processing device according to claim 1, characterized in that, when the mixed information indicates an application program that displays a composite image that combines a real image of the real space with the virtual image, the adjustment means performs control so as not to reduce the external sound.
5. The mixed information includes distance information from the head-mounted device to an object in the real space, The information processing apparatus according to claim 1 , wherein the adjustment means performs control to reduce external sounds from a sound source that is equal to or greater than a distance threshold represented by the distance information in the real space.
6. The information processing device described in Claim 5, characterized in that the adjustment means performs control to adjust the gain of the external sound according to the distance information.
7. The information processing device described in Claim 1, characterized in that the mixed information is the proportion of images based on the actual image in the synthetic image.
8. The mixed information includes area information of an image based on the real image and an image based on the virtual image, The information processing device according to claim 1, characterized in that the adjustment means performs control so as not to reduce the external sound when the proportion of the area occupied by the image based on the real image is larger than that of the virtual image in the synthetic image displayed on the display device provided in the head-mounted device.
9. The information processing device described in Claim 1, characterized in that the adjustment means adjusts the external sound so that it changes gradually over time.
10. Image acquisition means for acquiring a real image and a virtual image corresponding to a real space; a sound acquisition means for acquiring an external sound in the real space and a virtual sound associated with the virtual image; an adjustment unit that adjusts the relative volume between the external sound and the virtual sound based on at least one of the real image and the virtual image; an output means for outputting a composite image obtained by combining the real image and the virtual image based on the mixed information to a display device provided in a head-mounted device; and the adjustment means generates a synthetic sound by combining the virtual sound and the external sound based on at least one of the real image and the virtual image; the real image is an image obtained by capturing an image of the real space using an imaging device mounted on the head-mounted device, the blending information includes transparency information that determines a transparency of an image based on the virtual image; the adjusting means performs the adjustment in accordance with the transmission information, the mixed information includes the transparency information in which the transparency increases as the amount of movement of the head-mounted device in the real space increases, and The information processing device is characterized in that the adjustment means performs control to increase the volume of the external sound as the transparency of the transparency information increases.
11. An information processing device as described in Claim 1, characterized in that when the head-mounted device is outside the specified area of the real space, the adjustment means performs control to reduce the external sound.
12. An image acquisition step of acquiring a real image and a virtual image corresponding to a real space; a sound acquisition step of acquiring an external sound in the real space and a virtual sound associated with the virtual image; an adjustment step of adjusting the relative volume between the external sound and the virtual sound based on at least one of the real image and the virtual image; an output step of outputting a composite image obtained by combining the real image and the virtual image based on the mixture information to a display device provided in a head-mounted device; and In the adjustment step, a synthetic sound is generated by synthesizing the virtual sound and the external sound based on at least one of the real image and the virtual image, the real image is an image obtained by capturing an image of the real space using an imaging device mounted on the head-mounted device, the blending information includes transparency information that determines a transparency of an image based on the virtual image; In the adjusting step, the adjustment is performed in accordance with the transmission information, the mixed information includes the transparency information in which the transparency decreases as the amount of movement of the head-mounted device in the real space decreases, and The information processing method, wherein the adjusting step performs control such that the amount of reduction of the external sound increases as the transparency of the transmission information decreases.
13. A program that causes a computer to function as the information processing device described in claim 1.
Citation Information
Patent Citations
Information processing program, information processing method and program
JP2017069687A
Program, method, and information processing apparatus
JP2021002390A
User-based context sensitive hologram reaction
US20160266386A1
Filtering sounds for conferencing applications
US20160379660A1
Voxel-based, real-time acoustic adjustment
US20170165575A1