Information processing device, control method for information processing device, program, recording medium, and system
The system enhances virtual reality immersion by synchronizing sound and image playback with viewer attention and device orientation, addressing the disjointedness in existing technologies through dynamic sound adjustments.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-07
- Publication Date
- 2026-04-13
AI Technical Summary
Existing technologies fail to provide a realistic audio-visual experience in virtual reality environments by accurately synchronizing sound and image playback based on viewer attention and device orientation, leading to a disjointed and less immersive experience.
An information processing system that synchronizes sound playback with viewer attention and device orientation by adjusting sound volume and source based on viewer focus and device position, using multiple sound-collecting and imaging devices to generate immersive audio-visual experiences.
Enhances the realism and immersion of virtual reality experiences by dynamically adjusting sound levels and sources based on viewer attention and device orientation, providing a more engaging and realistic audio-visual experience.
Smart Images

Figure 0007844196000001 
Figure 0007844196000002 
Figure 0007844196000003
Abstract
Description
Technical Field
[0004] , , , , , ,
[0005] , , , , ,
[0003] , , , , ,
[0001] The present invention relates to an information processing device, a control method for an information processing device, a program, a recording medium, and a system.
Background Art
[0006] A first aspect of the present invention is: The system includes: motion image acquisition means for acquiring multiple motion images captured by multiple sound-collecting devices and multiple corresponding imaging devices; sound acquisition means for acquiring multiple sounds captured by the multiple sound-collecting devices in synchronization with the capture of the multiple motion images; device information acquisition means for acquiring device information relating to the imaging device selected by the viewer; control means for controlling the playback of motion images captured by the imaging device selected by the viewer based on the device information acquired by the device information acquisition means; state information acquisition means for acquiring state information relating to the viewer's attention state to the playback motion images; and generation means for generating sounds to be played together with the playback motion images based on the state information acquired by the state information acquisition means, wherein if the state information indicates that the viewer is viewing the playback motion images from above, the generation means sets the proportion of sounds captured by the multiple sound-collecting devices installed within the area being viewed by the viewer to be the same, and sets the proportion of sounds captured by the sound-collecting devices installed outside the area to be set The longer the distance from the area to the sound-collecting device, the smaller the volume of the multiple sounds is increased, and the multiple sounds are synthesized to generate a sound to be played back with the video. If the state information indicates that the viewer is paying attention to a specific object included in the video being played back, the shorter the distance from the object the viewer is paying attention to to the sound-collecting device, the larger the increase in volume of the sound picked up by the sound-collecting device, and the multiple sounds are synthesized to generate a sound to be played back with the video. If there is a change in the viewer's selection of the imaging device, before the change in the imaging device, the state information acquisition means acquires information as state information indicating that the viewer is paying attention to a specific object included in the video before the change, and the changed video includes the specific object, then the shorter the distance from the object the viewer is paying attention to to the sound-collecting device, the larger the increase in volume of the sound picked up by the sound-collecting device, and the multiple sounds are synthesized to generate a sound to be played back with the changed video. Prior to the change, the information processing apparatus obtains information as state information indicating that the viewer is paying attention to a specific object included in the original video, and if the modified video does not contain the specific object, the apparatus determines that the sound picked up by the sound pickup device corresponding to the modified imaging device will be played back together with the modified video, without synthesizing the multiple sounds. That is the case. A second aspect of the present invention includes: motion image acquisition means for acquiring multiple motion images captured by a plurality of sound-collecting devices and a plurality of imaging devices corresponding to each of them; sound acquisition means for acquiring multiple sounds captured by the plurality of sound-collecting devices in synchronization with the capture of the plurality of motion images; device information acquisition means for acquiring device information relating to an imaging device selected by a viewer; control means for controlling the playback of motion images captured by the imaging device selected by the viewer based on the device information acquired by the device information acquisition means; state information acquisition means for acquiring state information relating to the viewer's attention state to the playback motion images; and based on the state information acquired by the state information acquisition means, the playback together with the playback motion images. The information processing device comprises a sound generation means for generating sounds, wherein, when the state information indicates that the viewer is viewing the video to be played back, even if the region of the video being viewed by the viewer does not include a specific sound-gathering device which corresponds to the imaging device selected by the viewer, the generation means makes the ratio of multiple sounds picked up by multiple sound-gathering devices set within the region to the sound picked up by the specific sound-gathering device the same, and the ratio of sounds picked up by sound-gathering devices different from the specific sound-gathering device, which are installed outside the region, decreases as the distance from the region to the sound-gathering device increases, thereby synthesizing the multiple sounds to generate a sound to be played back together with the video.
[0007] This invention 3 The manner of is, The system includes: a video acquisition step of acquiring multiple video images captured by multiple sound-collecting devices and multiple corresponding imaging devices; a sound acquisition step of acquiring multiple sounds captured by the multiple sound-collecting devices in synchronization with the capture of the multiple video images; a device information acquisition step of acquiring device information relating to the imaging device selected by the viewer; a control step of controlling the playback of video images captured by the imaging device selected by the viewer based on the device information acquired in the device information acquisition step; a state information acquisition step of acquiring state information relating to the viewer's attention state to the video images to be played back; and a generation step of generating sounds to be played back together with the video images to be played back based on the state information acquired in the state information acquisition step, wherein in the generation step, if the state information indicates that the viewer is viewing the video images to be played back from above, the proportion of the video images that were captured by the multiple sound-collecting devices installed within the area being viewed by the viewer is set accordingly. If the viewer changes their selection of imaging device, if, before the change in imaging device, information indicating that the viewer is paying attention to a specific object in the pre-changed video is acquired as state information in the state information acquisition step, and if the post-changed video contains the specific object, if, before the change in imaging device, the proportion of sound acquired by the sound acquisition device is increased as the distance from the area to the sound acquisition device increases, the proportion of sound acquired by the sound acquisition device is increased as the distance from the area to the sound acquisition device increases, the proportion of sound acquired by the sound acquisition device is increased as the proportion of sound acquired by the sound acquisition device is increased as the distance from the object the viewer is paying attention to, the proportion of sound acquired by the sound acquisition device is increased as the distance from the object the viewer is paying attention to increases, the proportion of sound acquired by the sound acquisition device is increased as the proportion of sound acquired by the sound acquisition device is increased, the proportion of sound acquired by the sound acquisition device is increased as the distance from the object to the viewer increases, the proportion of sound acquired by the sound acquisition device is increased Control method for an information processing device, characterized by: increasing the volume of the collected sounds and synthesizing the multiple sounds to generate a sound to be played back with the modified video; before the change of the imaging device, in the state information acquisition step, information indicating that the viewer is paying attention to a specific object included in the video before the change is acquired as state information, and if the modified video does not include the specific object, then determining that the sound collected by the sound collection device corresponding to the modified imaging device will be played back with the modified video without synthesizing the multiple sounds; That is the case.
[0008] This invention 4 The embodiment is a program that causes a computer to function as one of the means of the information processing device described above.
[0009] This invention 5 The embodiment is a computer-readable recording medium that stores a program for causing the computer to function as each of the means of the information processing device described above.
[0010] This invention 6 The manner of is, A system comprising a plurality of imaging devices, a plurality of sound-collecting devices corresponding to each of the plurality of imaging devices, and an information processing device, wherein the information processing device includes: motion image acquisition means for acquiring a plurality of motion images captured by the plurality of imaging devices; sound acquisition means for acquiring a plurality of sounds collected by the plurality of sound-collecting devices in synchronization with the capture of the plurality of motion images; device information acquisition means for acquiring device information relating to an imaging device selected by a viewer; control means for controlling the playback of motion images captured by the imaging device selected by the viewer based on the device information acquired by the device information acquisition means; state information acquisition means for acquiring state information relating to the viewer's attention state to the playback motion images; and sound acquisition means for controlling the playback of sound images together with the playback motion images based on the state information acquired by the state information acquisition means. The system comprises a generation means for generating a sound, wherein, when the state information indicates that the viewer is viewing the video being played, the generation means equalizes the proportion of multiple sounds picked up by multiple sound-collecting devices installed within the area being viewed by the viewer, and decreases the proportion of sounds picked up by sound-collecting devices installed outside the area as the distance from the area to the sound-collecting device increases, thereby synthesizing the multiple sounds to generate a sound to be played with the video; and when the state information indicates that the viewer is focusing on a specific object included in the video being played, the generation means increases the volume of the sound picked up by the sound-collecting device by a larger increase as the distance from the object the viewer is focusing on to the sound-collecting device decreases, thereby synthesizing the multiple sounds to generate a sound to be played with the video. The system is characterized in that, if there is a change in the viewer's selection of the imaging device, before the change in the imaging device, the state information acquisition means acquires information as state information indicating that the viewer is paying attention to a specific object included in the video before the change, and if the video after the change includes the specific object, the volume of the sound picked up by the sound picking device is increased by a larger amount the shorter the distance from the object the viewer is paying attention to to the sound picking device, and the multiple sounds are synthesized to generate a sound to be played back with the video after the change, and before the change in the imaging device, the state information acquisition means acquires information as state information indicating that the viewer is paying attention to a specific object included in the video before the change, and if the video after the change does not include the specific object, the sound picked up by the sound picking device corresponding to the changed imaging device is determined to be played back with the video after the change without synthesizing the multiple sounds. That is the case. [Effects of the Invention]
[0011] According to the present invention, it is possible to provide a technology that can give viewers a sufficient sense of realism. [Brief explanation of the drawing]
[0012] [Figure 1] These are external view and block diagrams of a digital camera. [Figure 2] These include external view diagrams and block diagrams of the display control device. [Figure 3] This is a block diagram of the distribution device. [Figure 4] This is a block diagram of a system including an imaging device, a distribution device, and a display control device. [Figure 5]This is a diagram showing an example of a state where a viewer is watching a VR image. [Figure 6] This is a flowchart of the distribution process. [Figure 7] This is a flowchart of the determination process for the distribution audio based on the bird's-eye view state. [Figure 8] This is a flowchart of the determination process for the distribution audio based on the target of interest. [Figure 9] This is a flowchart of the determination process for the distribution audio based on the area of interest.
Embodiments for Carrying Out the Invention
[0013] Hereinafter, embodiments of the present invention will be described with reference to the drawings.
[0014] <Configuration> FIG. 1(a) is a front perspective view (appearance view) of a digital camera 100 (imaging device). FIG. 1(b) is a rear perspective view (appearance view) of the digital camera 100. The digital camera 100 is an omnidirectional camera (full spherical camera).
[0015] The barrier 102a is a protective window for the front camera unit with the front of the digital camera 100 as the shooting range. The front camera unit is, for example, a wide-angle camera unit with a shooting range covering a wide area of 180 degrees or more in the up, down, left, and right directions on the front side of the digital camera 100. The barrier 102b is a protective window for the rear camera unit with the rear of the digital camera 100 as the shooting range. The rear camera unit is, for example, a wide-angle camera unit with a shooting range covering a wide area of 180 degrees or more in the up, down, left, and right directions on the rear side of the digital camera 100.
[0016] The display unit 28 displays various information. The shutter button 61 is an operating unit (operating component) for issuing shooting instructions. The mode selector switch 60 is an operating unit for switching between various modes. The connection interface 25 is a connector for connecting a connection cable to the digital camera 100, and external devices such as smartphones, personal computers, and televisions are connected to the digital camera 100 using the connection cable. The operation unit 70 consists of various switches, buttons, dials, touch sensors, etc. that accept various operations from the user. The power switch 72 is a push button for switching the power on / off.
[0017] The light-emitting part 21 is a light-emitting element such as a light-emitting diode (LED), and notifies the user of various states of the digital camera 100 by light-emitting patterns and colors. The fixing part 40 is, for example This is a tripod screw hole, used to secure and mount the digital camera 100 with a tripod or other mounting device.
[0018] Figure 1(c) is a block diagram showing an example configuration of the digital camera 100.
[0019] The barrier 102a prevents dirt and damage to the imaging system of the front camera unit (photographic lens 103a, shutter 101a, imaging unit 22a, etc.) by covering the system. The photographic lens 103a is a lens group including a zoom lens and a focus lens, and is a wide-angle lens. The shutter 101a is a shutter with an aperture function that adjusts the amount of subject light incident on the imaging unit 22a. The imaging unit 22a is an image sensor (image sensor) composed of a CCD or CMOS element that converts an optical image into an electrical signal. The A / D converter 23a converts the analog signal output from the imaging unit 22a into a digital signal. Alternatively, the barrier 102a may be omitted, exposing the outer surface of the photographic lens 103a, and the photographic lens 103a may prevent dirt and damage to other imaging system components (shutter 101a and imaging unit 22a).
[0020] The barrier 102b prevents dirt and damage to the imaging system of the rear camera unit (photographing lens 103b, shutter 101b, imaging unit 22b, etc.) by covering the system. The photographic lens 103b is a lens group including a zoom lens and a focus lens, and is a wide-angle lens. The shutter 101b is a shutter with an aperture function that adjusts the amount of subject light incident on the imaging unit 22b. The imaging unit 22b is an image sensor composed of a CCD or CMOS element, etc., which converts an optical image into an electrical signal. The A / D converter 23b converts the analog signal output from the imaging unit 22b into a digital signal. Alternatively, the barrier 102b may be omitted, exposing the outer surface of the photographic lens 103b, and the photographic lens 103b may prevent dirt and damage to other imaging system components (shutter 101b and imaging unit 22b).
[0021] Virtual Reality (VR) images are captured by the imaging units 22a and 22b. A VR image is defined as an image that can be displayed in VR (displayed in the "VR View" display mode). VR images include omnidirectional images (spherical images) captured by an omnidirectional camera (spherical camera), and panoramic images that have a wider image range (effective image range) than the display range that can be displayed on the display unit at one time. VR images include not only still images, but also videos and live view images (images acquired from the camera in near real-time). VR images have an image range (effective image range) of up to 360 degrees vertically (vertical angle, angle from the zenith, elevation angle, depression angle, altitude angle, pitch angle) and 360 degrees horizontally (horizontal angle, azimuth angle, yaw angle).
[0022] Furthermore, VR images include images that have a wider field of view (view range) than that which can be captured by a normal camera, even if the vertical or horizontal field of view is less than 360 degrees, or an image range (effective image range) that is wider than the display range that can be displayed on a display unit at once. For example, an image taken with a 360-degree spherical camera capable of capturing a subject with a field of view (view range) of 360 degrees horizontally and 210 degrees vertically centered on the zenith is a type of VR image. Also, for example, an image taken with a camera capable of capturing a subject with a field of view (view range) of 180 degrees horizontally and 180 degrees vertically centered on the horizontal is a type of VR image. In other words, an image with an image range of 160 degrees (±80 degrees) or more vertically and horizontally, and an image range wider than the range that a human can see at once, is a type of VR image.
[0023] When this VR image is displayed in VR mode (using the "VR View" display mode), by changing the orientation of the display device (the device that displays the VR image) in the left-right rotation direction, you can view a seamless, omnidirectional image in the left-right direction (horizontal rotation direction). In the up-down direction (vertical rotation direction), you can view a seamless, omnidirectional image within a range of ±105 degrees from directly above (zenith). While this is possible, the area beyond 105 degrees from directly above becomes a blank region where no image exists. A VR image can also be described as "an image in which the image range is at least part of a virtual space (VR space)."
[0024] VR display (VR view) is a display method (display mode) that allows the display range to be changed, displaying images within a field of view that corresponds to the orientation of the display device. One type of VR display is "single-eye VR display (single-eye VR view)," which displays a single image by performing a transformation (distortion correction) that maps the VR image to a virtual sphere. Another type of VR display is "two-eye VR display (two-eye VR view)," which displays a VR image for the left eye and a VR image for the right eye side by side by performing a transformation that maps each to a virtual sphere. By performing "two-eye VR display" using VR images for the left eye and the right eye that have parallax with each other, it is possible to view these VR images in 3D. In any type of VR display, when viewing with a head-mounted display (HMD), the display device will show images within a field of view that corresponds to the orientation of the user's face. For example, suppose that in a VR image, at a certain point in time, the displayed image has a field of view centered at 0 degrees horizontally (a specific direction, e.g., north) and 90 degrees vertically (90 degrees from the zenith, i.e., horizontal). From this state, if the orientation of the display device is reversed (for example, changing the display surface from facing south to facing north), the display range of the same VR image changes to an image with a field of view centered on 180 degrees horizontally (opposite direction, e.g., south) and 90 degrees horizontally. In the case of a user viewing an HMD, if the user turns their face from north to south (i.e., turns their back), the image displayed on the HMD will also change from a northern image to a southern image. This type of VR display can provide the user with a visual sensation (sense of immersion) as if they were actually in the VR image (VR space). A smartphone attached to VR goggles (head-mounted adapter) can be considered a type of HMD.
[0025] The methods for displaying VR images are not limited to those described above. The display range may be moved (scrolled) in response to user operations such as touch panels or directional buttons, rather than changes in posture. When displaying in VR (in the "VR View" display mode), the display range may be changed not only in response to changes in posture, but also in response to touch movements on the touch panel, drag operations with a mouse, or presses of directional buttons.
[0026] The image processing unit 24 performs resizing and color conversion processing, such as predetermined pixel interpolation and reduction, on data from the A / D converters 23a and 23b, or from the memory control unit 15. The image processing unit 24 also performs predetermined calculation processing using the captured image data. The system control unit 50 performs exposure control and distance measurement control based on the calculation results obtained by the image processing unit 24. This enables TTL (through-the-lens) AF (autofocus), AE (automatic exposure), and EF (flash pre-flash) processing. The image processing unit 24 further performs predetermined calculation processing using the captured image data and performs TTL AWB (auto white balance) processing based on the obtained calculation results. The image processing unit 24 also applies basic image processing to the two images obtained from the A / D converters 23a and 23b (two fisheye images; two wide-angle images), and performs image merging processing to combine the two images with basic image processing applied to generate a single VR image. Furthermore, the image processing unit 24 performs image cropping, scaling, and distortion correction on the VR image during VR display in live view or playback, and renders the processing results to the VRAM in memory 32.
[0027] In image stitching processing, the image processing unit 24 uses one of the two images as a reference image and the other as a comparison image, calculates the amount of displacement between the reference image and the comparison image for each area using pattern matching processing, and detects the stitching position where the two images are joined based on the amount of displacement for each area. The image processing unit 24 considers the detected stitching position and the lens characteristics of each optical system and performs geometric transformations to connect each Image distortion is corrected, and each image is converted into a 360-degree spherical image (360-degree image format). Then, the image processing unit 24 generates a single 360-degree spherical image (VR image) by combining (blending) two 360-degree spherical images. The generated 360-degree spherical image is, for example, an image using equirectangular projection, and the position of each pixel in the 360-degree spherical image can be associated with the coordinates of the surface of the sphere (VR space).
[0028] The output data from the A / D converters 23a and 23b is written to the memory 32 via the image processing unit 24 and the memory control unit 15, or via the memory control unit 15 without going through the image processing unit 24. The memory 32 stores image data obtained by the imaging units 22a and 22b and converted into digital data by the A / D converters 23a and 23b, as well as image data for output to an external display from the connection I / F 25. The memory 32 has sufficient storage capacity to store a predetermined number of still images, a predetermined amount of video footage, and audio.
[0029] Furthermore, memory 32 also serves as memory for image display (video memory). The image display data stored in memory 32 can be output to an external display via the connection I / F 25.VR images captured by the imaging units 22a and 22b and generated by the image processing unit 24, which are stored in memory 32, can be sequentially transferred to an external display for display, thereby realizing the function of an electronic viewfinder and enabling live view display (LV display).Hereinafter, the images displayed in live view display will be referred to as live view images (LV images).In addition, live view display (remote LV display) can also be performed by transferring the VR images stored in memory 32 to an external device (such as a smartphone) wirelessly connected via the communication unit 54 and displaying them on the external device.
[0030] The non-volatile memory 56 is a memory recording medium that can be electrically erased and recorded, such as an EEPROM. Constants for the operation of the system control unit 50, programs, etc., are recorded in the non-volatile memory 56. The program referred to here is a computer program for executing various processes.
[0031] The system control unit 50 is a control unit having at least one processor or circuit, and controls the entire digital camera 100. The system control unit 50 performs each process by executing the program recorded in the non-volatile memory 56 mentioned above. The system memory 52 is, for example, RAM, and the system memory 52 stores constants, variables for the operation of the system control unit 50, the program read from the non-volatile memory 56, etc. The system control unit 50 also performs display control by controlling the memory 32, the image processing unit 24, the memory control unit 15, etc. The system timer 53 is a timing unit that measures the time used for various controls and the time of the built-in clock.
[0032] The mode selector switch 60, shutter button 61, operation unit 70, and power switch 72 are used to input various operation instructions to the system control unit 50.
[0033] The mode switch 60 switches the operating mode of the system control unit 50 to one of the following: still image recording mode, video recording mode, playback mode, communication connection mode, etc. Modes included in the still image recording mode include auto shooting mode, auto scene detection mode, manual mode, aperture priority mode (Av mode), shutter speed priority mode (Tv mode), and program AE mode. There are also various scene modes and custom modes that provide shooting settings for different shooting scenes. The user can switch directly to any of these modes using the mode switch 60. Alternatively, the user can switch to the shooting mode list screen using the mode switch 60, and then selectively switch to one of the multiple modes displayed on the display unit 28 using another operating element. Similarly, the video recording mode may also include multiple modes.
[0034] The shutter button 61 is equipped with a first shutter switch 62 and a second shutter switch 64. The first shutter switch 62 turns ON during the operation of the shutter button 61, so-called half-press (instruction to prepare for shooting), and generates a first shutter switch signal SW1. The system control unit 50 starts shooting preparation operations such as AF (autofocus) processing, AE (automatic exposure) processing, AWB (auto white balance) processing, and EF (flash pre-flash) processing in response to the first shutter switch signal SW1. The second shutter switch 64 turns ON when the operation of the shutter button 61 is completed, so-called full-press (instruction to shoot), and generates a second shutter switch signal SW2. The system control unit 50 starts a series of shooting processes from reading signals from the imaging units 22a and 22b to writing image data to the recording medium 90 in response to the second shutter switch signal SW2.
[0035] Furthermore, the shutter button 61 is not limited to an operating element that allows for two-stage operation (full press and half press), but may also be an operating element that allows for only one-stage pressing. In that case, the shooting preparation operation and the shooting process are performed consecutively by pressing the button once. This is the same operation as when a shutter button that allows for half-press and full press is fully pressed (when the first shutter switch signal SW1 and the second shutter switch signal SW2 occur almost simultaneously).
[0036] The operation unit 70 is assigned various functions as function buttons depending on the situation, by selecting various function icons and options displayed on the display unit 28. Examples of function buttons include an exit button, a back button, an image advance button, a jump button, a filter button, and an attribute change button. For example, when the menu button is pressed, various configurable menu screens are displayed on the display unit 28. The user can intuitively make various settings by operating the operation unit 70 while looking at the menu screen displayed on the display unit 28.
[0037] The power control unit 80 consists of a battery detection circuit, a DC-DC converter, a switch circuit for switching which blocks are energized, and detects whether a battery is installed, the type of battery, the remaining battery level, etc. The power control unit 80 also controls the DC-DC converter based on the detection results and instructions from the system control unit 50, supplying the necessary voltage to each part, including the recording medium 90, for the required period of time. The power supply unit 30 consists of primary batteries such as alkaline batteries and lithium batteries, secondary batteries such as NiCd batteries, NiMH batteries and Li batteries, an AC adapter, etc.
[0038] The recording medium I / F18 is an interface to a recording medium 90, such as a memory card or hard disk. The recording medium 90 is a recording medium such as a memory card for recording captured images, and is composed of semiconductor memory, optical disks, magnetic disks, etc. The recording medium 90 may be a replaceable recording medium that can be attached to or removed from the digital camera 100, or it may be a recording medium built into the digital camera 100.
[0039] The communication unit 54 transmits and receives video signals, audio signals, and other signals to and from external devices connected wirelessly or via wired cables. The communication unit 54 can also connect to a wireless LAN (Local Area Network) or the Internet. The communication unit 54 can transmit images (including LV images) captured by the imaging units 22a and 22b, as well as images recorded on the recording medium 90, and can receive images and other various information from external devices.
[0040] The attitude detection unit 55 detects the attitude of the digital camera 100 relative to the direction of gravity. Based on the attitude detected by the attitude detection unit 55, it is possible to determine whether the images captured by the imaging units 22a and 22b were taken with the digital camera 100 held horizontally or vertically. Furthermore, it is possible to determine how much the digital camera 100 was tilted in the three axis directions (rotational directions) of yaw, pitch, and roll when capturing images with the imaging units 22a and 22b. It is possible to determine whether an image has been captured. The system control unit 50 can add orientation information corresponding to the posture detected by the posture detection unit 55 to the image file of the VR image captured by the imaging units 22a and 22b, or rotate the image (adjust the orientation of the image to correct the tilt (zenith correction)) and record it. One or a combination of sensors from among an acceleration sensor, gyro sensor, geomagnetic sensor, compass sensor, and altitude sensor can be used as the posture detection unit 55. It is also possible to detect the movement of the digital camera 100 (pan, tilt, lift, whether it is stationary or not, etc.) using the acceleration sensor, gyro sensor, and compass sensor that constitute the posture detection unit 55.
[0041] Microphone 20 is a microphone that collects (records) sounds from the surroundings of the digital camera 100, which are recorded as audio for the VR image (VR video). Connection I / F25 is a connection plug to which HDMI® cables, USB cables, etc., are connected for transmitting and receiving video to external devices.
[0042] In Figures 1(a) to 1(c), the digital camera 100 is explained as an example of an omnidirectional camera. However, it may also be a twin-lens VR camera having a twin-lens unit capable of capturing right and left images with parallax. The twin-lens unit has a fisheye lens capable of capturing a range of approximately 180 degrees in both the right-eye optical system and the left-eye optical system. By using the twin-lens unit, one image containing two image regions with parallax can be acquired from two locations (optical systems), the right-eye optical system and the left-eye optical system. By dividing the acquired image into an image for the left eye and an image for the right eye and displaying them in VR, the user can view a three-dimensional VR image with a range of approximately 180 degrees.
[0043] Figure 2(a) is an external view of a display control device 200, which is a type of information processing device. The display control device 200 is a display device such as a smartphone. The display 205 is a display unit that displays images and various information. The display 205 is integrally configured with a touch panel 206a and is capable of detecting touch operations on the display surface of the display 205. The display control device 200 is capable of displaying VR images (VR content) in VR on the display 205. The operation unit 206b is a power button that accepts the operation of switching the power of the display control device 200 on and off. The operation units 206c and 206d are volume buttons that increase or decrease the volume of sound output from the speaker 212b or from earphones or external speakers connected to the audio output terminal 212a. The operation unit 206e is a home button for displaying the home screen on the display 205. The audio output terminal 212a is an earphone jack, which is a terminal that outputs an audio signal to earphones or external speakers. The speaker 212b is a built-in speaker that outputs sound.
[0044] Figure 2(b) is a block diagram showing an example configuration of the display control device 200. The CPU 201, memory 202, non-volatile memory 203, image processing unit 204, display 205, operation unit 206, recording medium I / F 207, external I / F 209, and communication I / F 210 are connected to the internal bus 250. The audio output unit 212 and attitude detection unit 213 are also connected to the internal bus 250. Each unit connected to the internal bus 250 is configured to exchange data with each other via the internal bus 250.
[0045] The CPU 201 is a control unit that controls the entire display control device 200 and consists of at least one processor or circuit. The memory 202 consists of, for example, RAM (volatile memory using semiconductor elements). The CPU 201 controls each part of the display control device 200 by using the memory 202 as work memory, for example, according to a program stored in the non-volatile memory 203. The non-volatile memory 203 stores image data, audio data, other data, and various programs for the operation of the CPU 201. Mori 203 is composed of components such as flash memory and ROM.
[0046] The image processing unit 204 performs various image processing operations on images stored in the non-volatile memory 203 and recording medium 208, video signals acquired via the external I / F 209, and images acquired via the communication I / F 210, based on the control of the CPU 201. The image processing operations performed by the image processing unit 204 include A / D conversion, D / A conversion, image data encoding, compression, decoding, resizing, noise reduction, and color conversion. It also performs various image processing operations such as panoramic unfolding, mapping, and conversion of VR images, which are omnidirectional images or wide-area images that have a wide range of video even if they are not omnidirectional. The image processing unit 204 may be composed of dedicated circuit blocks for performing specific image processing operations. Depending on the type of image processing, the CPU 201 may also perform image processing according to a program without using the image processing unit 204.
[0047] The display 205 displays images and GUI (Graphical User Interface) screens based on the control of the CPU 201. The CPU 201 generates display control signals according to the program and controls each part of the display control device 200 to generate video signals for display on the display 205 and output them to the display 205. The display 205 displays images based on the output video signals. The display control device 200 itself is configured to include an interface for outputting video signals for display on the display 205, and the display 205 may be an external monitor (television or HMD).
[0048] The operation unit 206 is an input device for receiving user input, including a keyboard or other text information input device, a mouse or touch panel or pointing device, buttons, dials, joysticks, touch sensors, and touchpads. In this embodiment, the operation unit 206 includes a touch panel 206a and operation units 206b, 206c, 206d, and 206e.
[0049] Recording medium I / F207 allows for the insertion and removal of recording media 208, such as memory cards, CDs, and DVDs. Based on the control of the CPU 201, the recording medium I / F207 reads data from the inserted recording media 208 and writes data to the recording media 208. The recording media 208 is a storage unit that stores data such as images for display on the display 205. The external I / F209 is an interface for connecting to external devices via wired cables (such as USB cables) or wirelessly, and for inputting and outputting video and audio signals (data communication). The communication I / F210 is an interface for communicating (wireless communication) with external devices and the Internet 211, and for sending and receiving various data such as files and commands (data communication).
[0050] The audio output unit 212 outputs audio from video and music data played by the display control device 200, as well as operation sounds, ringtones, and various notification sounds. The audio output unit 212 includes an audio output terminal 212a for connecting earphones, etc., and a speaker 212b, but the audio output unit 212 may also output audio data to an external speaker via wireless communication or other means.
[0051] The attitude detection unit 213 detects the attitude of the display control device 200 relative to gravity, as well as the attitude of the display control device 200 relative to each axis in the yaw, roll, and pitch directions, and notifies the CPU 201 of the attitude information. Based on the attitude detected by the attitude detection unit 213, it is possible to determine whether the display control device 200 is held horizontally, vertically, pointed upwards, pointed downwards, or in an oblique position. It is also possible to determine whether the display control device 200 is tilted in the rotational directions such as the yaw, pitch, and roll directions, and whether the display control device 200 has rotated in the direction of that rotation. Accelerometer, gyroscope One or a combination of sensors, such as a sensor, geomagnetic sensor, compass sensor, and altitude sensor, can be used as the attitude detection unit 213.
[0052] As described above, the operation unit 206 includes a touch panel 206a. The touch panel 206a is an input device that is superimposed on the display 205 to form a planar structure and outputs coordinate information corresponding to the position of contact. The CPU 201 can detect the following operations or states on the touch panel 206a. - A finger or pen that was not previously touching the touch panel 206a now touches the touch panel 206a, i.e., the start of a touch (hereinafter referred to as Touch-Down). • The state in which a finger or pen is touching the touch panel 206a (hereinafter referred to as Touch-On). • The finger or pen is moving while touching the touch panel 206a (hereinafter referred to as Touch-Move). - The finger or pen that was touching the touch panel 206a has been lifted off the touch panel 206a, i.e., the touch has ended (hereinafter referred to as Touch-Up). - When nothing is being touched on the touch panel 206a (hereinafter referred to as Touch-Off)
[0053] When a touchdown is detected, a touch-on is also detected simultaneously. After a touchdown, touch-ons are usually detected continuously unless a touch-up is detected. If a touch-move is detected, a touch-on is also detected simultaneously. Even if a touch-on is detected, a touch-move will not be detected if the touch position has not moved. A touch-off is detected when all fingers or pens that were touching the screen have been detected as having touched up.
[0054] These operations and states, as well as the position coordinates of the finger or pen touching the touch panel 206a, are notified to the CPU 201 via the internal bus. Based on the notified information, the CPU 201 determines what kind of operation (touch operation) was performed on the touch panel 206a. For touch moves, the direction of movement of the finger or pen moving on the touch panel 206a can also be determined for each vertical and horizontal component on the touch panel 206a based on the change in position coordinates. If a touch move of a predetermined distance or more is detected, it is determined that a slide operation has been performed.
[0055] A flick is an operation in which you touch the touch panel 206a with your finger, quickly move it a certain distance, and then release it. In other words, a flick is an operation in which you quickly trace the touch panel 206a with your finger as if flicking it. If a touch move of a certain distance or more at a certain speed or faster is detected, and a touch-up is detected immediately afterward, it can be determined that a flick has occurred (it can be determined that a flick occurred following a slide operation).
[0056] Furthermore, touching multiple points (for example, two points) simultaneously to bring them closer together is called a pinch-in, and touching them further apart is called a pinch-out. Pinch-out and pinch-in are collectively referred to as a pinch operation (or simply a pinch). The touch panel 206a may use any of the various types of touch panels, such as resistive, capacitive, surface acoustic wave, infrared, electromagnetic induction, image recognition, and optical sensor types. There are methods that detect a touch when there is contact with the touch panel, and methods that detect a touch when a finger or pen approaches the touch panel, and either method is acceptable.
[0057] The gaze detection unit 214 is used to detect changes in the direction of the user's gaze (changes in gaze position) and to determine where on the display surface of the display 205 the user is focusing their attention. The CPU 201 identifies the area (region) that the user is focusing on by associating the gaze position detected by the gaze detection unit 214 with the image displayed on the display 205.
[0058] Figure 2(c) is an external view of a VR goggle (head-mounted adapter) 230 to which the display control device 200 can be attached. The display control device 200 can also be used as a head-mounted display by attaching it to the VR goggle 230. The insertion slot 231 is an insertion slot for inserting the display control device 200. The entire display control device 200 can be inserted into the VR goggle 230 with the display surface of the display 205 facing the headband 232 side (i.e., the user side) for fixing the VR goggle 230 to the user's head. The user can view the display 205 of the display control device 200 without holding the display control device 200 with their hands while wearing the VR goggle 230 with the display control device 200 attached on their head. In this case, if the user moves their head or entire body, the posture of the display control device 200 will also change. The posture detection unit 213 detects the change in the posture of the display control device 200 at this time, and the CPU 201 performs processing for VR display based on this change in posture. In this case, the posture detection unit 213 detecting the posture of the display control device 200 is equivalent to detecting the posture of the user's head (the direction the user's gaze is directed). The display control device 200 itself may be an HMD that can be worn on the head without VR goggles.
[0059] Figure 3 is a block diagram showing an example configuration of a distribution device 300, which is a type of information processing device. The distribution device 300 distributes VR images captured by the digital camera 100 and the captured audio (sound) to the viewer's display control device 200. Based on the content that the viewer is paying attention to in the VR image displayed on the display control device 200, the distribution device 300 synthesizes multiple audios captured by multiple digital cameras 100 and distributes the synthesized audio to the viewer. The multiple audios synthesized by the distribution device 300 may include sounds captured by a sound-collecting device other than the digital camera 100 (for example, a sound-collecting device without an imaging unit).
[0060] The CPU 301 controls various parts of the distribution device 300 using memory 303 as work memory, according to the program stored in non-volatile memory 302. Non-volatile memory 302 stores various programs necessary for the operation of the CPU 301. Non-volatile memory 302 is composed of, for example, flash memory or ROM. Memory 303 functions as the main memory of the CPU 301. Memory 303 is composed of, for example, RAM (volatile memory using semiconductor elements).
[0061] The communication interface 304 is an interface for transmitting and receiving data with the digital camera 100 and with the viewer's display control device 200 by communicating with the internet 305 or the like via wireless or wired cable. The image processing unit 306 consists of a dedicated circuit block that performs image processing for distributing VR images. The audio processing unit 307 consists of a dedicated circuit block that performs audio processing for synthesizing multiple audio sources. The CPU 301, non-volatile memory 302, memory 303, communication interface 304, image processing unit 306, and audio processing unit 307 are connected to the internal bus 308. Each part connected to the internal bus 308 can exchange data with each other via the internal bus 308.
[0062] Figure 4 is a block diagram showing an example of a system configuration including an imaging device, a distribution device, and a display control device. The imaging unit 400 and the sound collection unit 401 are functional blocks of the digital camera (imaging device) 100. The imaging unit 400 (imaging units 22a and 22b) captures VR images (moving images) and transmits the captured VR images to the control unit 402. The sound collection unit 401, in synchronization with the capture of VR images, collects sound from the surrounding area where the digital camera 100 is installed and transmits the collected sound to the control unit 402.
[0063] The control unit 402 and the audio generation unit 403 are functional blocks of the distribution device 300. The control unit 402 controls the flow of data in the distribution device 300. The control unit 402 also performs the process of acquiring VR images captured by the imaging unit 400 (moving image acquisition process) and the process of acquiring audio captured by the sound pickup unit 401 (sound acquisition process). The control unit 402 also performs the process of acquiring state information regarding the viewer's attention state to the VR image based on the gaze information detected by the detection unit 405 (information acquisition process). The attention state includes the state of focusing on a specific object and the state of not focusing on any object (overhead view state). The audio generation unit 403 receives the audio captured by the sound pickup unit 401 and the state information based on the gaze information detected by the detection unit 405 via the control unit 402. The audio generation unit 403 synthesizes multiple sounds captured by multiple sound-gathering devices, including the digital camera 100, based on state information, to generate distributed audio (sound to be played along with the VR image) to be delivered to the viewer.
[0064] The viewing unit 404 and the detection unit 405 are functional blocks of the display control device 200. The viewing unit 404 provides the viewer with the VR image and streamed audio input from the control unit 402. The detection unit 405 detects the viewer's gaze information using the gaze detection unit 214 and transmits the gaze information to the control unit 402. The detection unit 405 may also generate state information based on the gaze state and transmit the state information to the control unit 402.
[0065] Figures 5(a) and 5(b) show examples of a viewer viewing a VR image. This embodiment shows an example of distributing VR images and audio recorded at an event venue equipped with a stage and audience seating to viewers via a network. Recording device A500 is configured using a digital camera 100 and is a device that combines the capture of VR images and the recording of audio. By installing devices similar to recording device A500 at multiple locations within the venue, both the capture of surrounding VR images and the recording of audio can be performed at each location. Recording device A500 is the recording device installed in the upper center (in front of stage 505) in Figures 5(a) and 5(b). Similarly, recording device B501 is installed in the upper left (front left of stage 505 as viewed from the audience seating), and recording device C502 is installed in the upper right (front right of stage 505 as viewed from the audience seating). Recording device D503 is located in the lower left section (rear left of the audience seating area), and recording device E504 is located in the lower right section (rear right of the audience seating area). Viewers can select one of these recording devices to view VR images and audio recorded at their desired location. The positional information of each of these recording devices is pre-stored in the non-volatile memory 302 of the distribution device 300. For example, the non-volatile memory 302 stores XY coordinate data, where the horizontal direction is defined as the X-axis and the vertical direction as the Y-axis in Figures 5(a) and 5(b).
[0066] Stage 505 is the stage on which the performers of this event stand. Performers 506 are the people who conduct the event's progress and perform on Stage 505. Audience members 507 are people who are positioned in a location where they can see Stage 505 and Performers 506, and who watch the event in person. Speaker 508 outputs audio related to Performers 506 and the event content, and conveys this content to Performers 506 and Audience Members 507. The audio output by Speaker 508 is also picked up by various recording devices installed in the venue. Viewing range 509 indicates the range of the VR image that a specific viewer is viewing (displayed on the display control device 200). In this embodiment, an example is described in which the viewer selects recording device B501 as the recording device to be viewed. The viewer can view a range of the VR image recorded by recording device B501, corresponding to the viewer's head or entire body posture, as the viewing range 509.
[0067] Figure 5(a) shows an example of a viewer having an overhead view of the VR image. The distribution device 300 shows the viewer having an overall overhead view of the audience 507, and is not focusing on any specific person or object. The system identifies that the user is not looking at the screen based on the gaze detection information detected by the gaze detection unit 214 of the display control device 200.
[0068] Figure 5(b) shows an example of a state in which a viewer is focusing on a specific object. While Figure 5(a) shows an example in which the viewer is looking over the audience 507, Figure 5(b) shows an example in which the viewer is focusing on a specific person. In Figure 5(b), the recording device B501 is selected, as in Figure 5(a), but the viewing range 509 in Figure 5(b) is facing the stage 505, not the audience 507. The distribution device 300 determines that the viewer has been looking at the performer 506 in the VR image for a predetermined amount of time and is focusing on the performer 506, based on the gaze detection information detected by the gaze detection unit 214 of the display control device 200.
[0069] <Delivery Processing> Figure 6 is a flowchart showing an example of the distribution process performed by the CPU 301 of the distribution device 300. In step S600, the CPU 301 processes the recording device (camera) selected by the viewer. The system obtains device information (information about the recording device the viewer wishes to view) related to the image device. At this time, the viewer can choose recording device A500, recording device B501, recording device C502, and recording device D One of the recording devices E504 can be selected as the device to be viewed. The CPU 301 acquires device information via the network. Processing starts from step S600 each time the viewer selects a recording device. When the viewer starts viewing, the device information may be predetermined default information or information about the recording device selected by the viewer.
[0070] In step S601, the CPU 301 determines whether the viewer was focusing on a specific object before changing the selection of the recording device. If the viewer was focusing on a specific object before changing the selection of the recording device, the CPU 301 proceeds to step S603; otherwise, it proceeds to step S602. If the viewer is not focusing on an object (overview state), the CPU 301 proceeds to step S602. In this embodiment, state information regarding the viewer's focus state is stored in memory 303. The CPU 301 can determine whether the viewer was focusing on a specific object based on the state information (information regarding the state before changing the recording device) stored in memory 303.
[0071] In step S602, the CPU 301 performs a determination process for the audio to be distributed based on the overview state. In step S603, the CPU 301 performs a determination process for the audio to be distributed based on the object of interest. The processing in steps S602 and S603 is performed by the audio generation unit 403 of the distribution device 300, and the details will be described later using a separate flowchart.
[0072] In step S604, the CPU 301 delivers the VR image recorded by the recording device selected by the viewer, along with the audio stream determined in step S602 or step S603, to the viewer. The delivery of the VR image and audio stream also includes controlling them to be played back (displayed and output) on the display control device 200 at the delivery destination. In this way, if the viewer changes their selection of the recording device (imaging device), the CPU 301 controls the playback to start from the synthesized audio, depending on whether the viewer was focusing on a specific object before the change.
[0073] The following steps are repeated until the viewer stops viewing. In step S605, the CPU 301 acquires information about the viewing range (viewing range information), which is the area of the recorded VR image that the viewer is viewing. The viewing range is based on information acquired by the attitude detection unit 213 of the display control device 200. For example, the display control device 20 The attitude detection unit 213 of unit 0 acquires information on the direction the viewer is facing and notifies (transmits) it to the distribution device 300. The CPU 301 of the distribution device 300 stores the received information in the memory 303.
[0074] In step S606, the CPU 301 acquires the viewer's gaze information. This gaze information is detected by the gaze detection unit 214 of the display control device 200 and is, for example, XY coordinate information corresponding to the display position of the display 205. Similar to the viewing range information, the CPU 301 stores the gaze information in the memory 303 of the distribution device 300.
[0075] In step S607, the CPU 301 determines the viewer's attention range. The attention range indicates the area within the viewing range that the viewer is paying attention to (where the viewer's gaze position is detected).
[0076] In step S608, the CPU 301 performs a determination process for the distributed audio based on the area of interest. Similar to steps S602 and S603, the processing in step S608 is performed by the audio generation unit 403 of the distribution device 300. Details will be described later using a separate flowchart.
[0077] In step S609, the CPU 301 delivers the audio determined in step S608 to the viewer along with the VR image recorded by the recording device selected by the viewer.
[0078] In step S610, the CPU 301 determines whether the viewer will continue watching. If the viewer will continue watching, the CPU 301 returns to step S605 and continues processing to deliver the VR image and audio to the viewer. On the other hand, if the viewer will not continue watching, the CPU 301 terminates the distribution process.
[0079] Referring to Figure 7, we will now describe in detail the process of determining the audio to be delivered based on the overview state, which is performed in step S602 of Figure 6. Figure 7 is a flowchart showing an example of the process of determining the audio to be delivered based on the overview state.
[0080] In step S700, the CPU 301 identifies recording devices that are within the viewing range (installed within the area being viewed). The CPU 301 identifies recording devices within the current viewing range by comparing the viewing range information stored in memory 303 with the position information of each recording device that is pre-stored in non-volatile memory 302. For example, if the CPU 301 holds the viewing range as angle information, it considers the sector-shaped area that matches the direction indicated by the angle information as the viewing range and determines whether the coordinates of each recording device are included within the viewing range. Whether a particular point is included within the sector-shaped area can be determined by finding the vector of the sector and the point, and the direction vector of the sector, and calculating their dot product.
[0081] In step S701, the CPU 301 equalizes the proportion of audio in the streamed audio that was captured by recording devices within the listening range. Note that the recording devices within the listening range may include recording devices selected by the viewer (recording device B501 in the example of Figure 5(a)).
[0082] In step S702, the CPU 301 calculates the distance from the viewing area to the recording device located outside the viewing area. The CPU 301 may calculate the distance from the center of the viewing area to the recording device, the distance from the edge of the viewing area to the recording device, or the shortest distance.
[0083] In step S703, the CPU 301 reduces the proportion of the streamed audio that was captured by a recording device located outside the listening range, as the distance from the listening range to the recording device increases (it decreases in proportion to the distance).
[0084] In step S704, the CPU 301 identifies a recording device located at a predetermined distance or greater from the area of interest. The distance from the area of interest to the recording device may be the distance from the center of the area of interest, the distance from the edge, or the shortest distance. If the viewer is in an overhead view, the area of interest may be the same as the viewing area.
[0085] In step S705, the CPU 301 detects similar sounds (sounds also recorded by other recording devices) at a predetermined volume level or higher from the sound recorded by the recording device identified in step S704. Whether the sounds are identical (similar) or not may be determined, for example, by determining feature quantities from the frequency components of the sound data and calculating the similarity of the sounds from the difference in feature quantities.
[0086] In step S706, the CPU 301 determines that the audio of the recording device containing the similar audio detected in step S705 is reduced by a predetermined percentage. In step S706, the CPU 301 may reduce the proportion of the audio of the recording device included in the distributed audio so that the volume of that recording device is relatively lower compared to the volumes of other recording devices, or it may reduce the volume of the recording device itself and include it in the distributed audio. This prevents the specific audio from being included in the distributed audio at an excessively high volume, even when a specific audio is picked up by multiple recording devices. For example, in the example in Figure 5(b), recording device B501, C The audio collected by 502 is generated in such a way that the audio emitted from speaker 508 does not stand out excessively.
[0087] In step S707, the CPU 301 synthesizes the multiple audio tracks captured by each recording device at the ratio determined in steps S701 and S703 to determine the audio to be played back together with the VR image. Once the CPU 301 has determined the audio to be played back, it terminates the audio determination process.
[0088] Referring to Figure 8, we will now describe in detail the process of determining the audio to be delivered based on the object of interest, which is performed in step S603 of Figure 6. Figure 8 is a flowchart showing an example of the process of determining the audio to be delivered based on the object of interest.
[0089] In step S800, the CPU 301 determines whether the object the viewer was paying attention to before the viewer changed the selection of the recording device (the object of attention) is displayed within the viewing range of the changed recording device (the current viewing range). In this embodiment, both state information indicating that the viewer is paying attention to a specific object and object information indicating that specific object are stored in the memory 303. The CPU 301 can determine the object of attention before the selection of the recording device was changed from the object information. If the object of attention is displayed within the viewing range, the CPU 301 proceeds to step S801; otherwise, it proceeds to step S807. Alternatively, the CPU 301 may proceed to step S801 regardless of whether the object of attention is displayed within the viewing range.
[0090] In step S801, the CPU 301 calculates the distance (e.g., the shortest distance) from the object of interest displayed in the viewing area to each recording device. The position of the object of interest can be calculated (identified), for example, based on the size of the object of interest recorded in the VR image captured by each recording device. Alternatively, sensors capable of identifying the position may be attached to candidate objects of interest, such as people. The object information described above includes information on the position of the object of interest.
[0091] In step S802, the CPU 301 increases the volume of the audio picked up by the recording device, with a larger increase the shorter the distance from the object of interest to the recording device. In step S802, the CPU 301 may also increase the proportion of the audio from the recording device included in the distributed audio so that the volume of the recording device becomes relatively larger compared to the volumes of other recording devices. Furthermore, the volume of the recording device itself may be increased and included in the streamed audio. In this way, when a viewer changes the recording device, if the subject of interest from before the change is still displayed after the change, the streamed audio is synthesized so that the sound picked up by the recording device near the subject of interest is louder. Since the streamed audio is synthesized so that the sound picked up by the recording device near the subject of interest is louder both before and after the change in recording device, it is possible to avoid abrupt changes in the streamed audio before and after the change in recording device. In addition, if the subject of interest (for example, the vocalist of a band) is holding a dedicated microphone, the CPU 301 may synthesize the streamed audio so that the sound picked up by the dedicated microphone is included at a high volume (for example, the loudest volume).
[0092] The processing in steps S803 to S806 is the same as the processing in steps S704 to S707 in Figure 7.
[0093] In step S807, the CPU 301 determines that the audio captured by the recording device currently selected by the viewer will be used as the streamed audio.
[0094] Referring to Figure 9, we will describe in detail the process of determining the distribution audio based on the area of interest, which is performed in step S608 of Figure 6. Figure 9 is a flowchart showing an example of the process of determining the distribution audio based on the area of interest.
[0095] In step S900, the CPU 301 determines whether the viewer is focusing on a specific object based on the viewing range. If the viewer is focusing on a specific object, the CPU 301 proceeds to step S905; otherwise, it proceeds to step S901. The CPU 301 may determine that the viewer is focusing on a specific object if the position of the specific object in the VR image and the viewing position (the viewing range determined based on the viewing position) coincide for a predetermined time or longer. For example, if the CPU 301 recognizes the face of a person in the VR image being viewed by the viewer, and the position of the person's face and the viewing position coincide for a predetermined time or longer, the CPU 301 determines that the viewer is focusing on that person. On the other hand, if this condition is not met and the viewer cannot be considered to be focusing on a specific object, the CPU 301 determines that the viewer is in an overhead viewing state, looking over the viewing range. At this time, the CPU 301 stores state information regarding the viewer's focusing state in memory 303. When the CPU 301 stores state information indicating that a viewer is paying attention to a specific object in memory 303, it also stores object information indicating that specific object in memory 303.
[0096] The processing in steps S901 to S904 is the same as the processing in steps S700 to S703 in Figure 7. Also, the processing in steps S905 to S910 is the same as the processing in steps S801 to S806 in Figure 8.
[0097] Furthermore, when CPU301 changes the proportion of audio to be included in the streamed audio, it may change the proportion to the target value instantly, or it may change the proportion to the target value gradually so that the streamed audio switches smoothly.
[0098] According to the embodiments of the present invention described above, it is possible to change the sound played according to the viewer's attention state within the VR image. This makes it possible to provide a more immersive video experience than before.
[0099] Furthermore, the various controls described above, which are performed by the distribution device 300, may be performed by a single piece of hardware, or multiple pieces of hardware (for example, multiple processors or circuits) may share the processing to control the entire device. Also, although the above-described embodiment used the application of the present invention to a VR image distribution device at an event venue as an example, the invention is not limited to this example and can be applied to any device related to viewing VR images. This method may be applied not only to VR images, but also to devices for viewing panoramic or multi-view images. Furthermore, the processing described as being performed by the distribution device may be performed as internal processing by a display control device (such as an HMD).
[0100] (Other embodiments) The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.
[0101] The above embodiments are merely examples, and configurations obtained by appropriately modifying or changing the configuration of the above embodiments within the scope of the gist of the present invention are also included in the present invention. Configurations obtained by appropriately combining the configurations of the above embodiments are also included in the present invention. [Explanation of symbols]
[0102] 402: Control Unit 403: Voice Generation Unit
Claims
1. A motion image acquisition means for acquiring multiple motion images captured by multiple sound-collecting devices and multiple imaging devices corresponding to each of them, Sound acquisition means that acquires multiple sounds picked up by the multiple sound pickup devices in synchronization with the acquisition of the multiple moving images, Device information acquisition means for acquiring device information related to the imaging device selected by the viewer, A control means that controls the playback of a video image captured by the imaging device selected by the viewer, based on the device information acquired by the device information acquisition means, A state information acquisition means for acquiring state information relating to the viewer's attention state to the video image being played back, A generation means that generates sound to be played together with the video to be played, based on the state information acquired by the state information acquisition means, It has, The generating means is If the state information indicates that the viewer is viewing the video being played, the proportion of multiple sounds picked up by multiple sound-collecting devices installed within the area being viewed by the viewer is made equal, and the proportion of sounds picked up by sound-collecting devices installed outside the area is made smaller as the distance from the area to the sound-collecting device increases, and the multiple sounds are synthesized to generate sound to be played along with the video. If the state information indicates that the viewer is focusing on a specific object included in the video being played back, the volume of the sound picked up by the sound picking device is increased by a larger amount the shorter the distance from the object the viewer is focusing on to the sound picking device, and the multiple sounds are synthesized to generate a sound to be played back together with the video. If there is a change in the selection of imaging device by the aforementioned viewer, Before the image capture device is modified, the state information acquisition means acquires information as state information indicating that the viewer is paying attention to a specific object included in the video before the modification, and if the modified video includes the specific object, the volume of the sound picked up by the sound picking device is increased by a larger amount the shorter the distance from the object the viewer is paying attention to to the sound picking device, and the multiple sounds are synthesized to generate a sound to be played back together with the modified video. Before the change of the imaging device, the state information acquisition means acquires information as state information indicating that the viewer is paying attention to a specific object included in the video before the change, and if the video after the change does not include the specific object, the sound acquired by the sound acquisition device corresponding to the changed imaging device is determined to be played back together with the video after the change, without synthesizing the multiple sounds. An information processing device characterized by the following:
2. The state information acquisition means acquires, as state information, information indicating whether or not the viewer is viewing the video being played from above. The information processing apparatus according to feature 1.
3. The state information acquisition means acquires, as state information, information indicating whether or not the viewer is paying attention to a specific object included in the video being played. The information processing apparatus according to claim 1 or 2.
4. When synthesizing the plurality of sounds, the generation means reduces the volume of sounds that are picked up at a volume higher than a predetermined volume from the range that the viewer is paying attention to, among the sounds picked up by a sound-picking device installed at a predetermined distance or more away from the range that the viewer is paying attention to, and then synthesizes the plurality of sounds. The information processing apparatus according to any one of claims 1 to 3.
5. If, before the change of the imaging device, the status information acquisition means acquires information indicating that the viewer is viewing the video image before the change, The generation means equalizes the proportion of multiple sounds picked up by multiple sound-collecting devices installed within the area viewed by the viewer, and decreases the proportion of sounds picked up by sound-collecting devices installed outside the area as the distance from the area to the sound-collecting device increases, and synthesizes the multiple sounds to generate sound to be played back together with the modified video. The information processing apparatus according to any one of claims 1 to 4.
6. The state information acquisition means acquires the state information based on the viewer's line of sight. The information processing apparatus according to any one of claims 1 to 5.
7. The system further includes determination means for determining whether the state information indicates that the viewer is viewing the video being played from above, or that the viewer is focusing on a specific object. The information processing apparatus according to feature 1.
8. If the viewer is not focusing on any object, the state information indicates that the viewer is viewing the video being played from above. The information processing apparatus according to feature 1.
9. If the viewer cannot be considered to be focusing on the specific object, the state information indicates that the viewer is viewing the video being played from above. The information processing apparatus according to feature 1.
10. If the position of the specific object included in the video being played back coincides with the viewer's line of sight for a predetermined period of time or longer, the state information indicates that the viewer is paying attention to the specific object. The information processing apparatus according to feature 1.
11. If the position of the person's face in the video being played back coincides with the viewer's line of sight for a predetermined period of time or longer, the state information indicates that the viewer is paying attention to the person. The information processing apparatus according to feature 1.
12. A determination means for determining a characteristic quantity of each of the plurality of sounds from the frequency components of the sound, and determining the similarity of the plurality of sounds based on the difference in the characteristic quantities among the plurality of sounds, A detection means for detecting similar sounds that are sounds picked up by two or more of the aforementioned plurality of sound-picking devices, based on the similarity of the plurality of sounds, It further possesses, When synthesizing the multiple sounds, the generation means reduces the volume of similar sounds that have been recorded at a volume higher than a predetermined volume, and then synthesizes the multiple sounds. The information processing apparatus according to any one of claims 1 to 11.
13. When the specific object is holding a sound-collecting device, when synthesizing the plurality of sounds, the generating means synthesizes the sounds picked up by the sound-collecting device held by the specific object at the highest volume. The information processing apparatus according to any one of claims 1 to 12.
14. A motion image acquisition means for acquiring multiple motion images captured by multiple sound-collecting devices and multiple imaging devices corresponding to each of them, Sound acquisition means that acquires multiple sounds picked up by the multiple sound pickup devices in synchronization with the acquisition of the multiple moving images, Device information acquisition means for acquiring device information related to the imaging device selected by the viewer, A control means that controls the playback of a video image captured by the imaging device selected by the viewer, based on the device information acquired by the device information acquisition means, A state information acquisition means for acquiring state information relating to the viewer's attention state to the video image being played back, A generation means that generates sound to be played together with the video to be played, based on the state information acquired by the state information acquisition means, It has, When the state information indicates that the viewer is viewing the video being played back, the generation means, even if the region of the video being viewed by the viewer does not contain a specific sound-gathering device corresponding to the imaging device selected by the viewer, will equalize the ratio of multiple sounds picked up by multiple sound-gathering devices set within the region to the sound picked up by the specific sound-gathering device, and will decrease the ratio of sounds picked up by sound-gathering devices different from the specific sound-gathering device, which are installed outside the region, as the distance from the region to the sound-gathering device increases, thereby synthesizing the multiple sounds to generate a sound to be played back with the video. An information processing device characterized by the following:
15. A video acquisition step of acquiring multiple video images captured by multiple sound-collecting devices and multiple corresponding imaging devices, A sound acquisition step in which multiple sounds are acquired by the multiple sound acquisition devices in synchronization with the acquisition of the multiple moving images, A device information acquisition step to obtain device information about the imaging device selected by the viewer, A control step which controls the playback of a video image captured by the imaging device selected by the viewer, based on the device information acquired in the device information acquisition step, A state information acquisition step, which acquires state information relating to the viewer's attention state to the video image being played back, A generation step which generates sound to be played together with the video to be played, based on the state information acquired in the state information acquisition step, It has, In the above generation step, If the state information indicates that the viewer is viewing the video being played, the proportion of multiple sounds picked up by multiple sound-collecting devices installed within the area being viewed by the viewer is made equal, and the proportion of sounds picked up by sound-collecting devices installed outside the area is made smaller as the distance from the area to the sound-collecting device increases, and the multiple sounds are synthesized to generate sound to be played along with the video. If the state information indicates that the viewer is focusing on a specific object included in the video being played back, the volume of the sound picked up by the sound picking device is increased by a larger amount the shorter the distance from the object the viewer is focusing on to the sound picking device, and the multiple sounds are synthesized to generate a sound to be played back together with the video. If there is a change in the selection of imaging device by the aforementioned viewer, Before the change of the imaging device, in the state information acquisition step, information is acquired as state information indicating that the viewer is paying attention to a specific object included in the video before the change, and if the video after the change includes the specific object, the volume of the sound picked up by the sound picking device is increased by a larger amount the shorter the distance from the object the viewer is paying attention to to the sound picking device, and the multiple sounds are synthesized to generate a sound to be played back together with the video after the change. Before the change of the imaging device, in the state information acquisition step, information is acquired as state information indicating that the viewer is paying attention to a specific object included in the video before the change, and if the video after the change does not include the specific object, then, without synthesizing the multiple sounds, the sound picked up by the sound pickup device corresponding to the changed imaging device is determined to be played back together with the video after the change. A control method for an information processing device characterized by the following features.
16. A program for causing a computer to function as one of the means of an information processing apparatus according to any one of claims 1 to 14.
17. A computer-readable recording medium storing a program for causing a computer to function as one of the means of an information processing device according to any one of claims 1 to 14.
18. A system comprising a plurality of imaging devices, a plurality of sound receiving devices corresponding to each of the plurality of imaging devices, and an information processing device, The aforementioned information processing device is A motion image acquisition means for acquiring multiple motion images captured by the aforementioned multiple imaging devices, Sound acquisition means that acquires multiple sounds picked up by the multiple sound pickup devices in synchronization with the acquisition of the multiple moving images, Device information acquisition means for acquiring device information related to the imaging device selected by the viewer, A control means that controls the playback of a video image captured by the imaging device selected by the viewer, based on the device information acquired by the device information acquisition means, A state information acquisition means for acquiring state information relating to the viewer's attention state to the video image being played back, A generation means that generates sound to be played together with the video to be played, based on the state information acquired by the state information acquisition means, It has, The generating means is If the status information indicates that the viewer is viewing the video being played, then multiple sound-gathering devices installed within the area of the video being viewed by the viewer will be used to record the video. The proportions of multiple sounds picked up are kept the same, and the proportion of sounds picked up by a sound-picking device installed outside the area is reduced as the distance from the area to the sound-picking device increases, and the multiple sounds are synthesized to generate sound to be played back together with the video. If the state information indicates that the viewer is focusing on a specific object included in the video being played back, the volume of the sound picked up by the sound picking device is increased by a larger amount the shorter the distance from the object the viewer is focusing on to the sound picking device, and the multiple sounds are synthesized to generate a sound to be played back together with the video. If there is a change in the selection of imaging device by the aforementioned viewer, Before the image capture device is modified, the state information acquisition means acquires information as state information indicating that the viewer is paying attention to a specific object included in the video before the modification, and if the modified video includes the specific object, the volume of the sound picked up by the sound picking device is increased by a larger amount the shorter the distance from the object the viewer is paying attention to to the sound picking device, and the multiple sounds are synthesized to generate a sound to be played back together with the modified video. Before the change of the imaging device, the state information acquisition means acquires information as state information indicating that the viewer is paying attention to a specific object included in the video before the change, and if the video after the change does not include the specific object, the sound acquired by the sound acquisition device corresponding to the changed imaging device is determined to be played back together with the video after the change, without synthesizing the multiple sounds. A system characterized by the following features.
Citation Information
Patent Citations
Viewing device, and underwater space viewing system and method for viewing the same
JP2018113653A
Presentation system and presentation method pertaining to large scale repair construction work for multi-dwelling housing
JP2021068265A
Sound collection and reproduction system, sound collection and reproduction apparatus, sound collection and reproduction method, sound collection and reproduction program, sound collection system, and reproduction system
US20160021478A1
Selective audio reproduction
US20170347219A1
Spatial audio processing emphasizing sound sources close to a focal distance
US20190174246A1