Information processing apparatus, control method of the same, and program

JP2024056580A5Active Publication Date: 2025-10-14CANON KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022163588
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-10-11
Publication Date
2025-10-14
Estimated Expiration
2042-10-11

AI Technical Summary

Technical Problem

Performers have difficulty grasping the position of a virtual viewpoint outside their visual field, making it challenging to perform and direct effectively in virtual viewpoint video productions.

Method used

An information processing device that generates stereophonic sound data based on the virtual viewpoint position, allowing performers to perceive the virtual viewpoint through audio cues using headphones, thereby presenting the virtual viewpoint position outside their visual field.

Benefits of technology

Enables performers to accurately perceive the virtual viewpoint position through audio, enhancing their performance and direction in virtual viewpoint video productions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To allow a user to perceive a virtual viewpoint position outside a user's visual field.SOLUTION: An information processing apparatus in a system for generating a virtual viewpoint image from images acquired by a plurality of imaging devices has: an acquisition unit for acquiring position information of a virtual viewpoint; a generation unit for generating acoustic data corresponding to sound emitted from a position corresponding to the acquired virtual viewpoint based on the acquired position information of the virtual viewpoint; and a reproduction unit for reproducing acoustic data to a reproduction device.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an information processing apparatus, a control method thereof, and a program. [Background technology]

[0002] In recent years, a technology has become known in which an object is imaged from multiple starting positions, and an image of the object seen from a virtual viewpoint is generated using the images from the multiple viewpoints obtained (hereafter referred to as a virtual viewpoint image or virtual viewpoint video). For example, by filming the performance of an actor who is the object and generating a virtual viewpoint video, it is possible to generate an image from a position that cannot be approached by a real camera or from a dynamic camera angle. It is also possible to synthesize the virtual viewpoint video of the object with a virtual background space (for example, a background generated by computer graphics) to generate an image in which the object appears to be in a different space. In this way, the technology of virtual viewpoint video makes it possible to provide viewers with images that have a higher sense of realism than normal images. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2022-50305 Summary of the Invention [Problem to be solved by the invention]

[0004] However, since the virtual viewpoint does not exist in real space, it is difficult for the performer, who is the subject, to grasp the position of the virtual viewpoint or the line of sight at the virtual viewpoint. Therefore, it is difficult for the performer to perform and direct while being conscious of the virtual viewpoint. To solve this problem, Patent Document 1 discloses a technology that visually presents the position of the virtual viewpoint to the subject, but the subject can only know the position of the virtual viewpoint within his or her own field of vision.

[0005] The present invention has been made in consideration of such problems, and aims to provide a technique for allowing a user to perceive a virtual viewpoint position that is outside the user's field of view. [Means for solving the problem]

[0006] In order to solve this problem, for example, an information processing device of the present invention has the following arrangement. An information processing device in a system that generates a virtual viewpoint video from images acquired by a plurality of imaging devices, An acquisition means for acquiring position information of a virtual viewpoint; a generating means for generating, based on the acquired position information of the virtual viewpoint, sound data corresponding to a sound emitted from a position corresponding to the acquired virtual viewpoint; and a playback means for causing a playback device to play back the audio data. Effect of the Invention

[0007] According to the present invention, it is possible to allow a user to perceive a virtual viewpoint position that is outside the user's visual field. [Brief description of the drawings]

[0008] [Figure 1] FIG. 2 is a diagram showing the configuration of a system for generating a virtual viewpoint video according to the first embodiment. [Diagram 2] FIG. 2 is a functional configuration diagram of an information processing device according to the first embodiment. [Diagram 3] FIG. 4A is a diagram showing the processing flow of the stereophonic sound generating unit of the first embodiment, and FIG. 4B is a diagram related to the definition of the positional relationship between a virtual viewpoint and a subject. [Figure 4] FIG. 2 is a hardware configuration diagram of the information processing apparatus according to the first embodiment. [Diagram 5] FIG. 11 is a functional configuration diagram of an information processing device according to a second embodiment. [Figure 6] FIG. 11 is a diagram for explaining a processing flow of an information processing device according to the second embodiment. [Figure 7] FIG. 13 is a functional configuration diagram of an information processing device according to a third embodiment. [Figure 8]FIG. 13 is a diagram for explaining processing of an input signal generating unit according to the third embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0009] Hereinafter, the embodiments will be described in detail with reference to the attached drawings. Note that the following embodiments do not limit the invention according to the claims. Although the embodiments describe a number of features, not all of these features are essential to the invention, and the features may be combined in any manner. Furthermore, in the attached drawings, the same reference numbers are used for the same or similar configurations, and duplicated descriptions are omitted.

[0010] [First embodiment] FIG. 1 is a system configuration diagram in the first embodiment. A plurality of cameras 11a to 11e are arranged to surround a subject (performer) 12. Each of the cameras 11a to 11e captures moving images at, for example, 30 frames per second. Each of the cameras 11a to 11e is connected to a network 15. This network 15 may be wired or wireless. An information processing device 1 is connected to the network 15, which generates a virtual viewpoint image representing a view from a virtual viewpoint 13 and presents (notifies) the subject 12 of the position of the virtual viewpoint 13. The subject 12 wears earphones (or headphones) 14.

[0011] In the illustrated system, the number of cameras is five, but there is no particular limit to this number. The more cameras there are, the fewer blind spots there will be, and the more accurate the virtual viewpoint video can be created. In addition, communication between the earphones 14 worn by the subject 12 and the information processing device 1 may be either wired or wireless, but in the embodiment, a wireless communication technology (e.g., Bluetooth (registered trademark)) is used from the viewpoint of the ease of movement of the subject.

[0012] In the system having the above configuration, during imaging to generate (or record) a virtual viewpoint video, the position of the virtual viewpoint 13 is presented to the subject, and in this embodiment, the subject is made to perceive the position (and therefore the direction) of the virtual viewpoint 13 using a stereophonic technology reproduced by the earphones 14. For this purpose, the information processing device 100 uses the stereophonic technology to generate sound data in a binaural manner so that a sound source exists at the position of the virtual viewpoint 12, and transmits the sound data to the earphones 14. More specifically, the information processing device 100 obtains a corresponding function from a head-related transfer function measured in advance based on the position information of both ears of the subject 12 and the position information or direction information of the virtual viewpoint, and converts an input signal that is a playback sound source into stereophonic data using the head-related transfer function. As a result, the subject 12 is able to perceive the position or direction of the sound source (the position or direction of the virtual viewpoint) by listening to the stereophonic sound from the earphones 14, rather than relying on vision.

[0013] FIG. 2 shows a functional configuration diagram of an information processing device 1 according to this embodiment.

[0014] The information processing device 1 has a video input unit 100, a shape estimation unit 101, a virtual viewpoint video rendering unit 102, a video output unit 103, a virtual viewpoint acquisition unit 200, a subject ear position information acquisition unit 301, an input signal acquisition unit 302, a stereophonic sound generation unit 303, and a sound output unit 304. Note that, although an example in which the generation process of the virtual viewpoint video and the generation process of the sound data representing the virtual viewpoint position are realized by one device is shown in the embodiment, this may be performed by independent devices.

[0015] Hereinafter, the operation of each component of the information processing device 1 will be described to explain the characteristic configuration of the embodiment.

[0016] The cameras 11a to 11e assign frame numbers and time codes to each frame constituting the video data obtained by synchronously shooting, and transmit the frames to the information processing device 1 via the network 15. The image input unit 100 receives image data for each frame from the cameras 11a to 11e, and supplies the received image data to the shape estimation unit 101 and the subject's ear position information acquisition unit 301.

[0017] The shape estimation unit 101 estimates the three-dimensional shape of the subject using image data from the cameras 11a to 11e and camera information related to the installation position, imaging direction, and optics of each camera acquired in advance. The camera information is composed of external parameters such as the position and line of sight direction of the corresponding camera (any of 11a to 11e) in three-dimensional space, and internal parameters representing the optical center and focal length of the camera. The three-dimensional shape of this embodiment is a group of voxels as an element group representing the three-dimensional shape of the subject in three-dimensional space. The shape estimation process of the shape estimation unit 101 in this embodiment may use existing technology, or may be an estimation process using, for example, a visual volume intersection method or a stereo method. The shape estimation unit 101 supplies the three-dimensional shape of the subject obtained by the estimation process to the virtual viewpoint video drawing unit 102.

[0018] The virtual viewpoint acquisition unit 200 acquires virtual viewpoint information for drawing a virtual viewpoint. The virtual viewpoint information includes at least optical information such as the position (coordinates), attitude (including line of sight direction) and angle of view of the virtual camera in real space. The virtual viewpoint information is associated with a frame number or a time code attached to a captured image. The virtual viewpoint information is generated by, for example, a user operating an input device such as a mouse or a keyboard. Alternatively, virtual viewpoint information including time information describing the position, attitude, etc. of the virtual viewpoint along the time axis may be stored in a storage device such as a HDD in advance and used.

[0019] The virtual viewpoint video rendering unit 102 renders an image of a subject seen from a virtual viewpoint indicated by the virtual viewpoint information acquired by the virtual viewpoint acquisition unit 200, using the image data input by the video input unit 100 and the three-dimensional shape of the subject from the shape estimation unit 101. It is assumed that the virtual viewpoint video rendering unit 102 in this embodiment uses a known technique, and a description thereof will be omitted.

[0020] The subject's ear position information acquisition unit 301 acquires the positions of both ears of the subject 12 based on the captured image by the imaging device 100, and outputs the acquired positions to the stereophonic sound generation unit 303. The positions of both ears are estimated, for example, by detecting the areas of the subject's ears in the captured image and back-projecting the detected areas from each imaging camera to estimate the positions of the subject's ears in three-dimensional space.

[0021] Also, in cases where the ears of the subject 12 are covered by headphones or the like, the position of the subject's head or torso may be estimated, and the position information of both ears may be estimated based on the relative positional relationship between the ears and the head or torso. A pattern matching method or a deep neural network is used as a means for detecting the subject's ears, head, or torso from an image. Also, if headphones are worn, the position of the subject's ears can be estimated from the positions of the left and right speaker parts of the headphones, and this may be utilized. In this case, the left and right speaker parts of the headphones may be identified by, for example, making the left and right speaker parts different in color, or by using the speaker parts to have appropriate marks for identifying the left and right. Also, when multiple subjects are present, a number or coordinate information of the subject is added to the acquired ear position information so as to associate it with the subject.

[0022] The input signal acquisition unit 302 acquires an input signal for stereophonic sound generated by the stereophonic sound generation unit 303. The input signal is assumed to be a prepared audio file, but may be sampled data of a user's voice directly uttered into a microphone, or audio data (such as a beep) with a frequency within a specific audible range.

[0023] The stereophonic sound generating unit 303 generates stereophonic sound by binaurally generating audio signals from the input signal acquiring unit 302 based on the information on the position of the virtual viewpoint from the virtual viewpoint acquiring unit 200 and the position information of the subject's both ears from the subject's ear position information acquiring unit 301. The binaural method is a technology that simulates the propagation characteristics of sound from a sound source to the left and right ears for an input signal, allowing a listener to perceive that sound is coming from a sound source in any direction. In this embodiment, a head-related transfer function is used to generate audio data that represents the propagation characteristics of sound from the sound source to the left and right ears of the subject, with the position of the virtual viewpoint as the sound source.

[0024] The process of the stereophonic sound generating unit 303 in this embodiment will be described below with reference to the flowchart in FIG. 3(a) and the image diagram in FIG. 3(b).

[0025] In S301, the stereophonic sound generating unit 303 calculates the direction and distance from both ears of the subject to the position of the virtual viewpoint.

[0026] For example, in the case of the arrangement in FIG. 3(b), the stereophonic sound generator 303 calculates the directional angle θ, elevation angle φ, and distance γ from the center position O of the subject's ears to the sound source S representing the virtual viewpoint based on the coordinate information of both.

[0027] In S302, the stereophonic sound generating unit 303 obtains the head-related transfer functions of the left and right ears corresponding to the direction and distance calculated in S301 from a database of measurements taken in advance.

[0028] In S303, the stereophonic sound generating unit 303 generates audio signals for both ears in order to generate a stereophonic sound signal. The stereophonic sound generating unit 303 then outputs the generated audio signals for the left and right ears via the audio output unit 304 to the left and right speakers of the earphones 14 worn by the subject 12 for playback.

[0029] In this embodiment, the stereophonic sound generating unit 303 performs a Fourier transform on the signal to reduce the amount of calculation, and performs calculations in the frequency domain. First, the stereophonic sound generating unit 303 performs a Fourier transform on the audio signal acquired by the input signal acquiring unit 302 and the head related transfer function by S302. Then, the stereophonic sound generating unit 303 performs a convolution operation on the audio signal after the Fourier transform and the head related transfer function after the Fourier transform for each of the two ears. Finally, the stereophonic sound generating unit 303 performs an inverse Fourier transform on the result of the convolution operation, and generates an audio signal for each of the two ears. The overlap-add method or the overlap-safe method may be used for the convolution operation. In a database of head related transfer functions measured in advance, the measurements are generally not taken in all directions, but are taken by sampling at regular intervals. Therefore, there may be cases where there is no corresponding head related transfer function for the direction and distance information by S301. In that case, the stereophonic sound generating unit 303 acquires at least two head related transfer functions of the nearest sampled points, generates stereophonic signals for each of the two ears in S303, and synthesizes them into one audio signal by linear correction. When multiple subjects are present, the stereophonic sound generating section 303 generates a stereophonic sound signal corresponding to each subject based on the subject number or coordinate information associated with the ear position information from the subject ear position information acquiring section 301.

[0030] The audio output unit 304 transmits the stereophonic signals (sound signals of the left and right channels) generated by the stereophonic generation unit 303 to the earphones 14 worn by the subject for playback. In the embodiment, the earphones 14 and the information processing device 1 are connected wirelessly, and therefore the earphones 14 actually include a circuit, an amplifier, and a speaker for receiving wireless audio signals. However, this is not the main focus of the embodiment, and therefore a description thereof will be omitted.

[0031] In the above embodiment, the virtual viewpoint is one, but multiple virtual viewpoints can be applied. In other words, a virtual sound source corresponding to an independent virtual viewpoint may be generated for each subject. In this case, multiple virtual viewpoint acquisition units 200, multiple virtual viewpoint image rendering units 102, multiple input signal acquisition units 302, and multiple stereophonic sound generation units 303 are used.

[0032] Next, the hardware configuration of the information processing device 1 will be described with reference to Fig. 4. The information processing device 1 has a calculation unit for performing image processing and three-dimensional shape generation, including a GPU (Graphics Processing Unit) 410 and a CPU (Central Processing Unit) 411. The information processing device 1 also has a ROM (Read Only Memory) 412, a RAM (Random access memory) 413, and an auxiliary storage device 414. The information processing device 1 further has a display unit 415, an operation unit 416, a communication I / F 417, and a bus 418.

[0033] When the power supply of this device is turned on, the CPU 411 executes an OS (operating system) loader in the ROM 412, loads the OS from the auxiliary storage device 414 into the RAM 413, and starts the OS. As a result, the display unit 415 and the operation unit 416 start functioning. Then, under the control of the OS, the CPU 411 loads an application program from the auxiliary storage device 414 into the RAM 413 and executes it, thereby causing the CPU 411 to function as each of the functional units shown in FIG. 2. The CPU 411 also operates as a display control unit that controls the display unit 415, and an operation control unit that controls the operation unit 416.

[0034] The GPU 410 can perform efficient calculations by processing more data in parallel. Therefore, in the first embodiment, the GPU 410 is used in addition to the CPU 411 for some of the processes of the shape estimation unit 101, the virtual viewpoint video rendering unit 102, and the subject ear position information acquisition unit 301. When executing a program, calculations may be performed by only one of the CPU 411 and the GPU 410, or the CPU 411 and the GPU 410 may work together to perform calculations.

[0035] The information processing device 1 may have one or more dedicated hardware components different from the CPU 411, and the dedicated hardware components may execute at least a part of the processing performed by the CPU 411. Examples of the dedicated hardware components include an ASIC (application specific integrated circuit), an FPGA (field programmable gate array), and a DSP (digital signal processor).

[0036] The ROM 412 stores programs that do not require modification. The RAM 413 temporarily stores programs and data supplied from an auxiliary storage device 414, and data supplied from the outside via a communication I / F 417. The auxiliary storage device 414 is constituted by, for example, a hard disk drive, and stores various data such as image data and audio data.

[0037] The display unit 415 is configured with, for example, a liquid crystal display or LEDs, and displays a GUI (Graphical User Interface) for the user to operate the image generating device. The operation unit 416 is configured with, for example, a keyboard, a mouse, a joystick, a touch panel, and receives operations by the user and inputs various instructions to the CPU 411.

[0038] The communication I / F 417 is used for communication with devices external to the image generation device. For example, when the image generation device is connected to an external device via a wired connection, a communication cable is connected to the communication I / F 417. When the image generation device has a function for wireless communication with an external device, the communication I / F 417 includes an antenna. The bus 418 connects each part of the image generation device to transmit information.

[0039] In this embodiment, the subject wears an audio reproduction device (earphones 14) that allows the subject to hear stereophonic sound. The information processing device 1 in this embodiment generates stereophonic sound using a virtual viewpoint as a virtual sound source, and reproduces the stereophonic sound by the audio reproduction device worn by the subject. As a result, the subject can perform while understanding that the virtual viewpoint exists at the position of the sound source. Furthermore, the subject (performer) can recognize the virtual viewpoint in all directions, not limited to the visual field, and does not need to constantly track the visual information presented.

[0040] In the above embodiment, one device has the function of inputting a virtual viewpoint, the function of generating sound data, and the function of generating a virtual viewpoint image. However, among these functions, the function of generating sound data and the function of generating a virtual viewpoint image may be realized by an independent device. In the above embodiment, the subject's ear position information acquisition unit 301 has been described as estimating the positions of the subject's left and right ears from the images of the cameras 11a to 11e. However, for example, a gyro sensor, an ultrasonic sensor, or the like may be stored in the headphones worn on the subject's head, and the left and right ear positions of the subject may be estimated from the signals from these sensors. Although the distance between the actual ears of the subject varies from person to person, no particular problem occurs even if the distance is considered to be the average distance for an adult. The above also applies to the second and third embodiments described below.

[0041] [Second embodiment] In generating a virtual viewpoint video, it is necessary to process a plurality of captured images, and it may take time from the moment of capturing images until the virtual viewpoint video is rendered. In this embodiment, after a virtual viewpoint is presented to a subject during capturing images, the virtual viewpoint is delayed according to a processing speed, and the virtual viewpoint video is rendered. The configuration of an information processing device 2 according to this second embodiment will be described with reference to FIG. 5. Hereinafter, differences between this second embodiment and the first embodiment described above will be mainly described, and the same reference symbols will be used for the same configuration as the first embodiment, and description thereof will be omitted.

[0042] As shown in FIG. 5, the difference between the second embodiment and the first embodiment is that a delay adjustment unit 201 is added. For ease of explanation, the video input unit 100, the shape estimation unit 101, the virtual viewpoint video drawing unit 102, and the video output unit 103 are collectively referred to as a virtual viewpoint video generation unit 10, and the subject ear position information acquisition unit 301, the input signal acquisition unit 302, the stereophonic sound generation unit 303, and the sound output unit 304 are collectively referred to as a subject support unit 30. Note that this system may be configured by one electronic device or may be configured by multiple electronic devices. The second embodiment has the same hardware configuration of the video generation device as the first embodiment shown in FIG. 4.

[0043] Next, the operation of the delay adjustment unit 201 of the second embodiment will be described.

[0044] The delay adjustment unit 201 first accumulates the virtual viewpoint information (the position of the virtual viewpoint, the line of sight direction, the angle of view, etc.) acquired by the virtual viewpoint acquisition unit 200. Next, the delay adjustment unit 201 inputs time information such as a frame number or a time code attached to a three-dimensional shape or a captured image to be rendered by the virtual viewpoint video rendering unit 102. Then, the delay adjustment unit 201 searches for virtual viewpoint information having the same time information as the input time information from among the accumulated virtual viewpoint information, and outputs the corresponding virtual viewpoint information to the virtual viewpoint video rendering unit 102. In addition, the delay adjustment unit 201 may delay the virtual viewpoint by the virtual viewpoint acquisition unit 200 based on the processing speed of the virtual viewpoint video generation unit 10 and the speed of data transmission and reception between each component, and output the information to the virtual viewpoint video rendering unit 102. The delay amount may be a value determined in advance according to the performance of the processing device, or may change dynamically according to the amount of data of each component. Typically, the delay adjustment unit 201 may supply the next virtual viewpoint information to be drawn each time the virtual viewpoint video rendering unit 102 completes rendering of one frame of the virtual viewpoint video.

[0045] The flow of processing of the second information processing device will be described with reference to FIG. 6, for example, when the shape estimation unit 101 takes up the processing time of the virtual viewpoint video generation unit 10. Here, the shape estimation unit 101 has a plurality of calculation units, and the shape estimation process can be performed in parallel. For successive times T_1, T_2, T_n, T_(n+1), T_(n+2), etc., the imaging device 100, the virtual viewpoint acquisition unit 200, the subject support unit 30, the delay adjustment unit 201, the shape estimation unit 101, the virtual viewpoint video rendering unit 102, and the video output device 103 are processed at each time. The unit of time T is the time per frame according to the frame rate of the imaging device 100 (in the embodiment, the cameras 11a to 11e are assumed to capture images at 30 FPS, so 1 / 30 seconds).

[0046] For example, at time T_1, the video input unit 100 inputs images captured by the cameras 11a to 11e at time T_1. The virtual viewpoint acquisition unit 200 acquires a virtual viewpoint at time T_1, and the subject support unit 30 generates audio data based on the virtual viewpoint as if a sound source were at the virtual viewpoint position, and reproduces the audio data through the earphones 14 (audio reproduction device) worn by the subject 12. The virtual viewpoint at time T_1 is then input to the delay adjustment unit 201, and the processing of the shape estimation unit 101 is also started. The processing of the shape estimation unit 101 cannot be completed within T_1, and the processing ends at T_n. At T_n, the imaging device 100, the virtual viewpoint acquisition unit 200, and the subject support unit 30 process the data at time T_n, and the delay adjustment unit 201 accumulates the virtual viewpoint at time T_n obtained by the virtual viewpoint acquisition unit 200. Furthermore, at time Tn, the shape estimation unit 101 finishes processing the data at time T_1, and the virtual viewpoint video rendering unit 102 starts rendering at time T_1, and acquires and renders the virtual viewpoint at time T_1 to the delay adjustment unit 201. The video output unit 103 outputs the virtual viewpoint video rendered by the virtual viewpoint video rendering unit 102. The output destination of the video output unit 103 may be a display device, or may be a device that accumulates the virtual viewpoint video as a video file.

[0047] In order to reduce the overall processing delay, the number of processes that can be executed in parallel by the shape estimation unit 101 can be increased, so that the delay time of the virtual viewpoint video by the virtual viewpoint video rendering unit 102 can be substantially reduced to the time required for one shape estimation process (=Tn-T1). This allows the delay time from virtual viewpoint input to generation of the virtual viewpoint video to be reduced to the time required for one shape estimation process (=Tn-T1). If the time required for one shape estimation process by the shape estimation unit 101 can be further reduced, it becomes possible to set a sound source at the virtual viewpoint position in real time, play back audio data, and generate a virtual viewpoint video with an unnoticeable delay time.

[0048] [Third embodiment] Due to the mechanism of human hearing, humans will not mistakenly recognize the left or right position of a sound source using stereophonic sound, but there is a possibility that they may mistakenly recognize the front or back. In the third embodiment, stereophonic data is generated to enable the subject 12 to more accurately recognize information about the virtual viewpoint.

[0049] FIG. 7 is a functional configuration diagram of the information processing device 3 in the third embodiment. The difference from the first embodiment (FIG. 2) is in the input signal generating unit 305. Therefore, the configuration other than the input signal generating unit 305 is the same as that in FIG. 2 of the first embodiment, so it is indicated by the same reference numerals and the description thereof is omitted. Note that this system may be configured by one electronic device or may be configured by multiple electronic devices. Also, like the first embodiment, the third embodiment has the hardware configuration shown in FIG. 4. Also, the delay adjustment unit 201 of the second embodiment may be added to the third embodiment.

[0050] The operation of the input signal generating unit 305 in the third embodiment will be described.

[0051] The input signal generation unit 305 generates an input signal based on the information indicating the virtual viewpoint position acquired by the virtual viewpoint acquisition unit 200 and the information indicating the subject's ear position acquired by the subject's ear position information acquisition unit 301, and outputs the input signal to the stereophonic sound generation unit 303. In the third embodiment, there is no limitation on the type of input signal to be generated, but the following configuration may be used as an example.

[0052] As shown in FIG. 8(a), a virtual viewpoint S(θ,φ,γ) is located at a direction angle θ, an elevation angle φ, and a distance γ with respect to the center position of the positions of the ears of the subject 12. In this case, the input signal generating unit 305 estimates the gaze direction (or front direction) L(θ',φ') of the subject 12 based on the position information of the ears by the subject's ear position information acquiring unit 301 and the captured image. The input signal is a periodically sounded beep sound. The input signal generating unit 305 changes the frequency f of the beep sound to be output according to the magnitude of the difference between the gaze direction L(θ',φ') of the subject and the position S(θ,φ,γ) of the virtual viewpoint. For example, as shown in FIG. 8(b), the normal part may be set to the time when the difference is 0 according to the difference between the direction angle and the elevation angle. As a result, the closer the gaze of the subject is to the virtual viewpoint, the more frequently the beep sound is sounded, and the subject can determine that the virtual viewpoint is in the direction in which he or she is facing.

[0053] The third embodiment can avoid misrecognition caused by the mechanism of human hearing and improve the accuracy of recognizing the virtual viewpoint of the subject using the stereophonic sound. The strength of the sound to be reproduced may be changed depending on the difference between the line of sight L(θ',φ') of the subject and the position S(θ,φ,γ) of the virtual viewpoint.

[0054] <Other embodiments> In the above-described first to third embodiments, the virtual viewpoint is presented to the subject by an audio reproduction device, such as earphones 14, worn by the subject, but this can also be realized by other devices. For example, multiple speakers may be fixedly installed at preset positions in the real space where the subject may exist, and stereophonic sound may be reproduced by these speakers. In this case, the same virtual viewpoint position can be presented to each of the multiple subjects.

[0055] The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions.

[0056] The disclosure of this specification includes the following information processing device, control method thereof, and program. (Item 1) An information processing device in a system that generates a virtual viewpoint video from images acquired by a plurality of imaging devices, An acquisition means for acquiring position information of a virtual viewpoint; a generating means for generating, based on the acquired position information of the virtual viewpoint, sound data corresponding to a sound emitted from a position corresponding to the acquired virtual viewpoint; a playback means for playing back the audio data on a playback device; 13. An information processing device comprising: (Item 2) a first estimation means for estimating the positions of the left and right ears of the subject, The generating means generates acoustic data having a sound source at the position of the virtual viewpoint from the positions of the left and right ears estimated by the first estimating means and the position of the virtual viewpoint acquired by the acquiring means. 2. The information processing device according to item 1, (Item 3) The first estimation means estimates the positions of the left and right ears of the subject from the images obtained by the plurality of image capture devices. 3. The information processing device according to item 2. (Item 4) the playback device is an earphone or a headphone that is worn on the subject's head, The earphone or headphone has a sensor for detecting an orientation and a position, The first estimation means estimates the positions of the left and right ears of the subject based on the information from the sensor. 3. The information processing device according to item 2. (Item 5) The playback device is a plurality of speakers installed in the real space surrounding the subject. 2. The information processing device according to item 1, (Item 6) A second estimation means for estimating a gaze direction of the subject, The generating means generates acoustic data according to a difference between the line of sight of the subject estimated by the second estimating means and a direction from the head of the subject toward the virtual viewpoint. 6. The information processing device according to any one of items 1 to 5. (Item 7) Further, a means for generating a virtual viewpoint image representing a view from the position of the virtual viewpoint from the images obtained by the plurality of image pickup devices. 7. The information processing device according to any one of items 1 to 6, comprising: (Item 8) A method for controlling an information processing device in a system that generates a virtual viewpoint video from images acquired by a plurality of imaging devices, comprising: An acquisition step of acquiring position information of a virtual viewpoint; a generating step of generating, based on the acquired position information of the virtual viewpoint, acoustic data corresponding to a sound emitted from a position corresponding to the acquired virtual viewpoint; a reproduction step of reproducing the sound data on a reproduction device; 13. A method for controlling an information processing apparatus comprising the steps of: (Item 9) A program that, when read and executed by a computer, causes the computer to execute each step of the method according to item 8.

[0057] The invention is not limited to the above-described embodiments, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]

[0058] 1, 2, 3... Information processing device, 10... Virtual viewpoint video generation unit, 30... Subject support unit, 100... Video input unit, 101... Shape estimation unit, 102... Virtual viewpoint video drawing unit, 103... Video output unit, 200... Virtual viewpoint acquisition unit, 201... Delay adjustment unit, 301... Subject ear position information acquisition unit, 302, 305... Input signal acquisition unit, 303... Stereophonic sound generation unit, 304... Sound output unit

Claims

1. An information processing device in a system that generates a virtual viewpoint video from images acquired by a plurality of imaging devices, an acquisition means for acquiring position information of a virtual viewpoint; a generating means for generating, based on the acquired position information of the virtual viewpoint, sound data corresponding to a sound emitted from a position corresponding to the acquired virtual viewpoint; a playback means for causing a playback device to play back the sound data; An information processing device comprising:

2. Further comprising a first estimation means for estimating the positions of the left and right ears of the subject; The generating means generates acoustic data having a sound source at the position of the virtual viewpoint from the positions of the left and right ears estimated by the first estimating means and the position of the virtual viewpoint acquired by the acquiring means.

2. The information processing apparatus according to claim 1, wherein:

3. The first estimation means estimates the positions of the left and right ears of the subject from the images acquired by the plurality of image capture devices.

3. The information processing apparatus according to claim 2, wherein:

4. the playback device is an earphone or a headphone worn on the subject's head, The earphone or headphone has a sensor for detecting direction and position, The first estimation means estimates the positions of the left and right ears of the subject based on information from the sensor.

3. The information processing apparatus according to claim 2, wherein:

5. The playback device is a plurality of speakers installed in the real space surrounding the subject.

2. The information processing apparatus according to claim 1, wherein:

6. Further comprising a second estimation means for estimating the gaze direction of the subject, The generating means generates acoustic data according to a difference between the line of sight of the subject estimated by the second estimating means and a direction from the head of the subject toward the virtual viewpoint.

2. The information processing apparatus according to claim 1, wherein:

7. Furthermore, a means for generating a virtual viewpoint image representing a view from the position of the virtual viewpoint from the images obtained by the plurality of image pickup devices.

2. The information processing apparatus according to claim 1, further comprising:

8. A control method for an information processing device in a system that generates a virtual viewpoint video from images acquired by a plurality of imaging devices, comprising: an acquisition step of acquiring position information of a virtual viewpoint; a generating step of generating, based on the acquired position information of the virtual viewpoint, sound data corresponding to a sound emitted from a position corresponding to the acquired virtual viewpoint; a playback step of playing back the sound data on a playback device; 1. A method for controlling an information processing device, comprising:

9. A program that, when read and executed by a computer, causes the computer to execute each step of the method according to claim 8.