Information processing device, information processing method, and program
Patent Information
- Application Number
- JP2023004704
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-01-16
- Publication Date
- 2026-01-08
AI Technical Summary
Existing technologies do not effectively reproduce sound data, such as advertisements, when generating virtual viewpoint images from virtual viewpoints.
An information processing device that includes viewpoint acquisition, management, generating, and output means to synchronize virtual viewpoint images with sound data from virtual sound sources arranged in a three-dimensional space, allowing for the reproduction of sound data corresponding to virtual advertisements.
Enables the playback of sound data, particularly advertisements, in conjunction with virtual viewpoint images, enhancing user experience and advertising effectiveness.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to an information processing device, an information processing method, and a program, and more particularly to a technology for generating a virtual viewpoint video accompanied by audio data. [Background technology]
[0002] A technology for generating a virtual viewpoint image from a virtual viewpoint in a three-dimensional space is attracting attention. Such a three-dimensional space can be constructed based on a plurality of captured images of a subject (for example, Patent Document 1). For example, a plurality of captured images of a subject in a real space can be obtained using an imaging device such as a camera arranged at a plurality of different positions. Then, an image processing unit such as a server can construct a 3D model of the subject based on the plurality of captured images. In addition, the image processing unit can perform rendering processing of the 3D model of the subject and the virtual viewpoint image based on the virtual viewpoint. Then, the obtained virtual viewpoint image is delivered to a user terminal and displayed on the user terminal. With such a configuration, a user can observe a subject from various angles. Using such a technology, a virtual viewpoint video composed of a plurality of virtual viewpoint images can also be generated.
[0003] Meanwhile, advertisement delivery in virtual reality space has been attracting attention. For example, Patent Document 2 relates to a technology in which a user operates an avatar in a virtual reality space. In Patent Document 2, audio data such as an advertisement is output to a user depending on whether the user's avatar corresponds to a target. Specifically, the delivery of advertisement audio data is controlled depending on the area in which the avatar is located or the attributes of the avatar. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] JP 2019-36790 A [Patent Document 2] Patent Publication No. 2022-12034 Summary of the Invention [Problem to be solved by the invention]
[0005] When generating a virtual viewpoint image from a virtual viewpoint, it is also desirable to reproduce sound data such as advertisements.
[0006] An object of the present invention is to reproduce sound data such as advertisements when a virtual viewpoint image from a virtual viewpoint is displayed. [Means for solving the problem]
[0007] An information processing device according to an embodiment includes: A viewpoint acquisition means for acquiring information specifying a virtual viewpoint in a three-dimensional space in which one or more virtual sound sources are arranged; a management means for managing sound data corresponding to each of the one or more virtual sound sources; a generating means for generating a virtual viewpoint image of the three-dimensional space from the virtual viewpoint; an output means for outputting the virtual viewpoint image and the sound data corresponding to at least one of the one or more virtual sound sources; Equipped with. Effect of the Invention
[0008] When a virtual viewpoint image from a virtual viewpoint is displayed, sound data such as advertisements can be reproduced. [Brief description of the drawings]
[0009] [Figure 1] FIG. 1 is a block diagram showing an example of a functional configuration of an information processing device according to an embodiment. [Diagram 2] FIG. 4 is a diagram showing the data structure of data corresponding to a virtual sound source. [Diagram 3] FIG. 1 is a block diagram showing an example of a hardware configuration of an information processing apparatus according to an embodiment. [Figure 4] 1 is a flowchart of an information processing method according to an embodiment. [Diagram 5]1 is a diagram showing a method for selecting a virtual sound source based on the line of sight of a virtual viewpoint. [Figure 6] FIG. 13 illustrates a method for selecting a virtual sound source based on its distance from a virtual viewpoint. [Figure 7] 1 is a flowchart of an information processing method according to an embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0010] Hereinafter, the embodiments will be described in detail with reference to the attached drawings. Note that the following embodiments do not limit the invention according to the claims. Although the embodiments describe a number of features, not all of these features are essential to the invention, and the features may be combined in any manner. Furthermore, in the attached drawings, the same reference numbers are used for the same or similar configurations, and duplicated descriptions are omitted.
[0011] In one embodiment, one or more virtual sound sources are placed in a three-dimensional space. The virtual sound sources may be associated with advertisements. For example, in a distribution service of virtual viewpoint images, a video of a sport can be distributed. In this case, advertising can be performed on the video. In this case, a virtual advertisement can be placed at an intended position in the three-dimensional space. For example, a virtual object can be placed on a wall of a model such as a stadium placed in the three-dimensional space, or on a predetermined area such as a seating area. Furthermore, a CG image of an advertisement can be attached to the virtual object. This makes it possible to include a virtual advertisement in the virtual viewpoint image that the viewer watches.
[0012] Furthermore, in one embodiment, sound data corresponding to each of one or more virtual sound sources is managed. Therefore, when a virtual viewpoint image is displayed, it is possible to play back the associated sound data. According to such an embodiment, it is possible to provide richer content in a system that displays a virtual viewpoint image. In particular, in a configuration in which sound data is linked to a virtual advertisement, the advertising effect can be improved by outputting a combination of a visual advertisement and an audio advertisement.
[0013] (Embodiment 1) In one embodiment in which one or more virtual sound sources are arranged in a three-dimensional space, selected sound data from among the sound data corresponding to each virtual sound source can be reproduced. In the first embodiment, at least one virtual sound source is selected from the one or more virtual sound sources based on information specifying a virtual viewpoint. Then, the sound data corresponding to the selected virtual sound source is output.
[0014] 1 is a block diagram showing an example of the configuration of an image processing system according to an embodiment of the present invention. The image processing system includes an imaging system 1 and an information processing device 2.
[0015] The imaging system 1 includes a plurality of imaging devices. Each of the plurality of imaging devices is arranged at a different position so as to surround an imaging area to be imaged. The plurality of imaging devices can acquire a multi-viewpoint image by synchronously performing imaging. The imaging system 1 may have a control unit for synchronizing the captured images. The synchronization method is not particularly limited. The plurality of imaging devices may not be installed over the entire circumference of the imaging area. For example, due to restrictions on the installation location, they may be installed only in an area in a specific direction from the imaging area. The number of imaging devices is not limited. Furthermore, the imaging system 1 may include imaging devices having different functions, such as a telephoto camera and a wide-angle camera.
[0016] Furthermore, the imaging system 1 may have a sound collection unit such as a microphone. The imaging system 1 can generate sound data based on data collected by the sound collection unit. One or more or all of the multiple imaging devices may have such a sound collection unit. Furthermore, the imaging system 1 may have a sound collection unit independent of the multiple imaging devices.
[0017] The multi-viewpoint images obtained by the imaging system 1 are transmitted to a model generating unit 101 of the information processing device 2. Hereinafter, the multiple imaging devices included in the imaging system 1 may be simply referred to as the imaging system 1.
[0018] The information processing device 2 includes a model generation unit 101, a sound source management unit 102, a sound source placement unit 103, a viewpoint setting unit 105, a sound source selection unit 106, a sound data control unit 107, and an output unit .
[0019] The model generation unit 101 generates a 3D model of the subject from the multiple viewpoint images received from the imaging system 1. Furthermore, the model generation unit 101 transmits the generated 3D model of the subject to the output unit .
[0020] In one embodiment, the 3D model refers to a foreground image and a background image extracted from each of the multiple viewpoint images. A 3D model of an object can be constructed based on the foreground image for the multiple viewpoint images. The foreground image and the background image can be generated, for example, in the following manner. First, the model generation unit 101 extracts a foreground image from a foreground region corresponding to a predetermined object from each of the multiple viewpoint images. In addition, the model generation unit 101 extracts a background image from a background region that is a region other than the foreground region.
[0021] The subject extracted as the foreground region may be a dynamic object (moving body) that moves (its absolute position or shape may change) when capturing images in a time series from the same direction. Such a subject may be a person such as a player or a referee on the field where a game is played. In a ball game, the subject may include a ball. The subject may also be a singer, a player, a performer, or a presenter in a concert or entertainment.
[0022] The background region may be an object that remains stationary or nearly stationary when time-series images are captured from the same direction. Such an object may be, for example, a stage for a concert or the like, or a stadium where an event such as a competition is held. Furthermore, such an object may be a structure such as a goal used in a ball game, or a field, etc. The background region is at least a region different from the foreground region. The multi-viewpoint image may include other objects, etc., in addition to the subject corresponding to the foreground region and the object corresponding to the background region.
[0023] The sound source management unit 102 manages sound data corresponding to one or more virtual sound sources. In this embodiment, one or more virtual sound sources are arranged in a three-dimensional space. The sound source management unit 102 can further manage the positions (e.g., coordinates) of each of the one or more virtual sound sources. The virtual sound source is arranged in the same three-dimensional space as a virtual viewpoint described later. The coordinates of the virtual sound source can indicate coordinates in the same three-dimensional space as the virtual viewpoint image. In this way, the sound source management unit 102 can manage sound data associated with the position of the virtual sound source for each virtual sound source. In one embodiment, the position of the virtual sound source is set as the generation position of the sound data corresponding to the virtual sound source.
[0024] Furthermore, the sound source management unit 102 can manage models arranged at the positions of one or more virtual sound sources. Hereinafter, a model arranged at the position of a specific virtual sound source may be simply referred to as a virtual sound source. The type of model is not particularly limited, and may be a 2D object or a 3D object. The model may be an object with CG texture data attached. The model may also be an avatar with material information attached. In one embodiment, at least one of the one or more virtual sound sources is associated with a model of an advertisement arranged in a three-dimensional space. That is, an object that is an advertisement is arranged at the position of the virtual sound source. On the other hand, at least one of the one or more virtual sound sources may not be associated with a model arranged in a three-dimensional space. That is, an object may not be arranged at the position of the virtual sound source. The position of the virtual sound source may also be determined in advance, regardless of the position of the subject or its 3D model.
[0025] FIG. 2 shows an example of a data structure showing data corresponding to each virtual sound source. An index number (also written as pattern No.) is assigned to each virtual sound source arranged in a three-dimensional space. The pattern No. is assigned according to the number of patterns of the virtual sound source to be displayed in the virtual viewpoint image. This data structure also includes position information (Position) and attitude information of the virtual sound source for each virtual sound source. The attitude information includes the object's orientation (Rotation) and scale (Scale). This data structure also includes information (OBJECT) indicating the object arranged at the position of the virtual sound source for each virtual sound source. Furthermore, this data structure also includes, for each virtual sound source, appearance information of the object such as CG texture data or material information applied to the object (CG / MATERIAL). Furthermore, this data structure stores sound data associated with each virtual sound source. Hereinafter, object attitude information, material information, texture data, and other object setting information will be collectively referred to as object information. Information such as the volume, length, or sampling rate of the sound data can also be stored together with the sound data.
[0026] The sound source placement unit 103 places a virtual sound source in a three-dimensional space. Specifically, the sound source placement unit 103 places a model corresponding to the virtual sound source at the position of each of one or more virtual sound sources. Specifically, the sound source placement unit 103 reads the position information and posture information of the virtual sound source, the object information, the appearance information of the object, and the sound data managed by the sound source management unit 102. After the reading is completed, the sound source placement unit 103 places the object to which the appearance information of the object is applied, and the sound data in the same three-dimensional space as the virtual viewpoint video. The sound source placement unit 103 transmits the placement information to the sound source selection unit 106 and the output unit 108.
[0027] The viewpoint setting unit 105 acquires information that specifies a virtual viewpoint in a three-dimensional space. The viewpoint setting unit 105 can acquire this information from, for example, the user input unit 104. The user input unit 104 is an input device connected to the information processing device 2 directly or via a network. The input device may be a general device used by a user to perform an input operation, such as a controller, a keyboard, or a mouse. The input device may also be a tracking controller such as a head mounted display (HMD).
[0028] The user can input information indicating the position and orientation of a virtual viewpoint (also called a virtual camera) via the user input unit 104. Specifically, the user can input a parameter set including a parameter indicating the position of the virtual camera in a three-dimensional space and a parameter indicating the orientation of the virtual camera in the pan, tilt, and roll directions. However, the content of the information input by the user is not limited to these. For example, the parameter set may include a parameter indicating the size of the field of view (angle of view) of the virtual camera. Note that the virtual camera is a virtual camera that is different from a plurality of imaging devices that are actually installed around the imaging area. The virtual camera is a concept for explaining a virtual viewpoint used to generate a virtual viewpoint video.
[0029] The viewpoint setting unit 105 sets a virtual viewpoint based on the parameter set of the virtual viewpoint received from the user input unit 104. That is, the viewpoint setting unit 105 can place a virtual viewpoint having an attitude indicated in the parameter set at a position indicated in the parameter set. Then, the viewpoint setting unit 105 transmits placement information of the set virtual viewpoint to the sound source selecting unit 106 and the output unit 108.
[0030] The sound source selection unit 106 selects at least one virtual sound source from one or more virtual sound sources. The sound source selection unit 106 can select at least one virtual sound source based on information that specifies a virtual viewpoint. In this embodiment, sound data is reproduced from the virtual sound source selected by the sound source selection unit 106. In this embodiment, the sound source selection unit 106 selects a virtual sound source based on at least one of the position and the orientation of the virtual viewpoint. Since the virtual viewpoint image changes according to the virtual viewpoint, such a configuration makes it possible to output sound data that matches the virtual viewpoint image.
[0031] Specifically, the sound source selection unit 106 receives placement information of the object and sound data placed at the position of the virtual sound source from the sound source placement unit 103. The sound source selection unit 106 also receives placement information of the virtual viewpoint from the viewpoint setting unit 105. Then, the sound source selection unit 106 selects a virtual sound source based on the received information. The sound source selection unit 106 transmits selection information indicating the selected virtual sound source to the sound data control unit 107.
[0032] In one embodiment, sound data corresponding to a virtual sound source included in the field of view of the virtual viewpoint selected from one or more virtual sound sources is output. For this purpose, the sound source selection unit 106 can select at least one virtual sound source included in the field of view of the virtual viewpoint. With this configuration, sound data related to the generated virtual viewpoint image can be reproduced, improving the user experience.
[0033] In one embodiment, sound data corresponding to a virtual sound source selected based on at least one of the position of a virtual viewpoint and the line of sight of the virtual viewpoint is output from one or more virtual sound sources. For example, the sound source selection unit 106 may select a virtual viewpoint based on the relationship between the line of sight of the virtual viewpoint and the direction from the virtual viewpoint to the virtual sound source. In one example, the direction from the virtual viewpoint to the one virtual sound source selected is closer to the line of sight of the virtual viewpoint than the direction from the virtual viewpoint to a virtual sound source other than the one virtual sound source selected from the multiple virtual sound sources. Note that the two directions being closer means that the angle (defined to be 180° or less) between the two directions is closer. According to such a configuration, sound data related to an area close to the center of the generated virtual viewpoint image, i.e., an area focused on by the user, can be reproduced, improving the user experience.
[0034] The sound source selection unit 106 may select a virtual viewpoint based on the positional relationship between the virtual viewpoint and the virtual sound source. In one example, the distance from the virtual viewpoint to the one virtual sound source selected is shorter than the distance from the virtual viewpoint to the virtual sound sources other than the one virtual sound source selected from the multiple virtual sound sources. With this configuration, sound data related to the position of the virtual viewpoint can be reproduced, improving the user experience.
[0035] Furthermore, the sound source selection unit 106 may switch the selection method of a plurality of virtual sound sources. For example, the sound source selection unit 106 may switch the selection method of at least one virtual sound source from one or a plurality of virtual sound sources depending on whether the position of the virtual viewpoint is within a predetermined area. With such a configuration, it becomes easier to select a virtual sound source according to the virtual viewpoint, particularly when the distribution of the virtual sound sources within the predetermined area is different from that outside the predetermined area. For example, such a configuration can be used when the virtual sound sources are arranged around the predetermined area. A specific example of the selection method of the virtual sound source will be described later with reference to the flowchart of FIG. 4.
[0036] The sound data control unit 107 outputs management information for controlling the reproduction of sound data to the output unit 108. In this embodiment, the sound data control unit 107 manages the reproduction and stopping of sound data when a virtual viewpoint image is displayed according to the selection information received from the sound source selection unit 106. For example, during reproduction of sound data corresponding to a certain virtual sound source, the selection information of the virtual sound source received from the sound source selection unit 106 may change to indicate another virtual sound source. In this case, the sound data control unit 107 can switch the sound data to be output from sound data corresponding to a first virtual sound source among the multiple virtual sound sources to sound data corresponding to a second virtual sound source among the multiple virtual sound sources. That is, the sound data control unit 107 can transmit management information to the output unit 108 so as to switch the virtual sound source and the sound data corresponding thereto. In addition, the sound data control unit 107 may control the volume of the sound data to be output when switching the sound data. For example, when switching the sound data, the sound data control unit 107 may control the reproduction of the sound data so as to impart a sound effect such as fade-out to the sound data to be terminated. Furthermore, the sound data control unit 107 may control the reproduction of sound data so as to impart a sound effect, such as a fade-in, to newly reproduced sound data.
[0037] The sound data control unit 107 may control the sound data to be output based on the location information of the virtual sound source and the location information of the virtual viewpoint. For example, the sound data control unit 107 may control the volume of the sound data based on the distance between the virtual sound source and the virtual viewpoint. The sound data control unit 107 may also control the output of the sound data so that the sound data is heard from the direction of the virtual sound source. However, it is not essential that the sound data control unit 107 performs these controls. For example, the volume of the sound data corresponding to the virtual sound source selected by the sound source selection unit 106 does not need to change according to the positional relationship between the virtual sound source and the virtual viewpoint.
[0038] The output unit 108 outputs a virtual viewpoint image and sound data corresponding to at least one virtual sound source among one or more virtual sound sources. This sound data is data that can be reproduced together with the virtual viewpoint image. The output unit 108 can output a virtual viewpoint video including a time-series virtual viewpoint image and sound data. Here, the output unit 108 can output sound data corresponding to at least one virtual sound source selected by the sound source selection unit 106. The output unit 108 can output the obtained virtual viewpoint image to the display device 3. The display device 3 is, for example, a user terminal. The display device can display the received virtual viewpoint image. In addition, the output unit 108 can output the sound data to the display device 3 or other devices.
[0039] Specifically, the output unit 108 receives information from the model generation unit 101, the sound source placement unit 103, the sound data control unit 107, and the viewpoint setting unit 105. Then, the output unit 108 renders a virtual viewpoint image based on the received 3D model and virtual viewpoint placement information. The output unit 108 can also output sound data corresponding to the virtual sound source selected by the sound source selection unit 106. Specifically, the output unit 108 can output sound data according to the virtual sound source placement information and sound data management information.
[0040] The output unit 108 can generate a virtual viewpoint image, for example, by the following method. As described above, a virtual sound source is arranged in the three-dimensional space. Also, a 3D model of a subject in a real space is arranged in the three-dimensional space. The 3D model of the subject can be generated based on captured images of the subject captured by each of a plurality of imaging devices. Such captured images can be obtained by using the imaging system 1 described above. Also, a background model may be further arranged in the three-dimensional space. The output unit 108 can render a virtual viewpoint image from a virtual viewpoint arranged in such a three-dimensional space by using a ray tracing method or the like.
[0041] Specifically, the output unit 108 generates a foreground model representing a three-dimensional shape of a predetermined subject based on the foreground image. The method of generating the foreground model is not particularly limited, and may be, for example, a volume intersection method. The output unit 108 also generates texture data used to color the foreground model based on the foreground image. Furthermore, the output unit 108 generates texture data used to color a background model representing a three-dimensional shape of the background of a stadium or the like based on the background image. The output unit 108 then colors the foreground model and the background model by mapping the texture data. The output unit 108 then generates a virtual viewpoint image by performing rendering according to the virtual viewpoint indicated by the virtual viewpoint arrangement information. However, the method of generating the virtual viewpoint image is not limited to such a method. For example, the output unit 108 may generate a virtual viewpoint image by projective transformation of a captured image without using a three-dimensional model.
[0042] 3 is a block diagram showing an example of a hardware configuration of the information processing device 2. The information processing device 2 includes a CPU 301, a ROM 302, a RAM 303, an auxiliary storage device 304, a display unit 305, an operation unit 306, a communication I / F 307, and a bus 308.
[0043] The CPU 301 realizes each function of the information processing device 2 shown in Fig. 2 by controlling the entire information processing device 2 using a computer program or data stored in the ROM 302 or the RAM 303. The information processing device 2 may have one or more dedicated hardware pieces different from the CPU 301. In this case, at least a part of the processing by the CPU 301 can be executed by the dedicated hardware. Examples of the dedicated hardware include an ASIC (application specific integrated circuit), an FPGA (field programmable gate array), and a DSP (digital signal processor).
[0044] ROM 302 is a non-volatile memory. ROM 302 can store programs that do not require modification. RAM 303 is a primary storage memory. RAM 303 can temporarily store programs or data supplied from auxiliary storage device 304, or data supplied from the outside via communication I / F 307. Auxiliary storage device 304 is, for example, a hard disk drive. Auxiliary storage device 304 can store various data such as image data or sound data.
[0045] The display unit 305 can display characters or images. The display unit 305 is, for example, a liquid crystal display or an LED. The display unit 305 can display a GUI (Graphical User Interface) used by a user to operate the information processing device 2. The operation unit 306 can accept a user operation and input an instruction according to the user operation to the CPU 301. The operation unit 306 can be, for example, a keyboard, a mouse, a joystick, or a touch panel. The CPU 301 can also operate as a display control unit that controls the display unit 305 and an operation control unit that controls the operation unit 306. FIG. 3 illustrates the display unit 305 and the operation unit 306 as being present inside the information processing device 2. However, at least one of the display unit 305 and the operation unit 306 may be another device outside the information processing device 2.
[0046] The communication I / F 307 is used for the information processing device 2 to communicate with an external device. For example, when the information processing device 2 is connected to the external device by wire, a communication cable is connected to the communication I / F 307. When the information processing device 2 has a function of wirelessly communicating with the external device, the communication I / F 307 includes an antenna. The bus 308 connects each unit of the information processing device 2 and transmits information.
[0047] An information processing method according to an embodiment will be described below with reference to the flowchart shown in FIG. 4. In the following process, the information processing device 2 selects a virtual sound source that reproduces sound data from among virtual sound sources in the same three-dimensional space as the virtual viewpoint video. Then, the information processing device 2 renders the virtual viewpoint video. In the following process, a virtual viewpoint video including a multi-viewpoint image of a subject at one time and a time-series virtual viewpoint image of a subject based on a virtual viewpoint whose position or posture changes in time series is obtained. In this way, the information processing device 2 can generate a virtual viewpoint video of a stationary subject from a moving virtual viewpoint. On the other hand, the information processing device 2 may generate a virtual viewpoint video of a moving subject from a stationary virtual viewpoint. Also, the information processing device 2 may generate a virtual viewpoint video of a moving subject from a moving virtual viewpoint. In these cases, the model generation unit 101 can generate a 3D model of the subject at each time based on the multi-viewpoint image of the subject in time series in S401 described later. By using such a 3D model of the subject at each time, the output unit 108 can generate a virtual viewpoint video of the moving subject.
[0048] It is not necessary that the image capturing time of the subject and the sound data are synchronized. For example, the sound data may be reproduced together with a virtual viewpoint video of the subject at a specific time based on a virtual viewpoint whose position or posture changes over time. Also, the sound data may be reproduced together with a virtual viewpoint image (i.e., a still image) of the subject at a specific time based on a specific virtual viewpoint.
[0049] In the following example, the sound source selection unit 106 selects only one virtual sound source from among a plurality of virtual sound sources. Then, the output unit 108 outputs sound data corresponding to the selected one virtual sound source. However, the sound source selection unit 106 may select two or more virtual sound sources. In this case, the output unit 108 can output sound data corresponding to the selected two or more virtual sound sources as sound data to be simultaneously reproduced. Furthermore, the sound data control unit 107 can control the volume of the sound data corresponding to each of the two or more virtual sound sources. For example, the sound data control unit 107 may control the volume of the sound data corresponding to the virtual sound source based on the relationship between the line of sight of the virtual viewpoint and the direction from the virtual viewpoint to the virtual sound source, or based on the positional relationship between the virtual viewpoint and the virtual sound source. In a further embodiment, the sound data control unit 107 may control the reproduction of the sound data so that the sound data corresponding to the virtual sound source selected by the sound source selection unit 106 is reproduced at a louder volume than the sound data corresponding to the virtual sound source not selected by the sound source selection unit 106.
[0050] Furthermore, the sound source selection unit 106 may be capable of switching the virtual sound source in response to a user selection. In this case, the output unit 108 can output sound data corresponding to at least one virtual sound source selected by the user among one or more virtual sound sources. For example, after the sound source selection unit 106 selects a virtual sound source by the above-mentioned method, the sound source selection unit 106 may change the virtual sound source that reproduces sound data to the virtual sound source selected by the user.
[0051] In S401, the imaging system 1 transmits multiple viewpoint images to the model generation unit 101 of the information processing device 2. The model generation unit 101 generates a 3D model based on the received multiple viewpoint images. Then, the model generation unit 101 transmits the 3D model to the output unit .
[0052] In S402, the sound source management unit 102 sends information indicating an object corresponding to the position information of each stored virtual sound source, appearance information of the object, and sound data to the sound source placement unit 103. The sound source placement unit 103 places the object and sound data corresponding to the virtual sound source received from the sound source management unit 102 in the same three-dimensional space as the virtual viewpoint video.
[0053] In S403, the sound source arrangement unit 103 sends the position information and sound data of each virtual sound source arranged in S402 to the output unit 108 via the sound source selection unit 106 and the sound data control unit 107.
[0054] In S404, the viewpoint setting unit 105 receives a parameter set of the virtual viewpoint from the user input unit 104. The viewpoint setting unit 105 also sets the position and orientation of the virtual viewpoint based on the parameter set of the virtual viewpoint. Then, the viewpoint setting unit 105 transmits information indicating the position and orientation of the virtual viewpoint to the sound source selection unit 106 and the output unit 108.
[0055] Here, the viewpoint setting unit 105 determines whether the position of the virtual viewpoint is a predetermined area. The predetermined area is, for example, an imaging area. The imaging area is an area in which a foreground model in a three-dimensional space is arranged. The information processing device 2 can generate a foreground model of a subject in a predetermined area in real space based on an image captured by the imaging system 1, and the imaging area is an area corresponding to this predetermined area. This imaging area may be set in advance. For example, the imaging area may be set by parameters based on coordinates or the like set at the time of imaging. On the other hand, the imaging area may be specified separately.
[0056] If the viewpoint setting unit 105 determines that the position of the virtual viewpoint is within the imaging area, the process of S405 is executed. If not, the process of S406 is executed. In this way, the virtual sound source is selected in different ways depending on whether the position of the virtual viewpoint is within the imaging area or not.
[0057] In S405, the output unit 108 outputs a virtual viewpoint video. As already described, the output unit 108 can generate a virtual viewpoint video by performing rendering based on the 3D model, the arrangement information of the virtual sound source, the management information of the sound data, and the arrangement information of the virtual viewpoint. Here, the output unit 108 outputs sound data corresponding to at least one virtual sound source selected from one or more virtual sound sources.
[0058] In S405, the sound source selection unit 106 selects at least one virtual sound source included in the virtual viewpoint image. Specifically, the sound source selection unit 106 can determine the line of sight direction of the virtual viewpoint and the viewing angle of the virtual viewpoint based on the arrangement information received from the viewpoint setting unit 105. Based on such a line of sight direction and viewing angle of the virtual viewpoint, the sound source selection unit 106 can select a virtual sound source included in the viewing field of the virtual viewpoint from among the virtual sound sources arranged in the three-dimensional space. The sound source selection unit 106 may select a virtual sound source such that a model arranged at the position of the virtual sound source is at least partially included in the viewing field of the virtual viewpoint. Then, the sound source selection unit 106 transmits selection information indicating the selected virtual sound source to the sound data control unit 107. The sound data control unit 107 transmits management information to the output unit 108 so that sound data corresponding to the virtual sound source selected by the sound source selection unit 106 is reproduced according to the received selection information.
[0059] Further, the sound source selection unit 106 selects a virtual viewpoint based on the relationship between the line of sight of the virtual viewpoint and the direction from the virtual viewpoint to the virtual sound source. FIG. 5 is a diagram for explaining such a method of selecting a virtual sound source. The virtual viewpoint and the virtual sound source are arranged in a three-dimensional space, but for the sake of explanation, a two-dimensional diagram will be used below. FIG. 5 shows the position and attitude of a virtual viewpoint 501 according to an input in the user input unit 104. A line of sight direction 502 of the virtual viewpoint 501 corresponds to the front direction of the virtual viewpoint. A viewing area 503 corresponds to the viewing angle of the virtual viewpoint 501. The generated virtual viewpoint image is an image within the viewing area 503. The virtual sound sources 504, 505, and 506 are arranged at the corresponding positions by the sound source arrangement unit 103. The virtual advertisements 507, 508, and 509 are models corresponding to the virtual sound sources 504, 505, and 506 arranged at the positions of the virtual sound sources 504, 505, and 506.
[0060] The sound source selection unit 106 can select a virtual sound source close to the line of sight 502 of the virtual viewpoint 501. For example, the angle between a line from the virtual viewpoint toward the selected virtual sound source and the line of sight of the virtual viewpoint is smaller than the angle between a line from the virtual viewpoint toward another virtual sound source and the line of sight of the virtual viewpoint. The sound source selection unit 106 may determine a model that appears closest to the center of the field of view of the virtual viewpoint 501 in the virtual viewpoint image, and select a virtual sound source corresponding to this model.
[0061] 5, the sound source selection unit 106 first determines virtual advertisements 507 and 508 included in the viewing area 503. Then, the sound source selection unit 106 determines virtual sound sources 504 and 505 corresponding to the determined virtual advertisements 507 and 508. Furthermore, from among the virtual sound sources 504 and 505, the sound source 504 that is closest to the line of sight 502 of the virtual viewpoint is selected.
[0062] In S406, the output unit 108 outputs the virtual viewpoint video. The process of S406 can be performed in the same manner as S405, except for the method of selecting the virtual sound source. In S406, the sound source selection unit 106 selects the virtual sound source closest to the position of the virtual viewpoint. For example, the sound source selection unit 106 can select the virtual sound source closest to the position of the virtual viewpoint based on the position information of the virtual viewpoint received from the viewpoint setting unit 105.
[0063] FIG. 6 is a diagram for explaining a method for selecting a virtual sound source close to a virtual viewpoint. In FIG. 6, a virtual viewpoint 501, virtual sound sources 504 to 506, and virtual advertisements 507 to 509 similar to those in FIG. 5 are shown. The sound source selection unit 106 determines the distance from the virtual viewpoint 501 to each of the virtual sound sources 504 to 506. For example, in FIG. 6, the distance 601 is the shortest distance between the virtual viewpoint 501 and the virtual sound source 504. The sound source selection unit 106 can select the virtual sound source 505 closest to the virtual viewpoint 501 from among the virtual sound sources 504 to 506 based on the determined distance.
[0064] In S407, the sound source selection unit 106 determines whether the arrangement information of the virtual viewpoint received from the viewpoint setting unit 105 has changed. If the sound source selection unit 106 determines that the arrangement information of the virtual viewpoint has changed, the processing from S404 onwards is executed. In this case, the sound source selection unit 106 can reselect the virtual sound source. Also, the output unit 108 can output a virtual viewpoint video based on the new virtual viewpoint.
[0065] In S408, the output unit 108 determines whether or not to end the generation of the virtual viewpoint video. For example, when ending the reproduction or distribution of the virtual viewpoint video, the output unit 108 determines to end the generation of the virtual viewpoint video. If the output unit 108 determines to end the generation of the virtual viewpoint video, the process according to FIG. 4 ends. If not, the process of S407 is executed.
[0066] As described above, according to this embodiment, it is possible to select sound data to be played together with the virtual viewpoint image, thereby improving the user experience. In particular, in one embodiment, one or more models (e.g., advertising models) are arranged in a three-dimensional space. Then, sound data corresponding to the model shown in the virtual viewpoint image among the one or more models is output. With this configuration, it is possible to play back sound data (e.g., advertising audio) corresponding to the model shown in the virtual viewpoint image (e.g., advertising model). This further improves the user experience, and in particular, the advertising effect of the virtual advertisement can be increased.
[0067] (Embodiment 2) As described above, the information processing device 2 can output a virtual viewpoint image from a virtual viewpoint in a three-dimensional space in which a 3D model based on a plurality of captured images of a subject in a real space is arranged. In this embodiment, sound data such as an audio advertisement is output so as to be played together with the virtual viewpoint image. Here, the volume of the sound data is adjusted according to the volume of the collected sound data at the time of capturing the plurality of captured images. In this way, the collected sound data collected at the time of capturing is used as the environmental sound at the time of capturing. Then, the volume of the sound data is controlled according to the volume of the environmental sound.
[0068] An image processing system according to this embodiment will be described with reference to Fig. 1. A model generation unit 101, a sound source management unit 102, a sound source placement unit 103, a viewpoint setting unit 105, a sound source selection unit 106, and an output unit 108 have the same configurations as those in the first embodiment.
[0069] In this embodiment, the sound data control unit 107 acquires collected sound data when capturing a plurality of captured images in real space. The imaging system 1 can acquire collected sound data synchronized with the captured images. The synchronization method is not particularly limited. The imaging system 1 can transmit the collected sound data to the sound data control unit 107.
[0070] The sound data control unit 107 stores the collected sound data received from the imaging system 1. The sound data control unit 107 can also calculate the volume (e.g., noise level) of the imaging environment for each sound collection time from the stored collected sound data, and store the calculated noise level. When the imaging system 1 has a plurality of sound collection units, the sound data control unit 107 may calculate the noise level based on data obtained by averaging the collected sound data obtained by each sound collection unit. On the other hand, the sound data control unit 107 may calculate the noise level based on the collected sound data obtained by each sound collection unit, and calculate the average value of the power of the noise level for each collected sound data. Depending on the imaging environment, it is also possible that the collected sound data collected by some sound collection units or imaging devices is not used for calculating the noise level. The stored noise level for each time is linked to the time of the virtual viewpoint video as environmental sound information.
[0071] Furthermore, the sound data control unit 107 adjusts the volume of the sound data according to the volume of the collected sound data. For example, when the volume of the collected sound data is low, the sound data control unit 107 can reduce the volume of the sound data. For example, the sound data control unit 107 can control the volume of the sound data corresponding to at least one virtual sound source so as not to exceed the volume of the collected sound data at the time of capturing the captured image. In this way, the sound data control unit 107 may limit the volume of the sound data so as not to exceed the volume of the collected sound data. For example, the sound data control unit 107 can reduce the volume of the sound data so that the volume of the sound data is approximately the same as the volume of the collected sound data. On the other hand, when the volume of the collected sound data is high, the sound data control unit 107 may increase the volume of the sound data. For example, the sound data control unit 107 can increase the volume of the sound data so that the volume of the sound data is approximately the same as the volume of the collected sound data.
[0072] The sound data control unit 107 may adjust the volume of the sound data for each time according to the noise level of the collected sound data calculated for each time. When the output unit 108 generates a virtual viewpoint video of a subject at a specific time, the sound data control unit 107 can obtain the noise level of the collected sound data at this time. The sound data control unit 107 can also calculate the noise level of the sound data to be reproduced. The sound data control unit 107 can then control the volume of the sound data based on a comparison between the noise level of the sound data to be reproduced and the noise level of the collected sound data. The sound data control unit 107 can output management information indicating such volume control to the output unit 108. The sound data control unit 107 can compare the volume of the sound data with the noise level of the collected sound data for each frame of the virtual viewpoint video. On the other hand, the timing of the comparison does not necessarily have to be synchronized with the frame of the virtual viewpoint video. For example, a cycle of the comparison timing may be set.
[0073] In this embodiment, the sound data to be output may be determined in advance. In this case, the sound source management unit 102, the sound source arrangement unit 103, and the sound source selection unit 106 are not necessary. On the other hand, as in the first embodiment, sound data corresponding to at least one of one or more virtual sound sources may be output. In this case, the sound data control unit 107 can control the volume of the sound data corresponding to the virtual sound source selected by the sound source selection unit 106. In such an embodiment, the sound data control unit 107 controls the playback and stopping of the sound data to be played from the virtual sound source when the virtual viewpoint video is displayed, based on the arrangement information of the virtual sound source received from the sound source selection unit 106. Furthermore, the sound data control unit 107 controls the volume of the virtual sound source according to the time information.
[0074] In this embodiment, the output unit 108 may output collected sound data in addition to sound data.
[0075] The information processing method according to this embodiment will be described below with reference to the flowchart in Fig. 7. Note that, among the processes shown in Fig. 7, steps S401 to S408 are the same as those in Fig. 4, and therefore descriptions thereof will be omitted.
[0076] In S701, the sound data control unit 107 receives collected sound data acquired by the imaging system 1. Based on the acquired collected sound data, the sound data control unit 107 calculates the volume (e.g., noise level) of the environmental sound for each collected sound time as described above. Then, the calculated noise level is stored as environmental sound information linked to the time of the virtual viewpoint video.
[0077] In S702, the sound data control unit 107 calculates the volume (e.g., noise level) of the sound data reproduced from the virtual sound source selected in S405 or S406. The noise level of the sound data may be calculated in advance. For example, when storing the sound data in the sound source management unit 102, the noise level may be calculated and saved as a parameter in the sound source management unit 102. Next, the sound data control unit 107 acquires the volume (e.g., noise level) of the environmental sound at the reproduction time corresponding to the virtual viewpoint video generated in S405 or S406. Then, the sound data control unit 107 determines whether the volume of the sound data is lower than the volume of the environmental sound. If the sound data control unit 107 determines that the volume of the sound data is lower than the volume of the environmental sound, the process of S703 is performed. If not, the process of S704 is performed.
[0078] In S703, the sound data control unit 107 adjusts the volume of the sound data so that the volume of the sound data reproduced from the virtual sound source becomes similar to the volume of the environmental sound. For example, the sound data control unit 107 can increase the volume of the sound data. In addition, the sound data control unit 107 transmits management information indicating the adjusted volume to the output unit 108.
[0079] In S704, the sound data control unit 107 adjusts the volume of the sound data so that the noise level of the sound data reproduced from the virtual sound source does not exceed the noise level of the environmental sound. For example, the sound data control unit 107 can lower the volume of the sound data. In addition, the sound data control unit 107 transmits management information indicating the adjusted volume to the output unit 108.
[0080] The method of adjusting the volume of the sound data reproduced from the virtual sound source is not particularly limited. For example, the sound data control unit 107 can increase or decrease the overall volume of the sound data. On the other hand, the sound data control unit 107 may increase or decrease the volume of a part of the sound data being reproduced. In addition, the sound data control unit 107 may control the volume of the sound data corresponding to at least one virtual sound source according to the distance between the virtual viewpoint and at least one virtual sound source. For example, in S703 and S704, the sound data control unit 107 may further adjust the volume of the sound data reproduced from the virtual sound source, taking into account the attenuation of sound due to distance, according to the distance between the virtual sound source and the virtual viewpoint.
[0081] As described above, according to this embodiment, the volume of the sound data can be controlled when the sound data is reproduced. Therefore, the sound data can be reproduced without impeding the user experience. For example, when displaying a virtual viewpoint image of a sport, the volume of sound data such as advertisements can be reduced in a scene where spectators are concentrated and there is little noise. This allows the user to maintain concentration. Also, in a scene where there is a lot of noise, the volume can be increased so that the sound data is easier to hear. In such a scene, the user experience is unlikely to be impeded even if the volume of the sound data is increased.
[0082] (Other Examples) The information processing device 2 does not need to have all the processing units shown in FIG. 2. For example, the information processing device according to an embodiment may acquire a 3D model of a subject from another device. The information processing device according to an embodiment is a server that distributes virtual viewpoint images and sound data. Such an information processing device can output a virtual viewpoint image and sound data to a client device according to information specifying a virtual viewpoint received via a network. In addition, the information processing device according to an embodiment is a user terminal that generates and reproduces virtual viewpoint images and sound data. Such an information processing device may acquire data corresponding to a virtual sound source as shown in FIG. 2 from another device. In addition, the information processing device according to an embodiment of the present invention may be configured by, for example, a plurality of information processing devices connected via a network. For example, the information processing device according to an embodiment may be configured by a plurality of servers. In addition, the information processing device according to an embodiment may be configured by a combination of a server and a client.
[0083] The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions.
[0084] The disclosure of this specification includes the following information processing device, information processing method, and program. (Item 1) A viewpoint acquisition means for acquiring information specifying a virtual viewpoint in a three-dimensional space in which one or more virtual sound sources are arranged; a management means for managing sound data corresponding to each of the one or more virtual sound sources; a generating means for generating a virtual viewpoint image of the three-dimensional space from the virtual viewpoint; an output means for outputting the virtual viewpoint image and the sound data corresponding to at least one of the one or more virtual sound sources; An information processing device comprising: (Item 2) The information processing device according to item 1, further comprising a 3D model of a subject in real space arranged in the three-dimensional space, the 3D model being based on captured images of the subject captured by each of a plurality of imaging devices. (Item 3) A selection means for selecting the at least one virtual sound source from the one or more virtual sound sources based on information specifying the virtual viewpoint, 3. The information processing device according to item 1 or 2, wherein the output means outputs the sound data corresponding to the at least one virtual sound source selected by the selection means. (Item 4) The information processing device described in item 3, characterized in that the selection means switches a selection method of the at least one virtual sound source from the one or more virtual sound sources depending on whether the position of the virtual viewpoint is within a predetermined area or not. (Item 5) The information processing device according to any one of items 1 to 4, characterized in that the at least one virtual sound source is a virtual sound source selected from the one or more virtual sound sources and included in the field of view of the virtual viewpoint. (Item 6) The information processing device described in any one of items 1 to 5, characterized in that the at least one virtual sound source is a virtual sound source selected from the one or more virtual sound sources based on at least one of the position of the virtual viewpoint and the line of sight direction of the virtual viewpoint. (Item 7) the output means outputs the sound data corresponding to one virtual sound source among a plurality of virtual sound sources arranged in a three-dimensional space; The information processing device described in any one of items 1 to 6, characterized in that a direction from the virtual viewpoint to the one virtual sound source is closer to a line of sight direction of the virtual viewpoint than a direction from the virtual viewpoint to a virtual sound source other than the one virtual sound source among the multiple virtual sound sources. (Item 8) the output means outputs the sound data corresponding to one virtual sound source among a plurality of virtual sound sources arranged in a three-dimensional space; The information processing device described in any one of items 1 to 6, characterized in that a distance from the virtual viewpoint to the one virtual sound source is shorter than a distance from the virtual viewpoint to a virtual sound source other than the one virtual sound source among the multiple virtual sound sources. (Item 9) a 3D model based on captured images of a subject captured in a real space by each of a plurality of imaging devices is further disposed in the three-dimensional space; The information processing device according to any one of items 1 to 8, further comprising a control means for adjusting the volume of the sound data corresponding to the at least one virtual sound source based on the volume of the collected sound data at the time of capturing the captured image. (Item 10) Item 10. The information processing device according to item 9, wherein the control means controls the volume of the sound data corresponding to the at least one virtual sound source so as not to exceed the volume of the sound collection data at the time of capturing the captured image. (Item 11) 11. The information processing device according to any one of items 1 to 10, further comprising a control means for controlling a volume of the sound data corresponding to the at least one virtual sound source in accordance with a distance between the virtual viewpoint and the at least one virtual sound source. (Item 12) 12. The information processing device according to any one of items 1 to 11, further comprising a control means for controlling a volume of the sound data to be output when switching the sound data to be output from the sound data corresponding to a first virtual sound source among a plurality of virtual sound sources to the sound data corresponding to a second virtual sound source among the plurality of virtual sound sources. (Item 13) 13. The information processing device according to any one of items 1 to 12, characterized in that the output means outputs the sound data corresponding to at least one virtual sound source selected by a user from the one or more virtual sound sources. (Item 14) 14. The information processing device according to any one of items 1 to 13, characterized in that at least one of the one or more virtual sound sources is associated with a model of an advertisement placed in the three-dimensional space. (Item 15) 15. The information processing device according to any one of items 1 to 14, wherein at least one of the one or more virtual sound sources is not associated with a model arranged in the three-dimensional space. (Item 16) 16. The information processing device according to any one of items 1 to 15, wherein the sound data is data that is reproduced together with the virtual viewpoint image. (Item 17) A viewpoint acquisition means for acquiring information specifying a virtual viewpoint in a three-dimensional space in which one or more models are arranged; a generating means for generating a virtual viewpoint image of the three-dimensional space from the virtual viewpoint; an output means for outputting the virtual viewpoint image and sound data corresponding to a model shown in the virtual viewpoint image among the one or more models; An information processing device comprising: (Item 18) an image output means for outputting a virtual viewpoint image from a virtual viewpoint in a three-dimensional space in which a 3D model based on a plurality of captured images of a subject in a real space is arranged; a sound output means for outputting sound data to be reproduced together with the virtual viewpoint image, the sound output means being configured to adjust a volume of the sound data in accordance with a volume of collected sound data at the time of capturing the plurality of captured images in the real space; An information processing device comprising: (Item 19) An information processing method performed by an information processing device, obtaining information identifying a virtual viewpoint in a three-dimensional space in which one or more virtual sound sources are located; sound data is managed so as to correspond to each of the one or more virtual sound sources; The information processing method further comprises: generating a virtual viewpoint image of the three-dimensional space from the virtual viewpoint; outputting the virtual viewpoint image and the sound data corresponding to at least one virtual sound source among the one or more virtual sound sources; 13. An information processing method comprising: (Item 20) 19. A program for causing a computer to function as the information processing device according to any one of items 1 to 18.
[0085] The invention is not limited to the above-described embodiments, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]
[0086] 1: imaging system, 2: information processing device, 3: display device, 101: model generation unit, 102: sound source management unit, 103: sound source placement unit, 104: user input unit, 105: viewpoint setting unit, 106: sound source selection unit, 107: sound data control unit, 108: output unit
Claims
1. A viewpoint acquisition means for acquiring information specifying a virtual viewpoint in a three-dimensional space in which one or more virtual sound sources are arranged; a management means for managing sound data corresponding to each of the one or more virtual sound sources; a generating means for generating a virtual viewpoint image of the three-dimensional space from the virtual viewpoint; an output means for outputting the virtual viewpoint image and the sound data corresponding to at least one of the one or more virtual sound sources; An information processing device comprising:
2. 2. The information processing device according to claim 1, further comprising a 3D model of a subject in real space arranged in the three-dimensional space, the 3D model being based on captured images of the subject captured by each of a plurality of imaging devices.
3. A selection unit is further provided for selecting the at least one virtual sound source from the one or more virtual sound sources based on information specifying the virtual viewpoint, 2. The information processing apparatus according to claim 1, wherein said output means outputs said sound data corresponding to said at least one virtual sound source selected by said selection means.
4. The information processing apparatus according to claim 3 , wherein the selection means switches a method of selecting the at least one virtual sound source from the one or more virtual sound sources depending on whether the position of the virtual viewpoint is within a predetermined area or not.
5. The information processing apparatus according to claim 1 , wherein the at least one virtual sound source is a virtual sound source selected from the one or more virtual sound sources and included in a field of view of the virtual viewpoint.
6. The information processing device according to claim 1 , wherein the at least one virtual sound source is a virtual sound source selected from the one or more virtual sound sources based on at least one of a position of the virtual viewpoint and a line of sight direction of the virtual viewpoint.
7. the output means outputs the sound data corresponding to one virtual sound source among a plurality of virtual sound sources arranged in a three-dimensional space; The information processing device according to claim 1, characterized in that a direction from the virtual viewpoint to the one virtual sound source is closer to a line of sight direction of the virtual viewpoint than a direction from the virtual viewpoint to a virtual sound source other than the one virtual sound source among the plurality of virtual sound sources.
8. the output means outputs the sound data corresponding to one virtual sound source among a plurality of virtual sound sources arranged in a three-dimensional space; The information processing device according to claim 1 , wherein a distance from the virtual viewpoint to the one virtual sound source is shorter than a distance from the virtual viewpoint to a virtual sound source other than the one virtual sound source among the plurality of virtual sound sources.
9. a 3D model based on captured images of a subject captured in a real space by each of a plurality of imaging devices is further disposed in the three-dimensional space; The information processing apparatus according to claim 1 , further comprising a control means for adjusting a volume of the sound data corresponding to the at least one virtual sound source based on a volume of collected sound data at the time of capturing the captured image.
10. The information processing apparatus according to claim 9 , wherein the control means controls a volume of the sound data corresponding to the at least one virtual sound source so as not to exceed a volume of the sound data collected when the captured image was captured.
11. 2. The information processing apparatus according to claim 1, further comprising a control means for controlling a volume of the sound data corresponding to the at least one virtual sound source in accordance with a distance between the virtual viewpoint and the at least one virtual sound source.
12. 2. The information processing device according to claim 1, further comprising a control means for controlling a volume of the sound data to be output when switching the sound data to be output from the sound data corresponding to a first virtual sound source among a plurality of virtual sound sources to the sound data corresponding to a second virtual sound source among the plurality of virtual sound sources.
13. 2. The information processing apparatus according to claim 1, wherein said output means outputs said sound data corresponding to at least one virtual sound source selected by a user from said one or more virtual sound sources.
14. The information processing device according to claim 1 , wherein at least one of the one or more virtual sound sources is associated with a model of an advertisement arranged in the three-dimensional space.
15. The information processing apparatus according to claim 1 , wherein at least one of the one or more virtual sound sources is not associated with a model disposed in the three-dimensional space.
16. 2. The information processing apparatus according to claim 1, wherein the sound data is data that is reproduced together with the virtual viewpoint image.
17. A viewpoint acquisition means for acquiring information specifying a virtual viewpoint in a three-dimensional space in which one or more models are arranged; a generating means for generating a virtual viewpoint image of the three-dimensional space from the virtual viewpoint; an output means for outputting the virtual viewpoint image and sound data corresponding to a model shown in the virtual viewpoint image among the one or more models; An information processing device comprising:
18. an image output means for outputting a virtual viewpoint image from a virtual viewpoint in a three-dimensional space in which a 3D model based on a plurality of captured images of a subject in a real space is arranged; a sound output means for outputting sound data to be reproduced together with the virtual viewpoint image, the sound output means being configured to adjust a volume of the sound data in accordance with a volume of collected sound data at the time of capturing the plurality of captured images in the real space; An information processing device comprising:
19. An information processing method performed by an information processing device, obtaining information identifying a virtual viewpoint in a three-dimensional space in which one or more virtual sound sources are located; sound data is managed so as to correspond to each of the one or more virtual sound sources; The information processing method further comprises: generating a virtual viewpoint image of the three-dimensional space from the virtual viewpoint; outputting the virtual viewpoint image and the sound data corresponding to at least one virtual sound source among the one or more virtual sound sources; 13. An information processing method comprising:
20. A program for causing a computer to function as the information processing device according to any one of claims 1 to 18.