Reproduction device, reproduction method, information processing device, information processing method, and program
The playback device addresses the limitations of existing technologies by allowing users to select arbitrary listening positions for object-based audio data, ensuring sound reproduction aligns with the content producer's intentions and providing a highly flexible and user-friendly experience.
Patent Information
- Application Number
- JP2024107190
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-03-28
- Filing Date
- 2024-07-03
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2037-11-10
AI Technical Summary
Existing playback technologies for object-based audio data limit users to a predetermined sound localization based on pre-prepared metadata, restricting the ability to freely select the listening position and resulting in sound reproduction that may not reflect the content producer's intentions.
A playback device that acquires content with audio data and rendering parameters pre-generated for multiple assumed listening positions, allowing users to select a preferred listening position and render the audio data accordingly, while ensuring the sound reproduction aligns with the content producer's intentions.
Enables highly flexible playback of audio data by allowing users to select arbitrary listening positions while maintaining sound quality intended by the content producer, thereby enhancing user experience and content fidelity.
Smart Images

Figure 0007687494000003 
Figure 0007687494000004 
Figure 0007687494000005
Abstract
Description
Technical Field
[0001] The present technology particularly relates to a playback device, a playback method, an information processing device, an information processing method, and a program that can realize playback of audio data with a high degree of freedom during playback while reflecting the intentions of content producers.
Background Art
[0002] Videos included in teaching videos of instrument performances, etc. are generally videos that have been pre-cut and edited by content producers. Also, the sound is a sound in which a plurality of sound sources such as explanatory voices and instrument performance sounds are appropriately mixed by content producers into 2 channels or 5.1 channels, etc. Therefore, users can only view the content from the perspective of the video and sound intended by the content producer.
[0003] By the way, in recent years, object-based audio technology has attracted attention. Object-based audio data is composed of a waveform signal of object voices and metadata indicating localization information represented by relative positions from a reference viewpoint.
[0004] Playback of object-based audio data is performed by rendering the waveform signal into a signal with a desired number of channels according to the playback-side system based on the metadata. Examples of rendering techniques include VBAP (Vector Based Amplitude Panning) (for example, Non-Patent Documents 1 and 2).
Prior Art Documents
Non-Patent Documents
[0005]
Non-Patent Document 1
[0006] Also in object-based audio data, sound localization is determined by the metadata of each object. Therefore, the user can only view and listen to the content with the sound of the rendering result according to the pre-prepared metadata, in other words, only with the sound at the determined viewpoint (assumed listening position) and the localization relative to it.
[0007] Therefore, it is conceivable to be able to arbitrarily select the assumed listening position, correct the metadata according to the assumed listening position selected by the user, and perform rendering playback with the localization corrected using the corrected metadata.
[0008] However, in this case, the reproduced sound is a sound that mechanically reflects the change in the relative positional relationship of each object, and from the perspective of the content producer, it does not necessarily become a satisfactory sound, that is, the sound that the producer wants to express.
[0009] This technology has been made in view of such a situation, and it is to realize the reproduction of audio data with a high degree of freedom during reproduction while reflecting the intention of the content producer. [Means for Solving the Problems]
[0010] The playback device according to one aspect of the present technology includes an acquisition unit that acquires the content including the audio data of each audio object and the rendering parameters of the audio data, which are generated in advance before the acquisition of the content for each of a plurality of assumed listening positions, and a rendering unit that renders the audio data based on the rendering parameters for the assumed listening position selected from among the plurality of assumed listening positions.
[0011] In one aspect of the present technology, the content including the audio data of each audio object and the rendering parameters of the audio data, which are generated in advance before the acquisition of the content for each of a plurality of assumed listening positions, is acquired, and the audio data is rendered based on the rendering parameters for the assumed listening position selected from among the plurality of assumed listening positions.
Advantages of the Invention
[0012] According to the present technology, it is possible to realize the playback of audio data with a high degree of freedom during playback while reflecting the intention of the content producer.
[0013] Note that the effects described here are not necessarily limited, and any of the effects described in the present disclosure may be applicable.
Brief Description of the Drawings
[0014]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35
Figure 36
Figure 37
Figure 38
Figure 39
Embodiments for Carrying Out the Invention
[0015] Hereinafter, embodiments for carrying out the present technology will be described. The description will be made in the following order. ·First Embodiment 1. About Content 2. Configuration and Operation of the Reproduction Device 3. Other Configuration Examples of the Reproduction Device 4. Examples of Rendering Parameters 5. Examples of Free Viewpoints 6. Configuration and Operation of the Content Generation Device 7. Modification Examples ·Second Embodiment 1. Configuration Example of the Distribution System 2. Generation Example of Rendering Parameters 3. Other Configuration Examples of the Distribution System
[0016] <<First Embodiment>> <1. About Content> FIG. 1 is a diagram showing one scene of the content reproduced by the reproduction device according to an embodiment of the present technology.
[0017] The video of the content reproduced by the reproduction device is a video whose viewpoint can be switched. The content includes video data used to display videos from a plurality of viewpoints.
[0018] Also, the audio of the content reproduced by the reproduction device is audio whose viewpoint (assumed listening position) can be switched, for example, with the position of the viewpoint of the video as the listening position. When the viewpoint is switched, the sound localization is switched.
[0019] The audio of the content is prepared as object-based audio. The audio data included in the content includes the waveform data of each audio object and the metadata for localizing the sound source of each audio object.
[0020] The content composed of such video data and audio data is provided to the reproduction device in a multiplexed form by a predetermined method such as MPEG-H.
[0021] In the following, although the content to be played back is described as a teaching video of musical instrument performance, the present technology is applicable to various contents including object-based audio data. Examples of such contents include, for example, multi-viewpoint dramas including multi-viewpoint videos and voices composed of dialogues, background sounds, sound effects, BGM, etc. as audio objects.
[0022] The horizontally long rectangular area (screen) shown in FIG. 1 is displayed on the display of the playback device. In the example of FIG. 1, in order from the left, the performance by a band consisting of a person H1 playing the bass, a person H2 playing the drums, a person H3 playing the main guitar, and a person H4 playing the side guitar is shown. The video shown in FIG. 1 is a video with the position of viewing the entire band from the front as the viewpoint.
[0023] As shown in A of FIG. 2, in the content, the performances by the bass, drums, main guitar, and side guitar, and the explanatory voice by the instructor are each audio objects, and their respective independent waveform data are recorded.
[0024] In the following, it is described that the object of the teaching is the performance of the main guitar. The performances by the side guitar, bass, and drums are for accompaniment. An example of the viewpoint of the teaching video with the performance of the main guitar as the object of the teaching is shown in B of FIG. 2.
[0025] As shown in B of FIG. 2, viewpoint #1 has the position of viewing the entire band from the front as the viewpoint (FIG. 1). Viewpoint #2 has the position of viewing only the person H3 playing the main guitar from the front as the viewpoint.
[0026] Viewpoint #3 has the position of looking up at the vicinity of the left hand of the person H3 playing the main guitar as the viewpoint. Viewpoint #4 has the position of looking up at the vicinity of the right hand of the person H3 playing the main guitar as the viewpoint. Viewpoint #5 has the position of the person H3 playing the main guitar as the viewpoint. The content records the video data used for displaying the video at each viewpoint.
[0027] Figure 3 is a diagram showing an example of rendering parameters for each audio object with respect to viewpoint #1.
[0028] In the example of Figure 3, as rendering parameters for each audio object, localization information and gain information are shown. The localization information includes information indicating the azimuth angle and information indicating the elevation angle. The azimuth angle and the elevation angle are represented with the center plane and the horizontal plane being 0° respectively.
[0029] The rendering parameters in Figure 3 indicate that the sound of the main guitar is localized at 10° to the right, the sound of the side guitar is localized at 30° to the right, the sound of the bass is localized at 30° to the left, the sound of the drums is localized at 15° to the left, the explanatory voice is localized at 0°, and the gain is set to 1.0 for all.
[0030] Figure 4 is a diagram showing an image of the localization of each audio object with respect to viewpoint #1, realized using the parameters shown in Figure 3.
[0031] The positions P1 to P5 shown enclosed by circles in Figure 4 respectively indicate the positions where the performances by the bass, the drums, the explanatory voice, the main guitar, and the side guitar are localized.
[0032] By rendering the waveform data of each audio object using the parameters shown in Figure 3, the user will hear each performance and the explanatory voice localized as shown in Figure 4. Figure 5 is a diagram showing an example of the L / R gain distribution for each audio object with respect to viewpoint #1. In this example, the speakers used for audio output are a 2 - channel speaker system.
[0033] Such rendering parameters for each audio object are also prepared for each of viewpoints #2 to #5 as shown in Figure 6.
[0034] The rendering parameters for viewpoint #2 are parameters for playing back the sound centered on the main guitar in accordance with the viewpoint video that focuses on the main guitar. Regarding the gain information of each audio object, the gains of the side guitar, bass, and drums are suppressed compared to the gains of the main guitar and the explanatory voice.
[0035] The rendering parameters for viewpoints #3 and #4 are parameters for playing back a sound that focuses more on the main guitar than in the case of viewpoint #2 in accordance with the video that focuses on the fingering of the guitar.
[0036] Viewpoint #5 is a parameter for playing back sound with localization from the performer's viewpoint in accordance with the viewpoint video in which the user can pretend to be the person H3 who is the performer of the main guitar.
[0037] In this way, for the audio data of the content played back by the playback device, the rendering parameters of each audio object are prepared for each viewpoint. The rendering parameters for each viewpoint are determined in advance by the content producer and will be transmitted or held as metadata together with the waveform data of the audio object.
[0038] <2. Configuration and Operation of the Playback Device> FIG. 7 is a block diagram showing a configuration example of the playback device.
[0039] The playback device 1 in FIG. 7 is a device used for playing back multi-viewpoint content including object-based audio data with rendering parameters prepared for each viewpoint. The playback device 1 is, for example, a personal computer and is operated by the viewer of the content.
[0040] As shown in FIG. 7, a CPU (Central Processing Unit) 11, a ROM (Read Only Memory) 12, and a RAM (Random Access Memory) 13 are interconnected by a bus 14. An input / output interface 15 is further connected to the bus 14. An input unit 16, a display 17, a speaker 18, a storage unit 19, a communication unit 20, and a drive 21 are connected to the input / output interface 15.
[0041] The input unit 16 is composed of a keyboard, a mouse, etc. The input unit 16 outputs a signal representing the content of the user's operation.
[0042] The display 17 is a display such as an LCD (Liquid Crystal Display) or an organic EL display. The display 17 displays various kinds of information such as a selection screen used for viewpoint selection and the video of the reproduced content. The display 17 may be an integral display of the playback device 1 or an external display connected to the playback device 1.
[0043] The speaker 18 outputs the audio of the reproduced content. The speaker 18 is, for example, a speaker connected to the playback device 1.
[0044] The storage unit 19 is composed of a hard disk, a non-volatile memory, etc. The storage unit 19 stores various kinds of data such as programs executed by the CPU 11 and content to be reproduced.
[0045] The communication unit 20 is composed of a network interface, etc., and communicates with external devices via a network such as the Internet. Content distributed via the network may be received by the communication unit 20 and reproduced.
[0046] The drive 21 writes data to the attached removable media 22 and reads the data recorded on the removable media 22. In the playback device 1, the content read from the removable media 22 by the drive 21 is played back as appropriate.
[0047] FIG. 8 is a block diagram showing a functional configuration example of the playback device 1.
[0048] At least a part of the configuration shown in FIG. 8 is realized by a predetermined program being executed by the CPU 11 in FIG. 7. In the playback device 1, a content acquisition unit 31, a separation unit 32, an audio playback unit 33, and a video playback unit 34 are realized.
[0049] The content acquisition unit 31 acquires content such as the above-described didactic video including video data and audio data.
[0050] When the content is provided to the playback device 1 via the removable media 22, the content acquisition unit 31 controls the drive 21 and reads and acquires the content recorded on the removable media 22. Also, when the content is provided to the playback device 1 via the network, the content acquisition unit 31 acquires the content transmitted from an external device and received by the communication unit 20. The content acquisition unit 31 outputs the acquired content to the separation unit 32.
[0051] The separation unit 32 separates the video data and audio data included in the content supplied from the content acquisition unit 31. The separation unit 32 outputs the video data of the content to the video playback unit 34 and outputs the audio data to the audio playback unit 33.
[0052] The audio playback unit 33 renders the waveform data constituting the audio data supplied from the separation unit 32 based on the metadata and outputs the audio of the content from the speaker 18.
[0053] The video playback unit 34 decodes the video data supplied from the separation unit 32 and causes the display 17 to display the video of a predetermined viewpoint of the content.
[0054] FIG. 9 is a block diagram showing a configuration example of the audio playback unit 33 in FIG. 8.
[0055] The audio playback unit 33 includes a rendering parameter selection unit 51, an object data storage unit 52, a viewpoint information display unit 53, and a rendering unit 54.
[0056] The rendering parameter selection unit 51 selects, from the object data storage unit 52, the rendering parameters for the viewpoint selected by the user according to the input selected viewpoint information, and outputs the parameters to the rendering unit 54. When a predetermined viewpoint is selected by the user from viewpoints #1 to #5, the selected viewpoint information representing the selected viewpoint is input to the rendering parameter selection unit 51.
[0057] The object data storage unit 52 stores the waveform data of each audio object, the viewpoint information, and the rendering parameters of each audio object for each of the viewpoints #1 to #5.
[0058] The rendering parameters stored in the object data storage unit 52 are read out by the rendering parameter selection unit 51, the waveform data of each audio object is read out by the rendering unit 54, and the viewpoint information is read out by the viewpoint information display unit 53. Note that the viewpoint information is information indicating that viewpoints #1 to #5 are prepared as the viewpoints of the content.
[0059] The viewpoint information display unit 53 causes the display 17 to display a viewpoint selection screen, which is a screen used for selecting the viewpoint to be played back, according to the viewpoint information read out from the object data storage unit 52. The viewpoint selection screen shows that a plurality of viewpoints #1 to #5 are prepared in advance.
[0060] On the viewpoint selection screen, a plurality of viewpoints may be indicated by icons or characters, or may be indicated by thumbnail images representing each viewpoint. The user operates the input unit 16 to select a predetermined viewpoint from among the plurality of viewpoints. The selected viewpoint information representing the viewpoint selected by the user using the viewpoint selection screen is input to the rendering parameter selection unit 51.
[0061] The rendering unit 54 reads and acquires the waveform data of each audio object from the object data storage unit 52. Further, the rendering unit 54 acquires the rendering parameters for the viewpoint selected by the user, which are supplied from the rendering parameter selection unit 51.
[0062] The rendering unit 54 renders the waveform data of each audio object according to the rendering parameters acquired from the rendering parameter selection unit 51, and outputs the audio signals of each channel to the speaker 18.
[0063] For example, assume that the speaker 18 is a 2ch speaker system with a 30° opening to the left and right, and viewpoint #1 is selected. In this case, the rendering unit 54 obtains the gain distribution shown in FIG. 5 based on the rendering parameters of FIG. 3, and performs playback by allocating the audio signals of each audio object to each of the LR channels according to the obtained gain distribution. At the speaker 18, the audio of the content is output based on the audio signal supplied from the rendering unit 54, thereby realizing playback with localization as shown in FIG. 4.
[0064] When the speaker 18 is composed of a three-dimensional speaker system such as 5.1ch or 22.2ch, the rendering unit 54 generates audio signals for each channel according to each speaker system using a rendering method such as VBAP.
[0065] Here, referring to the flowchart of FIG. 10, the audio playback process of the playback device 1 having the above configuration will be described.
[0066] The process of FIG. 10 starts when the content to be played back is selected and the viewing point is selected by the user using the viewing point selection screen. The selected viewing point information representing the viewing point selected by the user is input to the rendering parameter selection unit 51. For video playback, the video playback unit 34 performs processing for displaying the video of the viewing point selected by the user.
[0067] In step S1, the rendering parameter selection unit 51 selects the rendering parameters for the selected viewing point from the object data storage unit 52 according to the input selected viewing point information. The rendering parameter selection unit 51 outputs the selected rendering parameters to the rendering unit 54.
[0068] In step S2, the rendering unit 54 reads and acquires the waveform data of each audio object from the object data storage unit 52.
[0069] In step S3, the rendering unit 54 performs rendering of the waveform data of each audio object according to the rendering parameters supplied from the rendering parameter selection unit 51.
[0070] In step S4, the rendering unit 54 outputs the audio signals of each channel obtained by performing rendering to the speaker 18 to output the voices of each audio object.
[0071] While the content is being played back, the above processing is repeatedly performed. For example, when the viewing point is switched by the user during the playback of the content, the rendering parameters used for rendering are also switched to the rendering parameters for the newly selected viewing point.
[0072] As described above, since the rendering parameters for each audio object are prepared for each viewpoint and playback is performed using them, the user can select a preferred viewpoint from among a plurality of viewpoints and view the content with the sound that matches the selected viewpoint. The sound reproduced using the rendering parameters prepared for the viewpoint selected by the user can be said to be a highly musical sound created by the content producer.
[0073] Suppose that one rendering parameter is prepared as something common to all viewpoints, and when a viewpoint is selected, the rendering parameter is corrected to mechanically reflect the change in the positional relationship of the selected viewpoint and used for playback. In this case, the sound may be a sound unintended by the content producer, but such a situation can be prevented.
[0074] That is, through the above processing, it is possible to realize highly flexible playback of audio data in that the user can select a viewpoint while reflecting the intention of the content producer.
[0075] <3. Other Configuration Examples of the Playback Device> FIG. 11 is a block diagram showing another configuration example of the audio playback unit 33.
[0076] The audio playback unit 33 shown in FIG. 11 has the same configuration as the configuration of FIG. 9. Redundant explanations will be omitted as appropriate.
[0077] In the audio playback unit 33 having the configuration of FIG. 11, it is possible to specify an audio object for which the localization is not desired to be changed depending on the viewpoint. Among the above-described audio objects, for example, for the explanatory voice, it may be preferable to fix the localization regardless of the position of the viewpoint.
[0078] Information representing a fixed object, which is an audio object for fixing the position, is input to the rendering parameter selection unit 51 as fixed object information. The fixed object may be specified by the user or may be specified by the content producer.
[0079] The rendering parameter selection unit 51 in FIG. 11 reads the default rendering parameters from the object data storage unit 52 as the rendering parameters of the fixed object specified by the fixed object information, and outputs them to the rendering unit 54.
[0080] As the default rendering parameters, for example, the rendering parameters for viewpoint #1 may be used, or dedicated rendering parameters may be prepared.
[0081] Also, for audio objects other than the fixed object, the rendering parameter selection unit 51 reads the rendering parameters for the viewpoint selected by the user from the object data storage unit 52, and outputs them to the rendering unit 54.
[0082] The rendering unit 54 performs rendering of each audio object based on the default rendering parameters supplied from the rendering parameter selection unit 51 and the rendering parameters for the viewpoint selected by the user. The rendering unit 54 outputs the audio signals of each channel obtained by performing the rendering to the speaker 18.
[0083] For all audio objects, rendering may be performed using the default rendering parameters instead of the rendering parameters corresponding to the selected viewpoint.
[0084] FIG. 12 is a block diagram showing still another configuration example of the audio playback unit 33.
[0085] The configuration of the audio playback unit 33 shown in FIG. 12 is different from the configuration of FIG. 9 in that a switch 61 is provided between the object data storage unit 52 and the rendering unit 54.
[0086] In the audio playback unit 33 having the configuration of FIG. 12, it is possible to specify an audio object to be played or an audio object not to be played. Information representing the audio object necessary for playback is input to the switch 61 as playback object information. The object necessary for playback may be specified by the user or may be specified by the content producer.
[0087] The switch 61 in FIG. 12 outputs the waveform data of the audio object specified by the playback object information to the rendering unit 54.
[0088] The rendering unit 54 performs rendering of the waveform data of the audio object necessary for playback based on the rendering parameters for the viewpoint selected by the user, which are supplied from the rendering parameter selection unit 51. That is, the rendering unit 54 does not perform rendering on the audio objects that are not necessary for playback.
[0089] The rendering unit 54 outputs the audio signals of each channel obtained by performing rendering to the speaker 18.
[0090] As a result, the user can, for example, specify the main guitar as an audio object that is not necessary for playback, mute the sound of the main guitar as a model, and superimpose his or her own performance while watching the teaching video. In this case, only the waveform data of the audio objects other than the main guitar is supplied from the object data storage unit 52 to the rendering unit 54.
[0091] Rather than controlling the output of waveform data to the rendering unit 54, muting may be achieved by controlling the gain. In this case, the reproduction object information is input to the rendering unit 54. The rendering unit 54 sets, for example, the gain of the main guitar to 0 according to the reproduction object information, and adjusts the gains of other audio objects according to the rendering parameters supplied from the rendering parameter selection unit 51 to perform rendering.
[0092] In this way, regardless of the selected viewpoint, the localization can be fixed or only the necessary sounds can be reproduced, so that the user can reproduce the content according to their preferences.
[0093] <4. Examples of Rendering Parameters> Especially in the production of music content, the sound production of each musical instrument is performed by adjusting the sound quality with an equalizer or adding a reverberation component with reverb, in addition to adjustment by localization and gain. Such parameters used for sound production may also be added to the audio data as metadata together with the localization information and gain information and used for rendering.
[0094] Other parameters added to the localization information and gain information are also prepared for each viewpoint.
[0095] FIG. 13 is a diagram showing another example of rendering parameters.
[0096] In the example of FIG. 13, the rendering parameters include, in addition to the localization information and gain information, equalizer information, compressor information, and reverb information.
[0097] The equalizer information is composed of information on the filter type, the center frequency of the filter, sharpness, gain, and pre-gain used for acoustic adjustment by the equalizer. The compressor information is composed of information on the frequency bandwidth, threshold, ratio, gain, attack time, and release time used for acoustic adjustment by the compressor. The reverb information is composed of information on the initial reflection time, initial reflection gain, reverberation time, reverberation gain, Dumping, and Dry / Wet coefficient used for acoustic adjustment by the reverb.
[0098] It is also possible to use parameters other than the information shown in FIG. 13 as the parameters included in the rendering parameters.
[0099] FIG. 14 is a block diagram showing a configuration example of the audio playback unit 33 corresponding to the processing of the rendering parameters including the information shown in FIG. 13.
[0100] The configuration of the audio playback unit 33 shown in FIG. 14 is different from the configuration of FIG. 9 in that the rendering unit 54 is composed of an equalizer unit 71, a reverberation component addition unit 72, a compression unit 73, and a gain adjustment unit 74.
[0101] The rendering parameter selection unit 51 reads out the rendering parameters for the viewpoint selected by the user from the object data storage unit 52 according to the input selection viewpoint information, and outputs them to the rendering unit 54.
[0102] The equalizer information, compressor information, and reverb information included in the rendering parameters output from the rendering parameter selection unit 51 are respectively supplied to the equalizer unit 71, the reverberation component addition unit 72, and the compression unit 73. Also, the localization information and gain information included in the rendering parameters are supplied to the gain adjustment unit 74.
[0103] The object data storage unit 52 stores the waveform data of each audio object, the viewpoint information, and the rendering parameters for each audio object with respect to each viewpoint. The rendering parameters stored in the object data storage unit 52 include each piece of information shown in FIG. 13. The waveform data of each audio object stored in the object data storage unit 52 is supplied to the rendering unit 54.
[0104] The rendering unit 54 performs respective sound quality adjustment processes on the waveform data of each audio object according to each rendering parameter supplied from the rendering parameter selection unit 51. The rendering unit 54 performs gain adjustment on the waveform data obtained by performing the sound quality adjustment process, and outputs an audio signal to the speaker 18.
[0105] That is, the equalizer unit 71 of the rendering unit 54 performs equalizing processing based on equalizer information on the waveform data of each audio object, and outputs the waveform data obtained by the equalizing processing to the reverberation component addition unit 72.
[0106] The reverberation component addition unit 72 performs a process of adding a reverberation component based on reverb information, and outputs the waveform data with the reverberation component added to the compression unit 73.
[0107] The compression unit 73 performs compression processing based on compressor information on the waveform data supplied from the reverberation component addition unit 72, and outputs the waveform data obtained by the compression processing to the gain adjustment unit 74.
[0108] The gain adjustment unit 74 performs gain adjustment on the waveform data supplied from the compression unit 73 based on localization information and gain information, and outputs the audio signals of each channel obtained by performing the gain adjustment to the speaker 18.
[0109] By using the rendering parameters as described above, content producers can better reflect their own sound production in the rendering playback of audio objects for each viewpoint. For example, depending on the directivity of the sound, these parameters can reproduce how the timbre of the sound changes for each viewpoint. Also, it becomes possible to control the intentional mixing configuration of sounds by content producers, such that the sound of the guitar is intentionally suppressed at a certain viewpoint.
[0110] <5. Examples of Free Viewpoints> In the above, it was assumed that the selection of viewpoints can be made for a plurality of viewpoints for which rendering parameters are prepared, but it may also be possible to freely select any viewpoint. The arbitrary viewpoint here refers to a viewpoint for which rendering parameters are not prepared.
[0111] In this case, the rendering parameters for the selected arbitrary viewpoint are pseudo-generated using the rendering parameters for two viewpoints adjacent to that arbitrary viewpoint. By applying the generated rendering parameters as the rendering parameters for the arbitrary viewpoint, it becomes possible to render and play back the sound for that arbitrary viewpoint.
[0112] The number of rendering parameters used for generating pseudo-rendering parameters is not limited to two, and rendering parameters for three or more viewpoints may be used to generate the rendering parameters for an arbitrary viewpoint. Also, not limited to the rendering parameters of adjacent viewpoints, as long as they are the rendering parameters of a plurality of viewpoints in the vicinity of an arbitrary viewpoint, the rendering parameters of any viewpoint may be used to generate pseudo-rendering parameters.
[0113] Figure 15 is a diagram showing an example of rendering parameters for two viewpoints, viewpoint #6 and viewpoint #7.
[0114] In the example of FIG. 15, the rendering parameters for each audio object of the main guitar, side guitar, bass, drums, and commentary voice include localization information and gain information. As the rendering parameters shown in FIG. 15, it is also possible to use the rendering parameters including the information shown in FIG. 13.
[0115] Also, in the example of FIG. 15, the rendering parameters for viewpoint #6 indicate that the sound of the main guitar is localized 10° to the right, the sound of the side guitar is localized 30° to the right, the sound of the bass is localized 30° to the left, the sound of the drums is localized 15° to the left, and the commentary voice is localized at 0°.
[0116] On the other hand, the rendering parameters for viewpoint #7 indicate that the sound of the main guitar is localized 5° to the right, the sound of the side guitar is localized 10° to the right, the sound of the bass is localized 10° to the left, the sound of the drums is localized 8° to the left, and the commentary voice is localized at 0°.
[0117] The localization images of each audio object for each of viewpoints #6 and #7 are shown in FIGS. 16 and 17. As shown in FIG. 16, viewpoint #6 assumes a viewpoint from the front, and viewpoint #7 assumes a viewpoint from the right hand.
[0118] Here, assume that a viewpoint slightly to the right of the front, that is, the middle between viewpoints #6 and #7, is selected as arbitrary viewpoint #X. For arbitrary viewpoint #X, viewpoints #6 and #7 are adjacent viewpoints. Arbitrary viewpoint #X is a viewpoint for which no rendering parameters are prepared.
[0119] In this case, in the audio playback unit 33, pseudo-rendering parameters for arbitrary viewpoint #X are generated using the above-described rendering parameters for viewpoints #6 and #7. The pseudo-rendering parameters are generated by an interpolation process such as linear interpolation based on the rendering parameters for viewpoints #6 and #7, for example.
[0120] FIG. 18 is a diagram showing an example of pseudo-rendering parameters for arbitrary viewpoint #X.
[0121] In the example of FIG. 18, the rendering parameters for an arbitrary viewpoint #X indicate that the sound of the main guitar is localized 7.5° to the right, the sound of the side guitar is localized 20° to the right, the sound of the bass is localized 20° to the left, the sound of the drums is localized 11.5° to the left, and the explanatory voice is localized at 0°. Each value shown in FIG. 18 is the intermediate value of the values of the rendering parameters for viewpoints #6 and #7 shown in FIG. 15, and is obtained by linear interpolation processing.
[0122] The localization images of each audio object using the pseudo-rendering parameters shown in FIG. 18 are shown in FIG. 19. As shown in FIG. 19, the arbitrary viewpoint #X is a viewpoint seen slightly from the right with respect to viewpoint #6 shown in FIG. 16.
[0123] FIG. 20 is a block diagram showing a configuration example of an audio playback unit 33 having a function of generating pseudo-rendering parameters as described above.
[0124] The configuration of the audio playback unit 33 shown in FIG. 20 is different from the configuration of FIG. 9 in that a rendering parameter generation unit 81 is provided between a rendering parameter selection unit 51 and a rendering unit 54. Selection viewpoint information representing an arbitrary viewpoint #X is input to the rendering parameter selection unit 51 and the rendering parameter generation unit 81.
[0125] The rendering parameter selection unit 51 reads out, from the object data storage unit 52, the rendering parameters for a plurality of viewpoints adjacent to the arbitrarily selected viewpoint #X according to the input selection viewpoint information. The rendering parameter selection unit 51 outputs the rendering parameters for the plurality of adjacent viewpoints to the rendering parameter generation unit 81.
[0126] Based on the selected viewpoint information, the rendering parameter generation unit 81 identifies, for example, the relative positional relationship between an arbitrary viewpoint #X and a plurality of adjacent viewpoints for which rendering parameters are prepared. The rendering parameter generation unit 81 performs interpolation processing according to the identified positional relationship, and generates pseudo-rendering parameters based on the rendering parameters supplied from the rendering parameter selection unit 51. The rendering parameter generation unit 81 outputs the generated pseudo-rendering parameters to the rendering unit 54 as the rendering parameters for the arbitrary viewpoint #X.
[0127] The rendering unit 54 renders the waveform data of each audio object according to the pseudo-rendering parameters supplied from the rendering parameter generation unit 81. The rendering unit 54 outputs the audio signals of each channel obtained by rendering to the speaker 18, and outputs them as the sound of the arbitrary viewpoint #X.
[0128] Here, with reference to the flowchart of FIG. 21, the audio reproduction process by the audio reproduction unit 33 having the configuration of FIG. 20 will be described.
[0129] The process of FIG. 21 is started, for example, when an arbitrary viewpoint #X is selected by the user using the viewpoint selection screen displayed by the viewpoint information display unit 53. The selected viewpoint information representing the arbitrary viewpoint #X is input to the rendering parameter selection unit 51 and the rendering parameter generation unit 81.
[0130] In step S11, the rendering parameter selection unit 51 selects, according to the selected viewpoint information, the rendering parameters for a plurality of viewpoints adjacent to the arbitrary viewpoint #X from the object data storage unit 52. The rendering parameter selection unit 51 outputs the selected rendering parameters to the rendering parameter generation unit 81.
[0131] In step S12, the rendering parameter generation unit 81 generates pseudo-rendering parameters by performing interpolation processing according to the positional relationship between the arbitrary viewpoint #X and a plurality of adjacent viewpoints for which rendering parameters are prepared.
[0132] In step S13, the rendering unit 54 reads and acquires the waveform data of each audio object from the object data storage unit 52.
[0133] In step S14, the rendering unit 54 performs rendering of the waveform data of each audio object according to the pseudo-rendering parameters generated by the rendering parameter generation unit 81.
[0134] In step S15, the rendering unit 54 outputs the audio signals of each channel obtained by performing rendering to the speaker 18 to output the voices of each audio object.
[0135] Through the above processing, the playback device 1 can play back audio localized for an arbitrary viewpoint #X for which no rendering parameters are prepared. As a user, the user can arbitrarily select a free viewpoint and view the content.
[0136] <6. Configuration and Operation of Content Generation Device> FIG. 22 is a block diagram showing a functional configuration example of a content generation device 101 that generates content such as the above-described didactic video.
[0137] The content generation device 101 is, for example, an information processing device operated by a content producer. The content generation device 101 basically has the same hardware configuration as the playback device 1 shown in FIG. 7.
[0138] Hereinafter, the configuration shown in FIG. 7 will be appropriately cited as the configuration of the content generation device 101 and described. Each configuration shown in FIG. 22 is realized by a predetermined program being executed by the CPU 11 (FIG. 7) of the content generation device 101.
[0139] As shown in FIG. 22, the content generation device 101 includes a video generation unit 111, a metadata generation unit 112, an audio generation unit 113, a multiplexing unit 114, a recording control unit 115, and a transmission control unit 116.
[0140] The video generation unit 111 acquires a video signal input from the outside and generates video data by encoding the multi-viewpoint video signal using a predetermined encoding method. The video generation unit 111 outputs the generated video data to the multiplexing unit 114.
[0141] The metadata generation unit 112 generates rendering parameters for each audio object for each viewpoint according to the operations by the content producer. The metadata generation unit 112 outputs the generated rendering parameters to the audio generation unit 113.
[0142] In addition, the metadata generation unit 112 generates viewpoint information, which is information regarding the viewpoints of the content, according to the operations by the content producer and outputs it to the audio generation unit 113.
[0143] The audio generation unit 113 acquires an audio signal input from the outside and generates waveform data for each audio object. The audio generation unit 113 generates object-based audio data by associating the waveform data for each audio object with the rendering parameters generated by the metadata generation unit 112.
[0144] The audio generation unit 113 outputs the generated object-based audio data to the multiplexing unit 114 together with the viewpoint information.
[0145] The multiplexing unit 114 multiplexes the video data supplied from the video generation unit 111 and the audio data supplied from the audio generation unit 113 in a predetermined format such as MPEG-H to generate content. The audio data constituting the content also includes viewpoint information. The multiplexing unit 114 functions as a generation unit that generates content including object-based audio data.
[0146] When the content is provided via a recording medium, the multiplexing unit 114 outputs the generated content to the recording control unit 115, and when it is provided via a network, the multiplexing unit 114 outputs the generated content to the transmission control unit 116.
[0147] The recording control unit 115 controls the drive 21 and records the content supplied from the multiplexing unit 114 on the removable medium 22. The removable medium 22 on which the content is recorded by the recording control unit 115 is provided to the playback device 1.
[0148] The transmission control unit 116 controls the communication unit 20 and transmits the content supplied from the multiplexing unit 114 to the playback device 1.
[0149] Here, with reference to the flowchart of FIG. 23, the content generation process of the content generation device 101 having the above configuration will be described.
[0150] In step S101, the video generation unit 111 acquires the video signal input from the outside and generates video data including multi-viewpoint video signals.
[0151] In step S102, the metadata generation unit 112 generates the rendering parameters of each audio object for each viewpoint according to the operations by the content producer.
[0152] In step S103, the audio generation unit 113 acquires an audio signal input from the outside and generates waveform data for each audio object. Further, the audio generation unit 113 generates object-based audio data by associating the waveform data for each audio object with the rendering parameters generated by the metadata generation unit 112.
[0153] In step S104, the multiplexing unit 114 multiplexes the video data generated by the video generation unit 111 and the audio data generated by the audio generation unit 113 to generate content.
[0154] The content generated by the above processing is provided to the playback device 1 via a predetermined path and is played back on the playback device 1.
[0155] <7. Modification Example> Although the content played back by the playback device 1 includes video data and object-based audio data, it may be content consisting only of object-based audio data without including video data. When a predetermined listening position is selected from among the listening positions for which rendering parameters are prepared, each audio object is played back using the rendering parameters for the selected listening position.
[0156] In the above, it is assumed that the rendering parameters are determined by the content producer, but it may be possible for the user who views the content to determine them. Also, it may be possible to provide the rendering parameters for each viewpoint determined by the user himself / herself to other users via the Internet or the like.
[0157] By performing rendering playback using the rendering parameters provided in this way, the sound intended by other users will be played back. Note that the content producer may be able to limit the types and values of the parameters that can be set by the user.
[0158] Any two or more of the above-described embodiments can be used in appropriate combination. For example, as described with reference to FIG. 11, when it is possible to specify an audio object for which the localization is not to be changed, as described with reference to FIG. 12, it may be possible to specify the audio object necessary for reproduction.
[0159] <<Second Embodiment>> <1. Configuration Example of Distribution System> FIG. 24 is a diagram showing a configuration example of a distribution system that distributes content including object audio as described above, in which rendering parameters are prepared for each viewpoint.
[0160] In the distribution system of FIG. 24, the content generation device 101 managed by the content producer is installed at venue #1 where a music live is being held. On the other hand, the playback device 1 is installed in the user's home. The playback device 1 and the content generation device 101 are connected via the Internet 201.
[0161] The content generation device 101 generates content composed of video data including multi-viewpoint images and object audio including rendering parameters for each of the plurality of viewpoints. The content generated by the content generation device 101 is transmitted to, for example, a server (not shown), and provided to the playback device 1 via the server.
[0162] The playback device 1 receives the content transmitted from the content generation device 101 and plays back the video data of the viewpoint selected by the user. Further, the playback device 1 performs rendering of the object audio using the rendering parameters of the viewpoint selected by the user, and outputs the audio of the music live.
[0163] For example, the generation and transmission of content by the content generation device 101 are performed in real time following the progress of a music live performance. The user of the playback device 1 can watch the music live performance remotely almost in real time.
[0164] In the example of FIG. 24, only the playback device 1 is shown as the playback device that receives the content distribution, but in reality, many playback devices are connected to the Internet 201.
[0165] The user of the playback device 1 can freely select any viewpoint and listen to object audio. If the rendering parameters of the viewpoint selected by the user have not been transmitted from the content generation device 101, the playback device 1 generates the rendering parameters of the selected viewpoint and performs the rendering of the object audio.
[0166] In the above-described example, it is assumed that the rendering parameters are generated by linear interpolation, but in the playback device 1 of FIG. 24, they are generated using a parameter estimator configured by a neural network. The playback device 1 has a parameter estimator generated by learning using the audio data of the music live performance held at venue #1. The generation of the rendering parameters using the parameter estimator will be described later.
[0167] FIG. 25 is a block diagram showing a configuration example of the playback device 1 and the content generation device 101.
[0168] In FIG. 25, only a part of the configurations of the playback device 1 and the content generation device 101 are shown, but the playback device 1 has the configuration shown in FIG. 8. Also, the content generation device 101 has the configuration shown in FIG. 22.
[0169] The content generation device 101 includes an audio encoder 211 and a metadata encoder 212. The audio encoder 211 corresponds to the audio generation unit 113 (FIG. 22), and the metadata encoder 212 corresponds to the metadata generation unit 112.
[0170] The audio encoder 211 acquires an audio signal during a music live performance and generates waveform data for each audio object.
[0171] The metadata encoder 212 generates rendering parameters for each audio object for each viewpoint according to operations by the content producer.
[0172] The waveform data generated by the audio encoder 211 and the rendering parameters generated by the metadata encoder 212 are associated in the audio generation unit 113, thereby generating object-based audio data. The object-based audio data is multiplexed with video data in the multiplexing unit 114 and then transmitted by the transmission control unit 116 to the playback device 1.
[0173] The playback device 1 includes an audio decoder 221, a metadata decoder 222, and a playback unit 223. The audio decoder 221, the metadata decoder 222, and the playback unit 223 constitute the audio playback unit 33 (FIG. 8). In the content acquisition unit 31 of the playback device 1, the content transmitted from the content generation device 101 is acquired, and the object-based audio data and video data are separated by the separation unit 32.
[0174] Object-based audio data is input to the audio decoder 221. Also, rendering parameters for each viewpoint are input to the metadata decoder 222.
[0175] The audio decoder 221 decodes the audio data and outputs the waveform data for each audio object to the playback unit 223.
[0176] The metadata decoder 222 outputs rendering parameters corresponding to the viewpoint selected by the user to the playback unit 223.
[0177] The playback unit 223 performs rendering of the waveform data of each audio object according to the rendering parameters supplied from the metadata decoder 222, and outputs the audio corresponding to the audio signal of each channel from the speaker.
[0178] When the intervening configuration is omitted and shown, as shown in FIG. 25, the waveform data of each audio object generated by the audio encoder 211 is supplied to the audio decoder 221. Also, the rendering parameters generated by the metadata encoder 212 are supplied to the metadata decoder 222.
[0179] FIG. 26 is a block diagram showing a configuration example of the metadata decoder 222.
[0180] As shown in FIG. 26, the metadata decoder 222 is composed of a metadata acquisition unit 231, a rendering parameter selection unit 232, a rendering parameter generation unit 233, and an accumulation unit 234.
[0181] The metadata acquisition unit 231 receives and acquires the rendering parameters for each viewpoint transmitted in a form included in the audio data. The rendering parameters acquired by the metadata acquisition unit 231 are supplied to the rendering parameter selection unit 232, the rendering parameter generation unit 233, and the accumulation unit 234.
[0182] The rendering parameter selection unit 232 identifies the viewpoint selected by the user based on the input selection viewpoint information. When there is a rendering parameter for the viewpoint selected by the user among the rendering parameters supplied from the metadata acquisition unit 231, the rendering parameter selection unit 232 outputs the rendering parameter for the viewpoint selected by the user.
[0183] Also, when there is no rendering parameter for the viewpoint selected by the user, the rendering parameter selection unit 232 outputs the selection viewpoint information to the rendering parameter generation unit 233 to cause generation of the rendering parameter.
[0184] The rendering parameter generation unit 233 has a parameter estimator. The rendering parameter generation unit 233 uses the parameter estimator to generate the rendering parameter for the viewpoint selected by the user. For generating the rendering parameter, the current rendering parameter supplied from the metadata acquisition unit 231 and the past rendering parameter read from the storage unit 234 are used as inputs to the parameter estimator. The rendering parameter generation unit 233 outputs the generated rendering parameter. The rendering parameter generated by the rendering parameter generation unit 233 corresponds to the above-described pseudo rendering parameter.
[0185] Thus, the generation of the rendering parameter by the rendering parameter generation unit 233 is performed using also the rendering parameter that has been transmitted in the past from the content generation device 101. For example, when a music live is held every day at venue #1 and the distribution of the content is performed every day, the rendering parameter is transmitted from the content generation device 101 (metadata encoder 212) every day.
[0186] FIG. 27 is a diagram showing an example of the input and output of the parameter estimator included in the rendering parameter generation unit 233.
[0187] As shown by arrows A1 to A3, in addition to the information on the viewpoint selected by the user, the current (most recent) rendering parameters transmitted from the metadata encoder 212 and the past rendering parameters are input to the parameter estimator 233A.
[0188] Here, the rendering parameters include parameter information and rendering information. The parameter information is information including information indicating the type of audio object, the position information of the audio object, the viewpoint position information, and the date and time information. On the other hand, the rendering information is information regarding the characteristics of the waveform data, such as gain. Details of the information constituting the rendering parameters will be described later.
[0189] When such each information is input, as shown by arrow A4, the rendering information of the viewpoint selected by the user is output from the parameter estimator 233A.
[0190] The rendering parameter generation unit 233 appropriately performs learning of the parameter estimator 233A using the rendering parameters transmitted from the metadata encoder 212. Learning of the parameter estimator 233A is performed at a predetermined timing such as when new rendering parameters are transmitted.
[0191] The storage unit 234 stores the rendering parameters supplied from the metadata acquisition unit 231. The rendering parameters transmitted from the metadata encoder 212 are stored in the storage unit 234.
[0192] <2. Example of Generation of Rendering Parameters> Here, generation of the rendering parameters by the rendering parameter generation unit 233 will be described.
[0193] (1) Assume that there are a plurality of audio objects. The audio data of the object is defined as follows. x(n,i) where i = 0, 1, 2, …, L-1
[0194] n is the time index. Also, i represents the type of object. Here, the number of objects is L.
[0195] (2) Assume there are multiple viewpoints. The rendering information of the object corresponding to each viewpoint is defined as follows. r(i,j) where j = 0, 1, 2, …, M-1
[0196] j represents the type of viewpoint. The number of viewpoints is M.
[0197] (3) The audio data y(n,j) corresponding to each viewpoint is represented by the following formula (1).
Equation
[0198] Here, it is assumed that the rendering information r is gain (gain information). In this case, the value range of the rendering information r is 0 to 1. The audio data of each viewpoint is represented as the sum of the audio data of each object multiplied by the gain. The operation as shown in formula (1) is performed by the playback unit 223.
[0199] (4) When the viewpoint specified by the user is not any of j = 0, 1, 2, …, M-1, the rendering parameters of the viewpoint specified by the user are generated using the past rendering parameters and the current rendering parameters.
[0200] (5) The rendering information of the object corresponding to each viewpoint is defined as follows according to the type of object, the position of the object, the position of the viewpoint, and the time. r(obj_type, obj_loc_x, obj_loc_y, obj_loc_z, lis_loc_x, lis_loc_y, lis_loc_z, date_time)
[0201] obj_type is information indicating the type of the object, for example, indicating the type of musical instrument.
[0202] obj_loc_x, obj_loc_y, obj_loc_z are information indicating the position of the object in three-dimensional space.
[0203] lis_loc_x, lis_loc_y, lis_loc_z are information indicating the position of the viewing point in three-dimensional space.
[0204] date_time is information representing the date and time when the performance was conducted.
[0205] From the metadata encoder 212, such parameter information composed of obj_type, obj_loc_x, obj_loc_y, obj_loc_z, lis_loc_x, lis_loc_y, lis_loc_z, date_time is transmitted together with the rendering information r.
[0206] Specifically, it will be described below.
[0207] (6) For example, assume that each object of bass, drum, guitar, and vocal is arranged as shown in FIG. 28. FIG. 28 is a view of stage #11 in venue #1 from directly above.
[0208] (7) For venue #1, each axis of XYZ is set as shown in FIG. 29. FIG. 29 is a view of the entire venue #1 including stage #11 and the auditorium from an oblique direction. The origin O is the central position on stage #11. Viewing points 1 to 5 are set in the auditorium.
[0209] Let the coordinates of each object be represented as follows. The unit is meter. Coordinates of the base: x = -20, y = 0, z = 0 Coordinates of the drum: x = 0, y = -10, z = 0 Coordinates of the guitar: x = 20, y = 0, z = 0 Coordinates of the vocal: x = 0, y = 10, z = 0
[0210] (8) Let the coordinates of each viewpoint be represented as follows. Viewpoint 1: x = 0, y = 50, z = -1 Viewpoint 2: x = -20, y = 30, z = -1 Viewpoint 3: x = 20, y = 30, z = -1 Viewpoint 4: x = -20, y = 70, z = -1 Viewpoint 5: x = 20, y = 70, z = -1
[0211] (9) At this time, for example, the rendering information of each object at viewpoint 1 is represented as follows. Rendering information of the base : r(0, -20, 0, 0, 0, 50, -1, 2014.11.5.18.34.50) Rendering information of the drum : r(1, 0, -10, 0, 0, 50, -1, 2014.11.5.18.34.50) Rendering information of the guitar : r(2, 20, 0, 0, 0, 50, -1, 2014.11.5.18.34.50) Rendering information of the vocal : r(3, 0, 10, 0, 0, 50, -1, 2014.11.5.18.34.50)
[0212] Assume that the date and time when the music live was held is November 5, 2014, 18:34:50. Also, assume that the obj_type of each object takes the following values. Base: obj_type = 0 Drum: obj_type = 1 Guitar: obj_type = 2 Vocal: obj_type = 3
[0213] For each of viewpoints 1 to 5, the rendering parameters including the parameter information and rendering information represented as above are sent from the metadata encoder 212. The rendering parameters for each of viewpoints 1 to 5 are shown in FIGS. 30 and 31.
[0214] (10) At this time, from the above formula (1), when viewpoint 1 is selected, the audio data is represented as in the following formula (2). [Number]
[0215] However, for x(n, i), let i represent the following objects. i = 0: Bass object i = 1: Drum object i = 2: Guitar object i = 3: Vocal object
[0216] (11) Assume that viewpoint 6 indicated by the dashed line in FIG. 32 is specified by the user as the viewing position. The rendering parameters for viewpoint 6 have not been sent from the metadata encoder 212. Let the coordinates of viewpoint 6 be represented as follows. Viewpoint 6: x = 0, y = 30, z = -1
[0217] In this case, the rendering parameters for current viewpoints 1 to 5 (at 2014.11.5.18.34.50), and the rendering parameters of adjacent viewpoints sent in the past (before 2014.11.5.18.34.50) are used to generate the current rendering parameters for viewpoint 6. The past rendering parameters are read from the storage unit 234.
[0218] (12) For example, assume that the rendering parameters of viewpoints 2A and 3A shown in FIG. 33 have been sent in the past. Viewpoint 2A is located between viewpoints 2 and 4, and viewpoint 3A is located between viewpoints 3 and 5. Assume that the coordinates of viewpoints 2A and 3A are represented as follows. Viewpoint 2A: x = -20, y = 40, z = -1 Viewpoint 3A: x = 20, y = 40, z = -1
[0219] The rendering parameters of each of viewpoints 2A and 3A are shown in FIG. 34. Also in FIG. 34, the obj_type of each object takes the following values. Base: obj_type = 0 Drum: obj_type = 1 Guitar: obj_type = 2 Vocal: obj_type = 3
[0220] Thus, the position of the viewpoint from which the rendering parameters are sent from the metadata encoder 212 is not always a fixed position, but varies depending on the time. The storage unit 234 stores the rendering parameters when various positions in venue #1 are used as viewpoints.
[0221] Note that it is desirable that the configuration and position of each object of the base, drum, guitar, and vocal are the same in the current rendering parameters and the past rendering parameters used for estimation, but they may be different.
[0222] (13) Method for estimating rendering information of viewpoint 6 The following information is input to the parameter estimator 233A. · Parameter information and rendering information of viewpoints 1 to 5 (FIGS. 30 and 31) · Parameter information and rendering information of viewpoints 2A and 3A (FIG. 34) · Parameter information of viewpoint 6 (FIG. 35)
[0223] In FIG. 35, lis_loc_x, lis_loc_y, and lis_loc_z represent the position of viewpoint 6 selected by the user. Also, as data_time, 2014.11.5.18.34.50 representing the current date and time is used.
[0224] The parameter information of viewpoint 6 used as the input to parameter estimator 233A is generated by rendering parameter generation unit 233 based on, for example, the parameter information of viewpoints 1 to 5 and the position of the viewpoint selected by the user.
[0225] When such various information is input, parameter estimator 233A outputs the rendering information of each object of viewpoint 6 as shown in the rightmost column of FIG. 35. Rendering information of base (obj_type = 0) :r(0, -20, 0, 0, 0, 30, -1, 2014.11.5.18.34.50) Rendering information of drum (obj_type = 1) :r(1, 0, -10, 0, 0, 30, -1, 2014.11.5.18.34.50) Rendering information of guitar (obj_type = 2) :r(2, 20, 0, 0, 0, 30, -1, 2014.11.5.18.34.50) Rendering information of vocal (obj_type = 3) :r(3, 0, 10, 0, 0, 30, -1, 2014.11.5.18.34.50)
[0226] The rendering information output from parameter estimator 233A is supplied to playback unit 223 together with the parameter information of viewpoint 6 and is used for rendering. In this way, rendering parameter generation unit 233 generates and outputs a rendering parameter composed of the parameter information of a viewpoint for which the rendering parameter is not prepared and the rendering information estimated using parameter estimator 233A.
[0227] (14) Learning of Parameter Estimator 233A The rendering parameter generation unit 233 uses the rendering parameters transmitted from the metadata encoder 212 and stored in the storage unit 234 as learning data to perform learning of the parameter estimator 233A.
[0228] The learning of the parameter estimator 233A is performed using the rendering information r transmitted from the metadata encoder 212 as teacher data. For example, the rendering parameter generation unit 233 adjusts the coefficients so that the error (r^ - r) between the rendering information r and the output r^ of the neural network becomes small, thereby performing the learning of the parameter estimator 233A.
[0229] By performing learning using the rendering parameters transmitted from the content generation device 101, the parameter estimator 233A becomes an estimator for Hall #1, which is used to generate the rendering parameters when a predetermined position in Hall #1 is used as the viewpoint.
[0230] In the above, it is assumed that the rendering information r takes a value from 0 to 1, but it may include equalizer information, compressor information, and reverb information as described with reference to FIG. 13. That is, the rendering information r can be information representing at least any one of gain, equalizer information, compressor information, and reverb information.
[0231] Also, although the parameter estimator 233A is configured to take each information shown in FIG. 27 as input, it may be configured as a neural network that outputs the rendering information r when simply taking the parameter information of viewpoint 6 as input.
[0232] <3. Other Configuration Examples of the Distribution System> FIG. 36 is a diagram showing another configuration example of the distribution system. The same components as those described above are denoted by the same reference numerals. Redundant descriptions are omitted. The same applies to FIGS. 37 and later.
[0233] In the example of FIG. 36, there are venues #1-1 to #1-3 as the venues where music live performances are held. Content generation devices 101-1 to 101-3 are installed in venues #1-1 to #1-3 respectively. When there is no need to distinguish between content generation devices 101-1 to 101-3, they are collectively referred to as content generation device 101.
[0234] Content generation devices 101-1 to 101-3 each have the same functions as the content generation device 101 in FIG. 24. That is, content generation devices 101-1 to 101-3 each distribute the content recorded of the music live performance being held at each venue via the Internet 201.
[0235] The playback device 1 receives the content distributed by the content generation device 101 installed at the venue where the music live performance selected by the user is held, and performs playback of object-based audio data and the like as described above. The user of the playback device 1 can select a viewpoint and watch the music live performance being held at a predetermined venue.
[0236] In the example described above, the parameter estimator used for generating the rendering parameters is generated in the playback device 1, but in the example of FIG. 36, it is generated on the content generation device 101 side.
[0237] That is, content generation devices 101-1 to 101-3 each generate a parameter estimator by using the past rendering parameters as learning data as described above.
[0238] The parameter estimator generated by the content generation device 101-1 becomes a parameter estimator for venue #1-1 according to the acoustic characteristics of venue #1-1 and each viewing position. The parameter estimator generated by the content generation device 101-2 becomes a parameter estimator for venue #1-2, and the parameter estimator generated by the content generation device 101-3 becomes a parameter estimator for venue #1-3.
[0239] For example, when the playback device 1 plays the content generated by the content generation device 101-1, it acquires the parameter estimator for venue #1-1. When a user selects a viewpoint for which no rendering parameter is prepared, as described above, the playback device 1 inputs the current and past rendering parameters to the parameter estimator for venue #1-1 to generate a rendering parameter.
[0240] As described above, in the distribution system of FIG. 36, a parameter estimator for each venue is prepared on the content generation device 101 side and provided to the playback device 1. Since rendering parameters for an arbitrary viewpoint are generated using the parameter estimator for each venue, the user of the playback device 1 can view the music live of each venue by selecting an arbitrary viewpoint.
[0241] FIG. 37 is a block diagram showing a configuration example of the playback device 1 and the content generation device 101.
[0242] The configuration of the content generation device 101 shown in FIG. 37 is different from the configuration shown in FIG. 25 in that a parameter estimator learning unit 213 is provided. The content generation devices 101-1 to 101-3 in FIG. 36 each have the same configuration as the configuration of the content generation device 101 shown in FIG. 37.
[0243] The parameter estimator learning unit 213 uses the rendering parameter generated by the metadata encoder 212 as learning data to perform learning of the parameter estimator. The parameter estimator learning unit 213 transmits the parameter estimator to the playback device 1 at a predetermined timing such as before starting the distribution of the content.
[0244] The metadata acquisition unit 231 of the metadata decoder 222 of the playback device 1 receives and acquires the parameter estimator transmitted from the content generation device 101. The metadata acquisition unit 231 functions as an acquisition unit that acquires a parameter estimator corresponding to the venue.
[0245] The parameter estimator acquired by the metadata acquisition unit 231 is set in the rendering parameter generation unit 233 of the metadata decoder 222 and is appropriately used for generating rendering parameters.
[0246] FIG. 38 is a block diagram showing a configuration example of the parameter estimator learning unit 213 of FIG. 37.
[0247] The parameter estimator learning unit 213 is composed of a learning unit 251, an estimator DB 252, and an estimator providing unit 253.
[0248] The learning unit 251 uses the rendering parameters generated by the metadata encoder 212 as learning data to perform learning of the parameter estimator stored in the estimator DB 252.
[0249] The estimator providing unit 253 controls the transmission control unit 116 (FIG. 22) and transmits the parameter estimator stored in the estimator DB 252 to the playback device 1. The estimator providing unit 253 functions as a providing unit that provides the parameter estimator to the playback device 1.
[0250] In this way, it is possible to prepare a parameter estimator for each venue on the content generation device 101 side and provide it to the playback device 1 at a predetermined timing such as before the start of content playback.
[0251] In the example of FIG. 36, it is assumed that a parameter estimator for each venue is generated in the content generation device 101 installed in each venue, but it is also possible to generate it in a server connected to the Internet 201.
[0252] FIG. 39 is a block diagram showing still another configuration example of the distribution system.
[0253] The management server 301 in FIG. 39 receives the rendering parameters transmitted from the content generation devices 101-1 to 101-3 installed in Venues #1-1 to #1-3, and learns the parameter estimators for each venue. That is, the management server 301 has the parameter estimator learning unit 213 in FIG. 38.
[0254] When the management server 301 reproduces the content of a music live performance at a predetermined venue by the playback device 1, the management server 301 transmits the parameter estimator for that venue to the playback device 1. The playback device 1 appropriately uses the parameter estimator transmitted from the management server 301 to perform audio playback.
[0255] In this way, the parameter estimator may be provided via the management server 301 connected to the Internet 201. Note that the learning of the parameter estimator may be performed on the content generation device 101 side, and the generated parameter estimator may be provided to the management server 301.
[0256] The embodiments of the present technology are not limited to the above-described embodiments, and various modifications can be made without departing from the gist of the present technology.
[0257] For example, the present technology can adopt a cloud computing configuration in which one function is shared and jointly processed by a plurality of devices via a network.
[0258] In addition, each step described in the above flowchart can be executed by one device or can be shared and executed by a plurality of devices.
[0259] Furthermore, when a plurality of processes are included in one step, the plurality of processes included in that one step can be executed by one device or can be shared and executed by a plurality of devices.
[0260] The effects described in this specification are merely illustrative and not limiting, and there may be other effects.
[0261] · Regarding the program The above-described series of processes can be executed by hardware or by software. When the series of processes is executed by software, the program constituting the software is installed in a computer incorporated in dedicated hardware, or in a general-purpose personal computer or the like.
[0262] The installed program is recorded and provided on a removable medium 22 shown in FIG. 7, which includes an optical disk (such as a CD-ROM (Compact Disc-Read Only Memory), a DVD (Digital Versatile Disc), etc.) or a semiconductor memory. Further, it may be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital broadcasting. The program can be installed in advance in the ROM 12 or the storage unit 19.
[0263] Note that the program executed by the computer may be a program in which processing is performed in time series in accordance with the order described in this specification, or a program in which processing is performed in parallel or at a necessary timing such as when a call is made.
[0264] · Regarding the combination This technology can also be configured as follows. (1) An acquisition unit that acquires content including audio data of each audio object and rendering parameters of the audio data for each of a plurality of assumed listening positions; A rendering unit that performs rendering of the audio data based on the rendering parameters for a selected predetermined one of the assumed listening positions and outputs an audio signal A playback device comprising (2) The content further includes information regarding the preset assumed listening position, and further comprises a display control unit that displays a screen used for selecting the assumed listening position based on the information regarding the assumed listening position. The playback device according to (1) above. (3) The rendering parameters for each of the assumed listening positions include localization information indicating the position for localizing the audio object and gain information which is a parameter for gain adjustment of the audio data. The playback device according to (1) or (2) above. (4) The rendering unit performs the rendering of the audio data of the audio object selected as the audio object for fixing the sound source position based on rendering parameters different from the rendering parameters for the selected assumed listening position. The playback device according to any one of (1) to (3) above. (5) The rendering unit does not perform the rendering of the audio data of a predetermined audio object among the plurality of audio objects constituting the audio of the content. The playback device according to any one of (1) to (4) above. (6) It further comprises a generation unit that generates rendering parameters for each audio object with respect to the assumed listening position for which the rendering parameters are not prepared, based on the rendering parameters for the assumed listening position. The rendering unit performs the rendering of the audio data of each audio object using the rendering parameters generated by the generation unit. The playback device according to any one of (1) to (5) above. (7) The generation unit generates the rendering parameters for the assumed listening positions where the rendering parameters are not prepared, based on the rendering parameters for a plurality of the assumed listening positions in the vicinity where the rendering parameters are prepared. The playback device according to (6) above. (8) The generation unit generates the rendering parameters for the assumed listening positions where the rendering parameters are not prepared, based on the rendering parameters included in the content acquired in the past. The playback device according to (6) above. (9) The generation unit generates the rendering parameters for the assumed listening positions where the rendering parameters are not prepared, using an estimator. The playback device according to (6) above. (10) The acquisition unit acquires the estimator corresponding to the venue where the content is recorded, The generation unit generates the rendering parameters using the estimator acquired by the acquisition unit. The playback device according to (9) above. (11) The estimator is configured by learning using at least the rendering parameters included in the content acquired in the past. The playback device according to (9) or (10) above. (12) The content further includes video data used for displaying a video with the assumed listening position as the viewpoint position, The playback device further includes a video playback unit that plays back the video data and displays a video with a selected predetermined assumed listening position as the viewpoint position. The playback device according to any one of (1) to (11) above. (13) Acquire content including the audio data of each audio object and the rendering parameters of the audio data for each of a plurality of assumed listening positions. Render the audio data based on the rendering parameters for the selected predetermined assumed listening position, and output an audio signal A playback method including steps (14) Cause a computer to Obtain content including the audio data of each audio object and the rendering parameters for each of a plurality of assumed listening positions for the audio data Render the audio data based on the rendering parameters for the selected predetermined assumed listening position, and output an audio signal A program for executing a process including steps (15) A parameter generation unit that generates rendering parameters for the audio data of each audio object for each of a plurality of assumed listening positions, A content generation unit that generates content including the audio data of each audio object and the generated rendering parameters An information processing apparatus comprising (16) The parameter generation unit further generates information regarding the assumed listening position set in advance, The content generation unit generates the content further including the information regarding the assumed listening position The information processing apparatus according to (15) above (17) Further comprising a video generation unit that generates video data used for displaying a video with the assumed listening position as a viewpoint position, The content generation unit generates the content further including the video data The information processing apparatus according to (15) or (16) above (18) Further comprising a learning unit that generates an estimator used for generating the rendering parameters when a position other than the plurality of assumed listening positions where the rendering parameters are generated is used as a listening position The information processing apparatus according to any one of (15) to (17). (19) The information processing apparatus further includes a providing unit that provides the estimator to a playback apparatus that plays back the content. The information processing apparatus according to (18). (20) Generating rendering parameters of audio data of each audio object for each of a plurality of assumed listening positions, Generating content including the audio data of each audio object and the generated rendering parameters. An information processing method including steps.
Description of Signs
[0265] 1 playback apparatus, 33 audio playback unit, 51 rendering parameter selection unit, 52 object data storage unit, 53 viewpoint information display unit, 54 rendering unit
Claims
1. an acquisition unit that acquires the content including audio data of each audio object and rendering parameters for the audio data that are generated in advance for each of a plurality of assumed listening positions before acquiring the content; a rendering unit that performs rendering of the audio data based on the rendering parameters for the assumed listening position selected from the plurality of assumed listening positions; A playback device comprising:
2. The content further includes information regarding the anticipated listening position that is set in advance, The device further includes a display control unit that displays a screen used for selecting the expected listening position based on information about the expected listening position. The playback device according to claim 1 .
3. The rendering parameters for each of the assumed listening positions include localization information that indicates a position at which the audio object is to be localized, and gain information that is a parameter for adjusting the gain of the audio data.
3. The playback device according to claim 1 or 2.
4. The rendering unit renders the audio data of the audio object selected as the audio object for fixing a sound source position based on the rendering parameters different from the rendering parameters for the selected assumed listening position.
4. A reproducing apparatus according to claim 1.
5. The rendering unit does not render the audio data of a predetermined one of the plurality of audio objects that constitute the audio of the content.
5. A reproducing apparatus according to claim 1.
6. a generating unit that generates, based on the rendering parameters for the assumed listening position, the rendering parameters for each of the audio objects for which the rendering parameters are not provided, The rendering unit renders the audio data of each of the audio objects using the rendering parameters generated by the generation unit.
6. A reproducing apparatus according to claim 1.
7. The generation unit generates the rendering parameters for the assumed listening position for which the rendering parameters are not prepared, based on the rendering parameters for the plurality of assumed listening positions in the vicinity for which the rendering parameters are prepared. The playback device according to claim 6.
8. The generating unit generates the rendering parameters for the assumed listening position for which the rendering parameters are not prepared, based on the rendering parameters included in the content previously acquired. The playback device according to claim 6.
9. The generating unit generates the rendering parameters for the assumed listening position for which the rendering parameters are not prepared, using an estimator. The playback device according to claim 6.
10. The acquisition unit acquires the estimator according to a venue where the content is recorded, The generating unit generates the rendering parameters using the estimator acquired by the acquiring unit. The playback device according to claim 9.
11. The estimator is configured by learning using at least the rendering parameters included in the content previously acquired. The playback device according to claim 9 or 10.
12. the content further includes video data used to display an image with the assumed listening position as a viewpoint position; The video playback unit plays back the video data and displays a video image having the selected predetermined assumed listening position as a viewpoint. A reproducing apparatus according to any one of claims 1 to 11.
13. The playback device acquiring said content including audio data for each audio object and rendering parameters for said audio data, said rendering parameters being generated in advance for each of a plurality of intended listening positions prior to acquisition of said content; Rendering the audio data based on the rendering parameters for the assumed listening position selected from the plurality of assumed listening positions. How to play.
14. On the computer, acquiring said content including audio data for each audio object and rendering parameters for said audio data, said rendering parameters being generated in advance for each of a plurality of anticipated listening positions prior to acquiring said content; Rendering the audio data based on the rendering parameters for the assumed listening position selected from the plurality of assumed listening positions. A program for executing a process.
15. an encoding unit that encodes audio data of each of a plurality of audio objects for each of a plurality of assumed listening positions, and encodes rendering parameters of the audio data of each of the audio objects for each of the plurality of assumed listening positions; a generating unit for generating a bitstream including the encoded audio data obtained by the encoding unit and encoded rendering parameters; An information processing device comprising:
16. An information processing device, encoding audio data for each of a plurality of anticipated listening positions for each of the audio objects; encoding rendering parameters for the sound data of each of said audio objects for each of a plurality of said anticipated listening positions; A bitstream is generated that includes the encoded audio data and the encoded rendering parameters obtained by encoding. Information processing methods.
17. On the computer, encoding audio data for each of a plurality of anticipated listening positions for each of the audio objects; encoding rendering parameters for the sound data of each of said audio objects for each of a plurality of said anticipated listening positions; A bitstream is generated that includes the encoded audio data and the encoded rendering parameters obtained by encoding. A program for executing a process.
Citation Information
Patent Citations
IEC23008-3
Apparatus and method for calculating the driving coefficient of a speaker in a speaker system with respect to an audio signal associated with a virtual sound source.
JP2013510481A
Spatial audio rendering and encoding
JP2015509212A
Audio signal processor
JP2016134769A
Automatic multi-channel music mix from multiple audio stems
JP2016523001A