Reproduction method, apparatus, medium, information processing method, and apparatus

By acquiring, rendering, and generating rendering parameters of audio objects, the problem of inflexible sound positioning in object-based sound data is solved, and highly flexible sound data reproduction is achieved, reflecting the intention of the content creator and ensuring the musicality and quality of the reproduced sound.

CN114466279BActive Publication Date: 2025-10-14SONY GROUP CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210123474.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-03-28
Filing Date
2017-11-10
Publication Date
2025-10-14
Estimated Expiration
2037-11-10

AI Technical Summary

Technical Problem

In object-based sound data, users can only view the sound of a predetermined rendering result and cannot flexibly adjust the sound positioning according to the selected listening position, resulting in the reproduced sound not meeting the content creator's intention.

Method used

The acquisition unit acquires the sound data of the audio object and the rendering parameters of multiple assumed listening positions. The rendering unit renders the sound data based on the rendering parameters of the selected predetermined assumed listening position, and provides a display control unit to select the assumed listening position. The generation unit generates rendering parameters for the position for which the rendering parameters are not prepared, and the rendering parameters are learned using the estimator to achieve highly flexible sound data reproduction.

Benefits of technology

This enables flexible adjustment of sound positioning based on the user's selected listening position, reflecting the content creator's intent and ensuring the musicality and quality of the reproduced sound.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114466279B_ABST
    Figure CN114466279B_ABST
Patent Text Reader

Abstract

The present invention relates to a reproduction method, apparatus and medium, an information processing method and apparatus. The present technology relates to a reproduction apparatus, a reproduction method, an information processing apparatus, an information processing method, and a program that enable reproduction of audio data with high degrees of freedom while reflecting the intentions of a content creator. A reproduction apparatus according to one aspect of the present technology acquires content including sound data for each of audio objects and rendering parameters for each of a plurality of assumed listening positions, renders the sound data based on the rendering parameters for a predetermined assumed listening position selected, and outputs a sound signal. The present technology is applicable to an apparatus capable of reproducing object-based audio data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese patent application No. 201780071306.8, filed on May 17, 2019, entitled “Reproduction Method, Apparatus, and Medium, Information Processing Method and Apparatus.” The parent application has an international filing date of November 10, 2017, international application number PCT / JP2017 / 040617, and a priority date of November 25, 2016. Technical Field

[0002] The present technology relates to a reproduction device, a reproduction method, an information processing device, an information processing method, and a program, and in particular, to a reproduction device, a reproduction method, an information processing device, an information processing method, and a program that can achieve highly flexible reproduction of sound data while reflecting the intention of the content creator. Background Art

[0003] The images included in instructional videos of musical instrument performances, etc., are typically pre-created by the content creator through editing, etc. Furthermore, the sound is obtained by the content creator by appropriately mixing multiple sound sources, such as commentary voices and the sounds of the instrumental performance, for 2-channel, 5.1-channel, etc. Therefore, users can view content that only includes the images and sounds from the viewpoints intended by the content creator.

[0004] Incidentally, object-based audio technology has been attracting attention in recent years. Object-based sound data is composed of an audio waveform signal of an object and metadata indicating positioning information represented by a relative position from a reference viewpoint.

[0005] The object-based sound data is reproduced to render the waveform signal into a signal of the desired number of channels compatible with the system on the reproduction side based on metadata. Examples of rendering techniques include Vector-Based Amplitude Panning (VBAP) (e.g., Non-Patent Documents 1 and 2).

[0006] Reference List

[0007] Non-patent literature

[0008] Non-Patent Document 1: ISO / IEC 23008-3 Information technology - High efficiency coding and media delivery in heterogeneous environments - Part 3: 3D audio

[0009] Non-Patent Literature 2: Ville Pulkki, “Virtual Sound Source Positioning Using Vector Base Amplitude Panning,” AES Journal, Vol. 45, No. 6, pp. 456–466, 1997 Summary of the Invention

[0010] Problems to be solved by the present invention

[0011] Even with object-based sound data, sound localization is determined by metadata for each object. Therefore, users can view content with only sounds rendered according to metadata—in other words, content with only sounds at a specific viewpoint (presumed listening position) and their localization.

[0012] Therefore, it may be considered to enable selection of any assumed listening position, correct metadata according to the assumed listening position selected by the user, and perform rendering reproduction in which the localization position is modified by using the corrected metadata.

[0013] However, in this case, the reproduced sound becomes a sound that mechanically reflects changes in the relative positional relationship of each object, and does not always become a satisfactory sound, that is, a sound that the content creator desires to express.

[0014] The present technology has been made in view of such circumstances, and can realize highly flexible reproduction of sound data while reflecting the intention of the content creator.

[0015] Solutions to the Problem

[0016] According to one aspect of the present technology, a reproduction device includes an acquisition unit and a rendering unit, wherein the acquisition unit acquires content including sound data of each audio object and rendering parameters for the sound data for each of a plurality of assumed listening positions, and the rendering unit renders the sound data based on the rendering parameters for a selected predetermined assumed listening position and outputs a sound signal.

[0017] The content may further include information on a pre-set assumed listening position. In this case, a display control unit may be further provided that causes a screen for selecting the assumed listening position to be displayed based on the information on the assumed listening position.

[0018] The rendering parameters for each of the assumed listening positions may include localization information indicating a position at which an audio object is localized and gain information which is a parameter for gain adjustment of sound data.

[0019] The rendering unit may render the sound data of the audio object selected as the audio object whose sound source position is fixed regardless of the selected assumed listening position, based on rendering parameters different from the rendering parameters for the selected assumed listening position.

[0020] The rendering unit can not render the sound data of a predetermined audio object among the plurality of audio objects constituting the sound of the content.

[0021] A generation unit may also be provided that generates rendering parameters for each of the audio objects for an assumed listening position for which no rendering parameters are prepared, based on the rendering parameters for the assumed listening position. In this case, the rendering unit may render the sound data of each of the audio objects using the rendering parameters generated by the generation unit.

[0022] The generation unit may generate the rendering parameters for the assumed listening position for which the rendering parameters are not prepared, based on the rendering parameters for the plurality of nearby assumed listening positions for which the rendering parameters are prepared.

[0023] The generation unit may generate the rendering parameters for the assumed listening position for which the rendering parameters are not prepared, based on the rendering parameters included in the content acquired in the past.

[0024] The generation unit may generate rendering parameters for an assumed listening position for which no rendering parameters are prepared, by using the estimator.

[0025] The acquisition unit may acquire an estimator corresponding to a place where the content is recorded, and the generation unit may generate the rendering parameter by using the estimator acquired by the acquisition unit.

[0026] The estimator may be constructed by performing learning using at least rendering parameters included in content acquired in the past.

[0027] The content may also include video data for displaying an image from an assumed listening position as a viewpoint position. In this case, a video reproduction unit may also be provided that reproduces the video data and causes an image from a selected predetermined assumed listening position as a viewpoint position to be displayed.

[0028] According to one aspect of the present technology, content including sound data for each audio object and rendering parameters for the sound data for each of a plurality of assumed listening positions is obtained, the sound data is rendered based on the rendering parameters for a selected predetermined assumed listening position, and a sound signal is output. The sound data for an audio object selected as an audio object whose sound source position is fixed regardless of the selected assumed listening position is rendered based on rendering parameters different from the rendering parameters for the selected assumed listening position. According to one aspect of the present technology, a non-transitory computer-readable medium embodying a program thereon is disclosed, the program causing a computer to execute processing comprising the steps of: obtaining content including sound data for each audio object and rendering parameters for the sound data for each of a plurality of assumed listening positions; and rendering the sound data based on the rendering parameters for the selected predetermined assumed listening position, and outputting a sound signal. The sound data for an audio object selected as an audio object whose sound source position is fixed regardless of the selected assumed listening position is rendered based on rendering parameters different from the rendering parameters for the selected assumed listening position. According to one aspect of the present technology, an information processing device is disclosed, including: a parameter generation unit that generates rendering parameters of sound data of each of audio objects for each of a plurality of assumed listening positions; and a content generation unit that generates content including the sound data of each of the audio objects and the generated rendering parameters, wherein the rendering parameters of the sound data of the audio objects selected as the audio objects whose sound source positions are fixed regardless of the selected assumed listening positions are different from the rendering parameters for the selected assumed listening positions. According to one aspect of the present technology, an information processing method is disclosed, including the following steps: generating rendering parameters of sound data of each of audio objects for each of a plurality of assumed listening positions; and generating content including the sound data of each of the audio objects and the generated rendering parameters, wherein the rendering parameters of the sound data of the audio objects selected as the audio objects whose sound source positions are fixed regardless of the selected assumed listening positions are different from the rendering parameters for the selected assumed listening positions.

[0029] Effects of the Invention

[0030] According to the present technology, highly flexible reproduction of sound data can be achieved while reflecting the intention of the content creator.

[0031] Note that the effects described herein are not necessarily limiting and may be any of the effects described in the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 A view showing one content scene.

[0033] Figure 2 is a diagram illustrating an example of audio objects and viewpoints.

[0034] Figure 3 is a diagram showing an example of rendering parameters for viewpoint #1.

[0035] Figure 4 is a diagram showing the illustrative positioning of each audio object.

[0036] Figure 5 is a diagram showing an example of gain distribution for each audio object.

[0037] Figure 6 is a diagram showing an example of rendering parameters for viewpoints #2 to #5.

[0038] Figure 7 is a block diagram showing a configuration example of a reproducing apparatus.

[0039] Figure 8 is a block diagram showing a functional configuration example of a reproducing apparatus.

[0040] Figure 9 It shows Figure 8 A block diagram of a configuration example of an audio reproduction unit in FIG.

[0041] Figure 10 This is a flowchart for explaining the audio reproduction process of the reproduction device.

[0042] Figure 11 is a block diagram showing another configuration example of the audio reproduction unit.

[0043] Figure 12 is a block diagram showing still another configuration example of the audio reproduction unit.

[0044] Figure 13 is a diagram showing other examples of rendering parameters.

[0045] Figure 14 is a block diagram showing a configuration example of an audio reproduction unit.

[0046] Figure 15 is a diagram showing an example of rendering parameters for viewpoint #6 and viewpoint #7.

[0047] Figure 16 is a view showing illustrative positioning of each audio object for viewpoint #6.

[0048] Figure 17 is a view showing illustrative positioning of each audio object for viewpoint #7.

[0049] Figure 18is a diagram showing an example of pseudo rendering parameters for an arbitrary viewpoint #X.

[0050] Figure 19 is a diagram showing illustrative positioning of each audio object using pseudo rendering parameters.

[0051] Figure 20 is a block diagram showing a configuration example of an audio reproduction unit.

[0052] Figure 21 This is a flowchart for explaining another audio reproduction process of the reproduction device.

[0053] Figure 22 is a block diagram showing a functional configuration example of a content generating apparatus.

[0054] Figure 23 This is a flowchart for explaining content generation processing by the content generation device.

[0055] Figure 24 is a diagram showing a configuration example of a distribution system.

[0056] Figure 25 is a block diagram showing a configuration example of a reproduction device and a content generation device.

[0057] Figure 26 is a block diagram showing a configuration example of a metadata decoder.

[0058] Figure 27 is a diagram showing an example of the input and output of a parameter estimator.

[0059] Figure 28 is a view showing an arrangement example of each object.

[0060] Figure 29 This is a view of the site from an oblique direction.

[0061] Figure 30 3 is a diagram showing rendering parameters for viewpoints 1 to 5.

[0062] Figure 31 It is shown from Figure 30 Continuing graph of rendering parameters for Viewpoints 1 to 5.

[0063] Figure 32 is a diagram showing the position of the viewpoint 6 .

[0064] Figure 33 3A is a diagram showing the positions of viewpoint 2A and viewpoint 3A.

[0065] Figure 34 3A is a diagram showing rendering parameters for viewpoint 2A and viewpoint 3A.

[0066] Figure 35 is a diagram showing rendering parameters for viewpoint 6.

[0067] Figure 36 is a diagram showing another configuration example of the distribution system.

[0068] Figure 37 is a block diagram showing a configuration example of a reproduction device and a content generation device.

[0069] Figure 38 It shows Figure 37 Block diagram of a configuration example of the parameter estimator learning unit in .

[0070] Figure 39 is a block diagram showing still another configuration example of the distribution system. DETAILED DESCRIPTION

[0071] Hereinafter, a mode for carrying out the present technology will be described. The description will be given in the following order.

[0072] -First embodiment

[0073] 1. About the content

[0074] 2. Configuration and Operation of Reproduction Device

[0075] 3. Another Configuration Example of Reproduction Device

[0076] 4. Examples of rendering parameters

[0077] 5. Example of Free Viewpoint

[0078] 6. Configuration and Operation of Content Generating Devices

[0079] 7. Modify the example

[0080] - Second embodiment

[0081] 1. Configuration example of distribution system

[0082] 2. Example of generating rendering parameters

[0083] 3. Another configuration example of the distribution system

[0084] <<First embodiment>>

[0085] <1. About Content>

[0086] Figure 1 is a view showing one scene of content reproduced by the reproduction device according to one embodiment of the present technology.

[0087] The images of the content reproduced by the reproduction device are images whose viewpoints can be switched. The content includes video data for displaying images from multiple viewpoints.

[0088] Furthermore, the sound of the content reproduced by the reproduction device is a sound whose viewpoint (assumed listening position) can be switched, so that, for example, the position of the viewpoint of the image is set to the listening position. When the viewpoint is switched, the localization position of the sound is switched.

[0089] The sound of the content is prepared as object-based audio. The sound data included in the content includes waveform data of each audio object and metadata for locating the sound source of each audio object.

[0090] Content composed of such video data and sound data is provided to a reproduction device in a form multiplexed by a predetermined method such as MPEG-H.

[0091] The following description uses a video of an instructional instrument performance as the target content for reproduction. However, the present technology can be applied to various content including object-based sound data. For example, a multi-view drama including multi-view images and sounds, in which dialogue, background sounds, sound effects, background music, and the like are composed of audio objects, is considered such content.

[0092] Figure 1 The horizontally long rectangular area (screen) shown is displayed on the display of the reproduction device. Figure 1 The example in shows a performance of a band including, in order from the left, a person H1 who plays bass, a person H2 who plays drums, a person H3 who plays lead guitar, and a person H4 who plays side guitar. Figure 1 The image shown is an image whose viewpoint is at a position where the entire band is seen from the front.

[0093] like Figure 2 As shown in A, each independent waveform data is recorded in the content as each audio object of the performance of bass, drums, lead guitar and side guitar and the commentary voice of the teacher.

[0094] Given the following description, the instruction target is the performance of the lead guitar. The performances of the side guitar, bass, and drums are the accompaniment. Figure 2 B shows an example of the viewpoint of a teaching video in which the instruction target is a performance of a lead guitar.

[0095] like Figure 2 As shown in FIG. 1B , viewpoint #1 is a viewpoint at a position where the entire band is seen from the front ( Figure 1 ). Viewpoint #2 is a viewpoint at a position where only the person H3 playing the lead guitar is seen from the front.

[0096] Viewpoint #3 is a viewpoint at a position where a close-up of the left hand of the person playing the lead guitar H3 is seen. Viewpoint #4 is a viewpoint at a position where a close-up of the right hand of the person playing the lead guitar H3 is seen. Viewpoint #5 is a viewpoint at the position of the person playing the lead guitar H3. Video data for displaying images from each viewpoint is recorded in the content.

[0097] Figure 3 is a diagram showing an example of rendering parameters for each audio object for viewpoint #1.

[0098] Figure 3 The example in FIG shows positioning information and gain information as rendering parameters for each audio object. The positioning information includes information indicating the azimuth angle and information indicating the elevation angle. For the median plane and the horizontal plane, the azimuth angle and the elevation angle are respectively represented as 0°.

[0099] Figure 3 The rendering parameters shown indicate that the sound of the main guitar is positioned 10° to the right, the sound of the side guitar is positioned 30° to the right, the sound of the bass is positioned 30° to the left, the sound of the drums is positioned 15° to the left, and the commentary voice is positioned 0°, and all gains are set to 1.0.

[0100] Figure 4 is a diagram showing an illustrative positioning of each audio object for viewpoint #1 by using Figure 3 The parameters shown in are implemented.

[0101] exist Figure 4 The circled positions P1 to P5 indicate the positions where the bass performance, drum performance, commentary voice, lead guitar performance, and side guitar performance are located, respectively.

[0102] By using Figure 3 The parameters shown in the figure are used to render the waveform data of each audio object, and the user listens to the Figure 4 Each performance and commentary voice in the indicated locations. Figure 5 is a diagram showing an example of L / R gain distribution for each audio object for viewpoint # 1. In this example, the speaker for outputting sound is a 2-channel speaker system.

[0103] like Figure 6 As shown, such rendering parameters of each audio object are also prepared for each of viewpoints #2 to #5.

[0104] The rendering parameters for viewpoint #2 are used to primarily reproduce the sound of the main guitar, with the viewpoint image focused on the main guitar. Regarding the gain information for each audio object, the gains of the side guitar, bass, and drums are suppressed compared to those of the main guitar and commentary voice.

[0105] The rendering parameters for viewpoint #3 and viewpoint #4 are parameters for reproducing a sound that is more focused on the main guitar and an image that is more focused on the guitar fingering than in the case of viewpoint #2.

[0106] As for viewpoint #5, parameters are used to reproduce the sound localized at the player's viewpoint and the viewpoint image so that the user can pretend to be person H3, the lead guitar player.

[0107] Therefore, in the sound data of the content reproduced by the reproduction device, rendering parameters for each audio object are prepared for each viewpoint. The rendering parameters for each viewpoint are predetermined by the content creator and transmitted or stored as metadata along with the waveform data of the audio object.

[0108] <2. Configuration and Operation of Playback Device>

[0109] Figure 7 is a block diagram showing a configuration example of a reproducing apparatus.

[0110] Figure 7 The reproduction device 1 in the embodiment is a device for reproducing multi-viewpoint content including object-based sound data for which rendering parameters for each viewpoint are prepared. The reproduction device 1 is, for example, a personal computer and is operated by a content viewer.

[0111] like Figure 7 As shown, a central processing unit (CPU) 11, a read-only memory (ROM) 12, and a random access memory (RAM) 13 are connected to each other via a bus 14. The bus 14 is also connected to an input / output interface 15. An input unit 16, a display 17, a speaker 18, a storage unit 19, a communication unit 20, and a drive 21 are connected to the input / output interface 15.

[0112] The input unit 16 is composed of a keyboard, a mouse, etc. The input unit 16 outputs a signal indicating the content of the user's manipulation.

[0113] The display 17 is a display such as a liquid crystal display (LCD) or an organic EL display. The display 17 displays various information, such as a selection screen for selecting a viewpoint and an image of the reproduced content. The display 17 may be a display integrated with the reproduction device 1 or an external display connected to the reproduction device 1.

[0114] The speaker 18 outputs the sound of the reproduced content. The speaker 18 is, for example, a speaker connected to the reproduction device 1.

[0115] The storage unit 19 is constituted by a hard disk, a nonvolatile memory, etc. The storage unit 19 stores various data such as programs executed by the CPU 11 and reproduction target content.

[0116] The communication unit 20 is composed of a network interface and the like, and communicates with an external device via a network such as the Internet. Content distributed via the network can be received by the communication unit 20 and reproduced.

[0117] The drive 21 writes data in an attached removable medium 22 and reads data recorded on the removable medium 22. In the reproducing apparatus 1, content read out from the removable medium 22 by the drive 21 is appropriately reproduced.

[0118] Figure 8 is a block diagram showing a functional configuration example of the reproducing apparatus 1 .

[0119] By utilizing Figure 7 The CPU 11 in the system executes a predetermined program to implement Figure 8 In the reproduction apparatus 1, a content acquisition unit 31, a separation unit 32, an audio reproduction unit 33, and a video reproduction unit 34 are implemented.

[0120] The content acquisition unit 31 acquires content such as the above-mentioned teaching video including video data and sound data.

[0121] When content is provided to the reproduction apparatus 1 via the removable medium 22, the content acquisition unit 31 controls the drive 21 to read out and acquire the content recorded on the removable medium 22. Furthermore, when content is provided to the reproduction apparatus 1 via a network, the content acquisition unit 31 acquires content transmitted from an external device and received by the communication unit 20. The content acquisition unit 31 outputs the acquired content to the separation unit 32.

[0122] The separation unit 32 separates the video data and the sound data included in the content supplied from the content acquisition unit 31. The separation unit 32 outputs the video data of the content to the video reproduction unit 34, and outputs the sound data to the audio reproduction unit 33.

[0123] The audio reproduction unit 33 renders waveform data constituting the sound data supplied from the separation unit 32 based on the metadata, and causes the speaker 18 to output the sound of the content.

[0124] The video reproduction unit 34 decodes the video data supplied from the separation unit 32 and causes the display 17 to display an image of the content from a predetermined viewpoint.

[0125] Figure 9 It shows Figure 8 A block diagram of a configuration example of the audio reproduction unit 33 in FIG.

[0126] The audio reproduction unit 33 is configured from a rendering parameter selection unit 51 , an object data storage unit 52 , a viewpoint information display unit 53 , and a rendering unit 54 .

[0127] The rendering parameter selection unit 51 selects rendering parameters for the viewpoint selected by the user from the object data storage unit 52 based on the input selected viewpoint information, and outputs the rendering parameters to the rendering unit 54. When a predetermined viewpoint is selected from viewpoint #1 to viewpoint #5 by the user, the selected viewpoint information indicating the selected viewpoint is input to the rendering parameter selection unit 51.

[0128] The object data storage unit 52 stores waveform data of each audio object, viewpoint information, and rendering parameters of each audio object for each of viewpoints # 1 to # 5 .

[0129] The rendering parameters stored in the object data storage unit 52 are read out by the rendering parameter selection unit 51, and the waveform data of each audio object is read out by the rendering unit 54. The viewpoint information is read out by the viewpoint information display unit 53. Note that the viewpoint information is information indicating that viewpoints #1 to #5 are prepared as viewpoints of the content.

[0130] The viewpoint information display unit 53 causes the display 17 to display a viewpoint selection screen for selecting a reproduction viewpoint based on the viewpoint information read from the object data storage unit 52. The viewpoint selection screen shows a plurality of viewpoints prepared in advance, namely, viewpoints #1 to #5.

[0131] In the viewpoint selection screen, the presence of multiple viewpoints may be indicated by icons or characters, or by thumbnails representing each viewpoint. The user manipulates the input unit 16 to select a predetermined viewpoint from the multiple viewpoints. Selected viewpoint information representing the viewpoint selected by the user using the viewpoint selection screen is input to the rendering parameter selection unit 51.

[0132] The rendering unit 54 reads out and acquires the waveform data of each audio object from the object data storage unit 52. The rendering unit 54 also acquires the rendering parameters for the viewpoint selected by the user, which are supplied from the rendering parameter selection unit 51.

[0133] The rendering unit 54 renders the waveform data of each audio object according to the rendering parameters acquired from the rendering parameter selection unit 51 , and outputs the sound signal of each channel to the speaker 18 .

[0134] For example, the speaker 18 is a 2-channel speaker system, which is opened 30° to the left and 30° to the right, respectively, and viewpoint #1 is selected. In this case, the rendering unit 54 is based on Figure 3 The rendering parameters in Figure 5The gain distribution shown in FIG is obtained, and reproduction is performed according to the obtained gain distribution to distribute the sound signal of each audio object to each of the LR channels. At the speaker 18, the sound of the content is output based on the sound signal provided from the rendering unit 54. Thus, the sound of the content is output based on the sound signal provided from the rendering unit 54. Figure 4 Positioning reproduction shown.

[0135] In a case where the speaker 18 is constituted by a three-dimensional speaker system such as 5.1 channels or 22.2 channels, the rendering unit 54 generates a sound signal of each channel for each speaker system using a rendering technique such as VBAP.

[0136] Here, reference Figure 10 Referring to the flowchart in FIG. 1 , the audio reproduction process of the reproduction apparatus 1 having the above configuration will be described.

[0137] When the reproduction target content is selected and the user selects a viewing viewpoint using the viewpoint selection screen, the playback starts. Figure 10 The selected viewpoint information indicating the viewpoint selected by the user is input to the rendering parameter selection unit 51. Note that, as for the reproduction of the video, the process for displaying the image from the viewpoint selected by the user is performed by the video reproduction unit 34.

[0138] In step S1 , the rendering parameter selection unit 51 selects rendering parameters for the selected viewpoint from the object data storage unit 52 based on the input selected viewpoint information. The rendering parameter selection unit 51 outputs the selected rendering parameters to the rendering unit 54 .

[0139] In step S2 , the rendering unit 54 reads out and acquires the waveform data of each audio object from the object data storage unit 52 .

[0140] In step S3 , the rendering unit 54 renders the waveform data of each audio object according to the rendering parameters supplied from the rendering parameter selection unit 51 .

[0141] In step S4 , the rendering unit 54 outputs the sound signal of each channel obtained by the rendering to the speaker 18 , and causes the speaker 18 to output the sound of each audio object.

[0142] The above-described process is repeatedly executed while the content is being reproduced. For example, when the viewpoint is switched by the user during the reproduction of the content, the rendering parameters used for rendering are also switched to the rendering parameters for the newly selected viewpoint.

[0143] As described above, since rendering parameters for each audio object are prepared for each viewpoint and reproduction is performed using the rendering parameters, the user can select a desired viewpoint from multiple viewpoints and view content with sound that matches the selected viewpoint. The sound reproduced using the rendering parameters prepared for the user-selected viewpoint is highly musical and can be said to have been carefully crafted by the content creator.

[0144] Assume that a rendering parameter common to all viewpoints is prepared and a viewpoint is selected. If the rendering parameter is adjusted to mechanically reflect the change in the positional relationship of the selected viewpoint and then used for reproduction, the sound may become something that the content creator did not intend. However, this situation can be avoided.

[0145] In other words, highly flexible reproduction of sound data while reflecting the intention of the content creator can be achieved through the above-described processing in which a viewpoint can be selected.

[0146] <3. Another Configuration Example of Reproduction Device>

[0147] Figure 11 : is a block diagram showing another configuration example of the audio reproduction unit 33 .

[0148] Figure 11 The audio reproduction unit 33 shown in FIG has Figure 9 The configuration in is similar to the configuration in . Redundant description will be omitted appropriately.

[0149] In having Figure 11 In the audio reproduction unit 33 of the configuration shown, it is possible to specify an audio object whose localization position is not desired to change according to the viewpoint. Among the above audio objects, for example, in some cases, the localization position of the commentary voice is preferably fixed regardless of the viewpoint position.

[0150] Information indicating a fixed object, which is an audio object whose localized position is fixed, is input as fixed object information into the rendering parameter selection unit 51. The fixed object may be designated by a user or may be designated by a content creator.

[0151] Figure 11 The rendering parameter selection unit 51 in the image processing unit 54 reads out the default rendering parameters as the rendering parameters of the fixed object specified by the fixed object information from the object data storage unit 52 and outputs the default rendering parameters to the rendering unit 54 .

[0152] As for the default rendering parameters, for example, rendering parameters for viewpoint #1 may be used, or dedicated rendering parameters may be prepared.

[0153] Furthermore, for audio objects other than fixed objects, the rendering parameter selection unit 51 reads out rendering parameters for a viewpoint selected by the user from the object data storage unit 52 and outputs the rendering parameters to the rendering unit 54 .

[0154] The rendering unit 54 renders each audio object based on the default rendering parameters supplied from the rendering parameter selection unit 51 and the rendering parameters for the viewpoint selected by the user. The rendering unit 54 outputs the sound signal of each channel obtained by the rendering to the speaker 18.

[0155] For all audio objects, rendering may be performed using default rendering parameters instead of rendering parameters for the selected viewpoint.

[0156] Figure 12 3 is a block diagram showing still another configuration example of the audio reproduction unit 33 .

[0157] Figure 12 The configuration of the audio reproduction unit 33 shown is similar to Figure 9 The configuration in FIG. 5 is different in that a switch 61 is provided between the object data storage unit 52 and the rendering unit 54 .

[0158] In having Figure 12 In the audio reproduction unit 33 of the illustrated configuration, it is possible to specify an audio object to be reproduced or not to be reproduced. Information indicating the audio object to be reproduced is input as reproduction object information into the switch 61. The object to be reproduced can be specified by the user or by the content creator.

[0159] Figure 12 The switch 61 in outputs the waveform data of the audio object specified by the reproduction object information to the rendering unit 54.

[0160] The rendering unit 54 renders the waveform data of the audio object that needs to be reproduced based on the rendering parameters for the viewpoint selected by the user provided from the rendering parameter selection unit 51. In other words, the rendering unit 54 does not render the audio object that does not need to be reproduced.

[0161] The rendering unit 54 outputs the sound signal of each channel obtained by the rendering to the speaker 18 .

[0162] Therefore, for example, by specifying the main guitar as an audio object that does not need to be reproduced, the user can mute the sound of the main guitar as a model and superimpose his / her performance while watching the teaching video. In this case, only the waveform data of the audio objects other than the main guitar is provided from the object data storage unit 52 to the rendering unit 54.

[0163] Muting can be achieved by controlling the gain instead of controlling the output of waveform data to the rendering unit 54. In this case, the reproduction target information is input to the rendering unit 54. For example, the rendering unit 54 adjusts the gain of the lead guitar to 0 based on the reproduction target information and adjusts the gains of other audio objects based on the rendering parameters provided by the rendering parameter selection unit 51, and performs rendering.

[0164] By fixing the positioning position regardless of the selected viewpoint and reproducing only necessary sounds in this manner, the user can reproduce content according to his / her preference.

[0165] <4. Example of rendering parameters>

[0166] In particular, when creating music content, in addition to adjusting the localization position and gain, the sound reproduction of each instrument is performed using, for example, an equalizer to adjust the sound quality and add a reverberation component with reverberation. Such parameters for sound reproduction can also be added to the sound data as metadata along with the localization information and gain information and used for rendering.

[0167] Other parameters added to the positioning information and gain information are also prepared for each viewpoint.

[0168] Figure 13 is a diagram showing other examples of rendering parameters.

[0169] exist Figure 13 In the example, in addition to the positioning information and gain information, equalizer information, compressor information and reverb information are also included as rendering parameters.

[0170] The equalizer information is composed of each piece of information about the type of filter used for acoustic adjustment by the equalizer, the center frequency of the filter, the sharpness, the gain, and the pre-gain. The compressor information is composed of each piece of information about the frequency bandwidth, threshold, ratio, gain, attack time, and release time used for acoustic adjustment by the compressor. The reverb information is composed of each piece of information about the initial reflection time, initial reflection gain, reverb time, reverb gain, dumping, and dry / wet coefficient used for acoustic adjustment by reverb.

[0171] Parameters included in the render parameters can be anything except Figure 13 Information other than that shown in .

[0172] Figure 14 is shown and included Figure 13 The block diagram shows an example of a configuration of the audio reproduction unit 33 compatible with the processing of the rendering parameters of the information.

[0173] Figure 14 The configuration of the audio reproduction unit 33 shown is similar to Figure 9The configuration in is different in that the rendering unit 54 is composed of an equalizer unit 71 , a reverberation component adding unit 72 , a compression unit 73 , and a gain adjustment unit 74 .

[0174] The rendering parameter selection unit 51 reads out the rendering parameters for the viewpoint selected by the user from the object data storage unit 52 based on the input selected viewpoint information, and outputs the rendering parameters to the rendering unit 54 .

[0175] The equalizer information, reverberation information, and compressor information included in the rendering parameters output from the rendering parameter selection unit 51 are respectively supplied to the equalizer unit 71, the reverberation component adding unit 72, and the compression unit 73. Furthermore, the localization information and gain information included in the rendering parameters are supplied to the gain adjustment unit 74.

[0176] The object data storage unit 52 stores the rendering parameters of each audio object for each viewpoint together with the waveform data and viewpoint information of each audio object. The rendering parameters stored in the object data storage unit 52 include Figure 13 The waveform data of each audio object stored in the object data storage unit 52 is supplied to the rendering unit 54.

[0177] The rendering unit 54 performs individual sound quality adjustment processing on the waveform data of each audio object according to each rendering parameter supplied from the rendering parameter selection unit 51. The rendering unit 54 performs gain adjustment on the waveform data obtained by performing the sound quality adjustment processing and outputs a sound signal to the speaker 18.

[0178] In other words, the equalizer unit 71 of the rendering unit 54 performs equalization processing on the waveform data of each audio object based on the equalizer information, and outputs the waveform data obtained by the equalization processing to the reverberation component adding unit 72 .

[0179] The reverberation component adding unit 72 performs a reverberation component adding process based on the reverberation information, and outputs the waveform data to which the reverberation component is added to the compression unit 73 .

[0180] The compression unit 73 performs compression processing on the waveform data supplied from the reverberation component adding unit 72 based on the compressor information, and outputs the waveform data obtained by the compression processing to the gain adjustment unit 74 .

[0181] The gain adjustment unit 74 performs gain adjustment on the gain of the waveform data supplied from the compression unit 73 based on the localization information and the gain information, and outputs a sound signal of each channel obtained by performing the gain adjustment to the speaker 18 .

[0182] By using the rendering parameters described above, content creators can more accurately reflect their own sound reproduction in the rendered reproduction of audio objects for each viewpoint. For example, these parameters can be used to reproduce how the pitch of a sound changes for each viewpoint due to the sound's directionality. Furthermore, content creators can control the composition of the intended sound mix, such as intentionally suppressing the sound of a guitar for a certain viewpoint.

[0183] <5. Example of Free Viewpoint>

[0184] In the above, a viewpoint can be selected from a plurality of viewpoints for which rendering parameters are prepared, but an arbitrary viewpoint can also be freely selected. The arbitrary viewpoint here is a viewpoint for which rendering parameters are not prepared.

[0185] In this case, by using the rendering parameters for two viewpoints adjacent to the selected arbitrary viewpoint, pseudo rendering parameters for the selected arbitrary viewpoint are generated. By applying the generated rendering parameters as rendering parameters for the arbitrary viewpoint, rendering and reproduction of sound can be performed for the arbitrary viewpoint.

[0186] The number of viewpoints whose rendering parameters are used to generate pseudo rendering parameters is not limited to two, and rendering parameters for any viewpoint can be generated by using rendering parameters for three or more viewpoints. Furthermore, pseudo rendering parameters can be generated by using rendering parameters for any viewpoint in addition to rendering parameters for adjacent viewpoints, as long as the rendering parameters are for multiple viewpoints near the arbitrary viewpoint.

[0187] Figure 15 is a diagram showing an example of rendering parameters for two viewpoints, viewpoint #6 and viewpoint #7.

[0188] exist Figure 15 In the example of , positioning information and gain information are included as rendering parameters for each of the audio objects for the lead guitar, side guitar, bass, drums, and commentary voice. Figure 13 The information shown in the rendering parameters can also be used as Figure 15 Rendering parameters shown in .

[0189] In addition, Figure 15 In the example, the rendering parameters for viewpoint #6 indicate that the sound of the main guitar is positioned 10° to the right, the sound of the side guitar is positioned 30° to the right, the sound of the bass is positioned 30° to the left, the sound of the drums is positioned 15° to the left, and the commentary voice is positioned 0°.

[0190] Meanwhile, the rendering parameters for viewpoint #7 indicate that the sound of the main guitar is positioned 5° to the right, the sound of the side guitar is positioned 10° to the right, the sound of the bass is positioned 10° to the left, the sound of the drums is positioned 8° to the left, and the commentary voice is positioned 0°.

[0191] Figure 16 and Figure 17 Illustrative positioning of each audio object for respective viewpoints, namely viewpoint #6 and viewpoint #7, is shown. Figure 16 As shown, the viewpoint from the front is assumed to be viewpoint #6, and the viewpoint from the right hand is assumed to be viewpoint #7.

[0192] Here, we select the midpoint between viewpoints #6 and #7, or in other words, the viewpoint slightly to the right from the front, as arbitrary viewpoint #X. Viewpoints #6 and #7 are adjacent to arbitrary viewpoint #X. Arbitrary viewpoint #X is a viewpoint for which no rendering parameters have been prepared.

[0193] In this case, the audio reproduction unit 33 generates pseudo rendering parameters for arbitrary viewpoint #X by using the rendering parameters for viewpoint #6 and viewpoint #7. For example, the pseudo rendering parameters are generated based on the rendering parameters for viewpoint #6 and viewpoint #7 through interpolation processing such as linear interpolation.

[0194] Figure 18 is a diagram showing an example of pseudo rendering parameters for an arbitrary viewpoint #X.

[0195] exist Figure 18 In the example, the rendering parameters for any viewpoint #X indicate that the sound of the main guitar is positioned 7.5° to the right, the sound of the side guitar is positioned 20° to the right, the sound of the bass is positioned 20° to the left, the sound of the drums is positioned 11.5° to the left, and the commentary voice is positioned 0°. Figure 18 Each value shown is in Figure 15 The intermediate values ​​between the respective values ​​of the rendering parameters for viewpoint #6 and viewpoint #7 shown in , and are obtained by linear interpolation processing.

[0196] Figure 19 Shows the use Figure 18 The illustrative positioning of each audio object in the pseudo rendering parameters is shown in FIG. Figure 19 As shown, any viewpoint #X is from the Figure 16 Viewpoint #6 shown is a viewpoint looking slightly to the right.

[0197] Figure 20 : is a block diagram showing a configuration example of the audio reproduction unit 33 having the function of generating pseudo rendering parameters as described above.

[0198] Figure 20The configuration of the audio reproduction unit 33 shown in Figure 9 The configuration of is different in that a rendering parameter generating unit 81 is provided between the rendering parameter selecting unit 51 and the rendering unit 54. The selected viewpoint information representing the arbitrary viewpoint #X is input to the rendering parameter selecting unit 51 and the rendering parameter generating unit 81.

[0199] The rendering parameter selection unit 51 reads rendering parameters for multiple viewpoints adjacent to the arbitrary viewpoint #X selected by the user from the object data storage unit 52 based on the input selected viewpoint information. The rendering parameter selection unit 51 outputs the rendering parameters for the multiple adjacent viewpoints to the rendering parameter generation unit 81.

[0200] For example, the rendering parameter generation unit 81 identifies the relative positional relationship between the arbitrary viewpoint #X and a plurality of adjacent viewpoints for which rendering parameters have been prepared based on the selected viewpoint information. By performing interpolation processing based on the identified positional relationship, the rendering parameter generation unit 81 generates pseudo rendering parameters based on the rendering parameters provided by the rendering parameter selection unit 51. The rendering parameter generation unit 81 outputs the generated pseudo rendering parameters to the rendering unit 54 as rendering parameters for the arbitrary viewpoint #X.

[0201] The rendering unit 54 renders waveform data of each audio object according to the pseudo rendering parameters supplied from the rendering parameter generating unit 81. The rendering unit 54 outputs the sound signal of each channel obtained by the rendering to the speaker 18 and causes the speaker 18 to output the sound signal as sound for arbitrary viewpoint #X.

[0202] Here, we will refer to Figure 21 The flowchart in Figure 20 The audio reproduction processing of the audio reproduction unit 33 configured in the embodiment will be described.

[0203] For example, when the user selects an arbitrary viewpoint #X using the viewpoint selection screen displayed by the viewpoint information display unit 53, the Figure 21 The selected viewpoint information indicating the arbitrary viewpoint #X is input to the rendering parameter selection unit 51 and the rendering parameter generation unit 81.

[0204] In step S11 , the rendering parameter selection unit 51 selects rendering parameters for a plurality of viewpoints adjacent to the arbitrary viewpoint #X from the object data storage unit 52 based on the selected viewpoint information. The rendering parameter selection unit 51 outputs the selected rendering parameters to the rendering parameter generation unit 81 .

[0205] In step S12 , the rendering parameter generation unit 81 generates pseudo rendering parameters by performing interpolation processing according to the positional relationship between the arbitrary viewpoint #X and a plurality of adjacent viewpoints for which rendering parameters are prepared.

[0206] In step S13 , the rendering unit 54 reads out and acquires the waveform data of each audio object from the object data storage unit 52 .

[0207] In step S14 , the rendering unit 54 renders the waveform data of each audio object according to the pseudo rendering parameters generated by the rendering parameter generation unit 81 .

[0208] In step S15 , the rendering unit 54 outputs the sound signal of each channel obtained by the rendering to the speaker 18 , and causes the speaker 18 to output the sound of each audio object.

[0209] Through the above processing, the reproduction device 1 can reproduce audio positioned relative to an arbitrary viewpoint #X for which rendering parameters are not prepared. The user can freely select an arbitrary viewpoint and view the content.

[0210] <6. Configuration and Operation of Content Generating Device>

[0211] Figure 22 : is a block diagram showing a functional configuration example of the content generating apparatus 101 that generates content such as the instructional video as described above.

[0212] The content generation device 101 is, for example, an information processing device operated by a content creator. Figure 7 The hardware configuration of the reproduction apparatus 1 shown in FIG.

[0213] The following description will be given, with appropriate reference to Figure 7 The configuration shown in is used for the configuration of the content generating device 101. Figure 22 Each component shown is controlled by the CPU 11 ( Figure 7 ) to execute a predetermined program to achieve this.

[0214] like Figure 22 As shown, the content generating apparatus 101 is composed of a video generating unit 111 , a metadata generating unit 112 , an audio generating unit 113 , a multiplexing unit 114 , a recording control unit 115 , and a transmission control unit 116 .

[0215] The video generation unit 111 acquires an image signal input from the outside, and generates video data by encoding the multi-viewpoint image signal using a predetermined encoding method. The video generation unit 111 outputs the generated video data to the multiplexing unit 114.

[0216] The metadata generation unit 112 generates rendering parameters for each audio object for each viewpoint according to the manipulation of the content creator, and outputs the generated rendering parameters to the audio generation unit 113 .

[0217] Further, the metadata generation unit 112 generates viewpoint information, which is information about a viewpoint of the content, in accordance with the manipulation of the content creator, and outputs the viewpoint information to the audio generation unit 113.

[0218] The audio generation unit 113 acquires a sound signal inputted from the outside, and generates waveform data of each audio object. The audio generation unit 113 generates object-based sound data by associating the waveform data of each audio object with the rendering parameters generated by the metadata generation unit 112.

[0219] The audio generation unit 113 outputs the generated object-based sound data to the multiplexing unit 114 together with the viewpoint information.

[0220] The multiplexing unit 114 multiplexes the video data supplied from the video generation unit 111 and the sound data supplied from the audio generation unit 113 by a predetermined method such as MPEG-H, and generates a content. The sound data constituting the content also includes the viewpoint information. The multiplexing unit 114 functions as a generation unit that generates a content including object-based sound data.

[0221] In a case where the content is provided via a recording medium, the multiplexing unit 114 outputs the generated content to the recording control unit 115. In a case where the content is provided via a network, the multiplexing unit 114 outputs the generated content to the transmission control unit 116.

[0222] The recording control unit 115 controls the drive 21 and records the content supplied from the multiplexing unit 114 on the removable medium 22. The removable medium 22 on which the content is recorded by the recording control unit 115 is supplied to the reproducing apparatus 1.

[0223] The transmission control unit 116 controls the communication unit 20, and transmits the content supplied from the multiplexing unit 114 to the reproducing apparatus 1.

[0224] Here, the content generation processing of the content generation apparatus 101 having the above configuration will be described with reference to the flowchart in Figure 23

[0225] In step S101, the video generation unit 111 acquires an image signal inputted from the outside, and generates video data including a multi-viewpoint image signal.

[0226] In step S102, the metadata generation unit 112 generates rendering parameters of each audio object for each viewpoint in accordance with the manipulation of the content creator.

[0227] ​In step S103 , the audio generation unit 113 acquires a sound signal input from the outside and generates waveform data for each audio object. The audio generation unit 113 also generates object-based sound data by associating the waveform data of each audio object with the rendering parameters generated by the metadata generation unit 112 .

[0228] In step S104 , the multiplexing unit 114 multiplexes the video data generated by the video generating unit 111 and the sound data generated by the audio generating unit 113 , and generates content.

[0229] The content generated by the above-described processing is supplied to the reproduction device 1 via a predetermined path, and is reproduced in the reproduction device 1 .

[0230] <7. Modification Example>

[0231] The content reproduced by the reproduction device 1 includes video data and object-based sound data, but the content may be composed of object-based sound data without including video data. When a predetermined listening position is selected from among the listening positions for which rendering parameters are prepared, each audio object is reproduced using the rendering parameters for the selected listening position.

[0232] In the above, the rendering parameters are determined by the content creator, but can also be determined by the user viewing the content herself / himself. In addition, the rendering parameters for each viewpoint determined by the user himself / herself can be provided to other users via the Internet or the like.

[0233] By rendering reproduction using the rendering parameters provided in this manner, sounds intended by different users are reproduced. Note that the content creator may place restrictions on the types and values ​​of parameters that can be set by the user.

[0234] In each of the above embodiments, two or more of the embodiments may be used in combination as appropriate. Figure 11 In the case of specifying an audio object whose positioning position is not desired to be changed, as described above, reference can be made to Figure 12 The described audio object can be specified to be reproduced.

[0235] <<Second embodiment>>

[0236] <1. Example of Distribution System Configuration>

[0237] Figure 24 is a diagram showing a configuration example of a distribution system that distributes content including the target audio for which rendering parameters are prepared for each viewpoint as described above.

[0238] exist Figure 24 In the distribution system, a content generation device 101 managed by a content creator is placed at a venue #1 where a music live performance is being held. At the same time, a playback device 1 is placed at a user's home. The playback device 1 and the content generation device 101 are connected via the Internet 201.

[0239] The content generating device 101 generates content composed of video data including multi-viewpoint images and object audio including rendering parameters for multiple corresponding viewpoints. The content generated by the content generating device 101 is transmitted to a server (not shown) and provided to the reproduction device 1 via the server.

[0240] The reproduction device 1 receives content transmitted from the content generation device 101 and reproduces video data for a viewpoint selected by the user. In addition, the reproduction device 1 renders target audio by using rendering parameters for the viewpoint selected by the user and outputs the sound of a music live performance.

[0241] For example, the content generation device 101 generates content in real time as the music live performance progresses and transmits the content. The user of the reproduction device 101 can remotely watch the music live performance in substantially real time.

[0242] exist Figure 24 In the example of , only the reproducing device 1 is shown as a reproducing device that receives distributed content, but many reproducing devices are actually connected to the Internet 201.

[0243] The user of the reproduction device 1 can freely select any viewpoint and hear the target audio. Without receiving rendering parameters for the viewpoint selected by the user from the content generation device 101, the reproduction device 1 generates rendering parameters for the selected viewpoint and renders the target audio.

[0244] In the above example, the rendering parameters are generated by linear interpolation, but the rendering parameters can also be generated by using a parameter estimator composed of a neural network. Figure 24 The reproduction device 1 in the . The reproduction device 1 has a parameter estimator generated by learning using the sound data of the music live performance held at the venue # 1. The generation of rendering parameters by using the parameter estimator will be described later.

[0245] Figure 25 10 is a block diagram showing a configuration example of the reproduction device 1 and the content generation device 101 .

[0246] Figure 25 Only a part of the configuration of the reproduction device 1 and the content generation device 101 is shown, but the reproduction device 1 has Figure 8 In addition, the content generating device 101 has Figure 22 Configuration shown.

[0247] The content generating apparatus 101 has an audio encoder 211 and a metadata encoder 212. The audio encoder 211 corresponds to the audio generating unit 113 ( Figure 22 ), and the metadata encoder 212 corresponds to the metadata generation unit 112.

[0248] The audio encoder 211 acquires a sound signal during a music live performance and generates waveform data for each audio object.

[0249] The metadata encoder 212 generates rendering parameters of each audio object for each viewpoint according to manipulation by a content creator.

[0250] The audio generation unit 113 generates object-based sound data by associating the waveform data generated by the audio encoder 211 with the rendering parameters generated by the metadata encoder 212. The object-based sound data is multiplexed with the video data in the multiplexing unit 114 and then transmitted to the reproduction device 1 by the transmission control unit 116.

[0251] The reproduction device 1 has an audio decoder 221, a metadata decoder 222, and a reproduction unit 223. The audio decoder 221, the metadata decoder 222, and the reproduction unit 223 constitute an audio reproduction unit 33 ( Figure 8 ). The content transmitted from the content generating device 101 is acquired in the content acquiring unit 31 of the reproducing device 1, and the object-based sound data and video data are separated by the separating unit 32.

[0252] The object-based sound data is input to the audio decoder 221 . Furthermore, the rendering parameters for each viewpoint are input to the metadata decoder 222 .

[0253] The audio decoder 221 decodes the sound data and outputs waveform data of each audio object to the reproduction unit 223 .

[0254] The metadata decoder 222 outputs the rendering parameters for the viewpoint selected by the user to the reproduction unit 223 .

[0255] The reproduction unit 223 renders the waveform data of each audio object according to the rendering parameters supplied from the metadata decoder 222 and causes the speaker to output sound corresponding to the sound signal of each channel.

[0256] like Figure 25 As shown, with intervening constituent parts omitted and not shown, the waveform data of each audio object generated by the audio encoder 211 is supplied to the audio decoder 221. In addition, the rendering parameters generated by the metadata encoder 212 are supplied to the metadata decoder 222.

[0257] Figure 26 2 is a block diagram showing a configuration example of the metadata decoder 222 .

[0258] like Figure 26 As shown, the metadata decoder 222 is composed of a metadata acquisition unit 231 , a rendering parameter selection unit 232 , a rendering parameter generation unit 233 , and an accumulation unit 234 .

[0259] The metadata acquisition unit 231 receives and acquires the rendering parameters for each viewpoint transmitted in the form of being included in the sound data. The rendering parameters acquired by the metadata acquisition unit 231 are supplied to the rendering parameter selection unit 232 , the rendering parameter generation unit 233 , and the accumulation unit 234 .

[0260] The rendering parameter selection unit 232 identifies the viewpoint selected by the user based on the input selected viewpoint information. If there are rendering parameters for the user-selected viewpoint among the rendering parameters supplied from the metadata acquisition unit 231, the rendering parameter selection unit 232 outputs the rendering parameters for the user-selected viewpoint.

[0261] Furthermore, in the case where there are no rendering parameters for the viewpoint selected by the user, the rendering parameter selection unit 232 outputs the selected viewpoint information to the rendering parameter generation unit 233 and causes the rendering parameter generation unit 233 to generate rendering parameters.

[0262] The rendering parameter generation unit 233 includes a parameter estimator. Using the parameter estimator, the rendering parameter generation unit 233 generates rendering parameters for the viewpoint selected by the user. To generate the rendering parameters, the current rendering parameters provided by the metadata acquisition unit 231 and the past rendering parameters read from the accumulation unit 234 are used as input to the parameter estimator. The rendering parameter generation unit 233 outputs the generated rendering parameters. The rendering parameters generated by the rendering parameter generation unit 233 correspond to the aforementioned pseudo rendering parameters.

[0263] Therefore, the rendering parameters are generated by the rendering parameter generation unit 233 by also using the rendering parameters transmitted in the past from the content generation device 101. For example, in the case where a music live performance is held every day at venue #1 and its content is distributed every day, the rendering parameters are transmitted from the content generation device 101 (metadata encoder 212) every day.

[0264] Figure 27 ] is a diagram showing an example of input and output of a parameter estimator that the rendering parameter generation unit 233 has.

[0265] As indicated by arrows A1 to A3 , in addition to information on the viewpoint selected by the user, current (latest) rendering parameters and past rendering parameters sent from the metadata encoder 212 are input into the parameter estimator 233A.

[0266] Here, rendering parameters include parameter information and rendering information. Parameter information includes information indicating the type of audio object, its location, viewpoint location, and date and time. Meanwhile, rendering information includes information about waveform data characteristics, such as gain. Details of the information that constitutes rendering parameters will be described later.

[0267] In the case where each piece of such information is input, the parameter estimator 233A outputs rendering information of the viewpoint selected by the user as indicated by arrow A4.

[0268] The rendering parameter generating unit 233 appropriately performs learning of the parameter estimator 233A by using the rendering parameters transmitted from the metadata encoder 212. The learning of the parameter estimator 233A is performed at a predetermined timing, for example, when new rendering parameters are transmitted.

[0269] The accumulation unit 234 stores the rendering parameters supplied from the metadata acquisition unit 231. The accumulation unit 234 accumulates the rendering parameters transmitted from the metadata encoder 212.

[0270] <2. Example of generating rendering parameters>

[0271] Here, generation of rendering parameters by the rendering parameter generation unit 233 will be described.

[0272] (1) Assume that there are multiple audio objects.

[0273] The sound data of the object is defined as follows.

[0274] x(n,i)i=0,1,2,...,L-1

[0275] n is a time index. In addition, i represents the type of object. Here, the number of objects is L.

[0276] (2) Assume that there are multiple viewpoints.

[0277] The rendering information of an object for each viewpoint is defined as follows.

[0278] r(i,j)j=0,1,2,...,M-1

[0279] j represents the type of viewpoint, and the number of viewpoints is M.

[0280] (3) The sound data y(n, j) for each viewpoint is expressed by the following expression (1).

[0281] [Mathematical formula 1]

[0282]

[0283] Here, it is assumed that the rendering information r is a gain (gain information). In this case, the value of the rendering information r ranges from 0 to 1. The sound data for each viewpoint is expressed by multiplying the sound data of each object by the gain and adding the sound data of all objects. The calculation shown in expression (1) is performed by the reproduction unit 223.

[0284] (4) In a case where the viewpoint designated by the user is not any of the viewpoints j=0, 1, 2, ..., M-1, rendering parameters for the viewpoint designated by the user are generated by using the past rendering parameters and the current rendering parameters.

[0285] (5) Next, the rendering information of the object for each viewpoint is defined using the object type, object position, viewpoint position, and time.

[0286] r(obj_type,obj_loc_x,obj_loc_y,obj_loc_z,lis_loc_x,lis_loc_y,lis_loc_z,date_time)

[0287] obj_type is information indicating the type of an object, and indicates, for example, the type of a musical instrument.

[0288] obj_loc_x, obj_loc_y, and obj_loc_z are information indicating the position of an object in a three-dimensional space.

[0289] lis_loc_x, lis_loc_y, and lis_loc_z are information indicating the position of the viewpoint in the three-dimensional space.

[0290] date_time is information indicating the date and time when the performance is performed.

[0291] Such parameter information composed of obj_type, obj_loc_x, obj_loc_y, obj_loc_z, lis_loc_x, lis_loc_y, lis_loc_z, and date_time is transmitted from the metadata encoder 212 together with the rendering information r.

[0292] This will be described in detail below.

[0293] (6) For example, Figure 28 Arrangements for bass, drums, guitar, and vocals are shown for each object. Figure 28 This is a view of Stage #11 in Venue #1 from directly above.

[0294] (7) For location #1, if Figure 29 The settings of the XYZ axes are shown. Figure 29This is a view of the entire venue #1, including stage #11 and the auditorium, seen from an oblique direction. The origin O is the center of stage #11. Viewpoints 1 to 5 are set at the auditorium.

[0295] The coordinates of each object are expressed as follows. The unit is meters.

[0296] Bess's coordinates: x = -20, y = 0, z = 0

[0297] Coordinates of the drum: x=0, y=-10, z=0

[0298] Guitar's coordinates: x = 20, y = 0, z = 0

[0299] Vocal coordinates: x = 0, y = 10, z = 0

[0300] (8) The coordinates of each viewpoint are expressed as follows.

[0301] Viewpoint 1: x = 0, y = 50, z = -1

[0302] Viewpoint 2: x = -20, y = 30, z = -1

[0303] Viewpoint 3: x = 20, y = 30, z = -1

[0304] Viewpoint 4: x = -20, y = 70, z = -1

[0305] Viewpoint 5: x = 20, y = 70, z = -1

[0306] (9) At this time, for example, the rendering information of each object for viewpoint 1 is expressed as follows.

[0307] Rendering information of bass:

[0308] r(0,-20,0,0,0,50,-1,2014.11.5.18.34.50)

[0309] Drum rendering information:

[0310] r(1,0,-10,0,0,50,-1,2014.11.5.18.34.50)

[0311] Guitar rendering information:

[0312] r(2,20,0,0,0,50,-1,2014.11.5.18.34.50)

[0313] Vocal rendering information:

[0314] r(3,0,10,0,0,50,-1,2014.11.5.18.34.50)

[0315] The date and time of the music live performance is 18:34:50 on November 5, 2014. In addition, obj_type of each object has the following value.

[0316] Bass: obj_type=0

[0317] Drum: obj_type = 1

[0318] Guitar: obj_type=2

[0319] Vocal: obj_type=3

[0320] For each of viewpoints 1 to 5, the metadata encoder 212 transmits rendering parameters including parameter information and rendering information expressed as described above. Figure 30 and Figure 31 Rendering parameters for each viewpoint from viewpoint 1 to viewpoint 5 are shown in FIG.

[0321] (10) At this time, according to the above-mentioned expression (1), the sound data in the case where the viewpoint 1 is selected is expressed by the following expression (2).

[0322] [Mathematical formula 2]

[0323] y(n,1)=x(n,0)*r(0,-20,0,0,0,50,-1,2014.11.5.18.34.50)+x(n,1)*r(1,0,-10,0,0,50,-1,2014.11.5.18.34.50 )+x(n,2)*r(2,20,0,0,0,50,-1,2014.11.5.18.34.50)+x(n,3)*r(3,0,10,0,0,50,-1,2014.11.5.18.34.50)···(2)

[0324] However, for x(n, i), i represents the following object.

[0325] i=0: bass object

[0326] i=1: drum object

[0327] i=2: guitar object

[0328] i=3: Vocal object

[0329] (11) User specified by Figure 32 The viewpoint 6 indicated by the dotted line in FIG is used as the viewing position. The rendering parameters for the viewpoint 6 are not transmitted from the metadata encoder 212. The coordinates of the viewpoint 6 are expressed as follows.

[0330] Viewpoint 6: x = 0, y = 30, z = -1

[0331] In this case, rendering parameters for the current viewpoint 6 are generated by using current (2014.11.5.18.34.50) rendering parameters for viewpoints 1 to 5 and rendering parameters for nearby viewpoints transmitted in the past (before 2014.11.5.18.34.50). The past rendering parameters are read out from the accumulation unit 234 .

[0332] (12) For example, Figure 33 The rendering parameters for viewpoint 2A and viewpoint 3A shown are those sent in the past. Viewpoint 2A is located between viewpoint 2 and viewpoint 4, and viewpoint 3A is located between viewpoint 3 and viewpoint 5. The coordinates of viewpoint 2A and viewpoint 3A are expressed as follows.

[0333] Viewpoint 2A: x = -20, y = 40, z = -1

[0334] Viewpoint 3A: x = 20, y = 40, z = -1

[0335] exist Figure 34 The rendering parameters for each of the viewpoints 2A and 3A are shown in FIG. Figure 34 In the same way, the obj_type of each object has the following values.

[0336] Bass: obj_type=0

[0337] Drum: obj_type = 1

[0338] Guitar: obj_type=2

[0339] Vocal: obj_type=3

[0340] Therefore, the position of the viewpoint whose rendering parameters are sent from the metadata encoder 212 is not always a fixed position but a different position at that time. The accumulation unit 234 stores the rendering parameters for the viewpoints at various positions of the site #1.

[0341] Note that the configuration and position of each object in bass, drums, guitar, and vocals is expected to be the same in the current rendering parameters and the past rendering parameters used for estimation, but may be different.

[0342] (13) Estimation method for rendering information for viewpoint 6

[0343] The following information is input into the parameter estimator 233A.

[0344] -Parameter information and rendering information for viewpoints 1 to 5 ( Figure 30 and Figure 31 )

[0345] -Parameter information and rendering information for viewpoint 2A and viewpoint 3A ( Figure 34 )

[0346] -Parameter information for viewpoint 6 ( Figure 35 )

[0347] exist Figure 35 In

[0065] , lis_loc_x, lis_loc_y, and lis_loc_z represent the position of the viewpoint 6 selected by the user. Also, 2014.11.5.18.34.50 representing the current date and time is used as data_time.

[0348] The parameter information for viewpoint 6 serving as an input to the parameter estimator 233A is generated by the rendering parameter generation unit 233 based on the parameter information for viewpoints 1 to 5 and the position of the viewpoint selected by the user, for example.

[0349] When each piece of information is input, the parameter estimator 233A outputs Figure 35 The rendering information for each object at viewpoint 6 is shown in the right column of FIG.

[0350] Rendering information of bass (obj_type=0):

[0351] r(0,-20,0,0,0,30,-1,2014.11.5.18.34.50)

[0352] Drum rendering information (obj_type=1):

[0353] r(1,0,-10,0,0,30,-1,2014.11.5.18.34.50)

[0354] Guitar rendering information (obj_type=2):

[0355] r(2,20,0,0,0,30,-1,2014.11.5.18.34.50)

[0356] Vocal rendering information (obj_type=3):

[0357] r(3,0,10,0,0,30,-1,2014.11.5.18.34.50)

[0358] The rendering information output from the parameter estimator 233A is supplied to the reproduction unit 223 together with the parameter information for viewpoint 6 and used for rendering. Thus, the parameter generation unit 233 generates and outputs rendering parameters composed of the parameter information for the viewpoint for which rendering parameters are not prepared and the rendering information estimated by using the parameter estimator 233A.

[0359] (14) Learning of parameter estimator 233A

[0360] The rendering parameter generation unit 233 performs learning of the parameter estimator 233A by using the rendering parameters as learning data, the rendering parameters being transmitted from the metadata encoder 212 and accumulated in the accumulation unit 234 .

[0361] In the learning of the parameter estimator 233A, the rendering information r sent from the metadata encoder 212 is used as teaching data. For example, the rendering parameter generation unit 233 performs learning of the parameter estimator 233A by adjusting the coefficients so that the error (r^-r) between the rendering information r and the output r^ of the neural network becomes smaller.

[0362] By performing learning using the rendering parameters transmitted from the content generating apparatus 101 , the parameter estimator 233A becomes an estimator of the location # 1 that is used to generate rendering parameters in a case where a predetermined position in the location # 1 is set as a viewpoint.

[0363] In the above, the rendering information r is a gain having a value from 0 to 1, but may include Figure 13 In other words, the rendering information r may be information indicating at least any one of gain, equalizer information, compressor information, or reverb information.

[0364] In addition, the parameter estimator 233A input Figure 27 , but can be simply configured as a neural network that outputs rendering information r when parameter information for viewpoint 6 is input.

[0365] <3. Another Configuration Example of Distribution System>

[0366] Figure 36 is a diagram showing another configuration example of a distribution system. The same components as those described above are denoted by the same reference numerals. Redundant descriptions will be omitted. This also applies to Figure 37 and the accompanying drawings that follow.

[0367] exist Figure 36In the example shown in FIG. 1 , there are venues #1-1 to #1-3 where live music performances are being held. Content generating devices 101-1 to 101-3 are placed at venues #1-1 to #1-3, respectively. When it is not necessary to distinguish between content generating devices 101-1 to 101-3, they are collectively referred to as content generating device 101.

[0368] Each of the content generating devices 101-1 to 101-3 has Figure 24 In other words, the content generating apparatuses 101-1 to 101-3 distribute content including music live performances held at various venues via the Internet 201.

[0369] The playback device 1 receives content distributed by the content generation device 101 placed at a venue where a music live performance selected by the user is being held, and reproduces the object-based sound data described above. The user of the playback device 1 can select a viewpoint and watch the music live performance being held at the predetermined venue.

[0370] In the aforementioned example, a parameter estimator for generating rendering parameters is generated in the reproduction device 1, but in Figure 36 In the example of , a parameter estimator for generating rendering parameters is generated on the content generating device 101 side.

[0371] In other words, the content generating devices 101 - 1 to 101 - 3 individually generate parameter estimators by using past rendering parameters as learning data or the like as mentioned previously.

[0372] The parameter estimator generated by the content generation device 101-1 is a parameter estimator for the location #1-1 that is compatible with the acoustic characteristics of the location #1-1 and each viewing position. The parameter estimator generated by the content generation device 101-2 is a parameter estimator for the location #1-2, and the parameter estimator generated by the content generation device 101-3 is a parameter estimator for the location #1-3.

[0373] For example, when reproducing content generated by the content generation device 101-1, the reproduction device 1 obtains a parameter estimator for the location #1-1. When the user selects a viewpoint for which rendering parameters are not prepared, the reproduction device 1 inputs the current rendering parameters and past rendering parameters into the parameter estimator for the location #1-1 and generates rendering parameters as described above.

[0374] Therefore, in Figure 36In the distribution system of FIG. 1 , a parameter estimator for each venue is prepared on the content generation device 101 side and provided to the reproduction device 1. Since rendering parameters for an arbitrary viewpoint are generated by using the parameter estimator for each venue, the user of the reproduction device 1 can select an arbitrary viewpoint and watch a live music performance at each venue.

[0375] Figure 37 10 is a block diagram showing a configuration example of the reproduction device 1 and the content generation device 101 .

[0376] exist Figure 37 The configuration of the content generating apparatus 101 shown is similar to Figure 25 The configuration shown is different in that a parameter estimator learning unit 213 is provided. Figure 36 Each of the content generating devices 101-1 to 101-3 shown has Figure 37 The configuration of the content generating apparatus 101 shown is the same configuration.

[0377] The parameter estimator learning unit 213 performs learning of the parameter estimator by using, as learning data, the rendering parameters generated by the metadata encoder 212. The parameter estimator learning unit 213 transmits the parameter estimator to the reproducing apparatus 1 at a predetermined timing, for example, before starting to distribute content.

[0378] The metadata acquisition unit 231 of the metadata decoder 222 of the reproduction device 1 receives and acquires the parameter estimator transmitted from the content generation device 101. The metadata acquisition unit 231 functions as an acquisition unit that acquires the parameter estimator for the location.

[0379] The parameter estimator acquired by the metadata acquisition unit 231 is provided in the rendering parameter generation unit 233 of the metadata decoder 222 , and is used to appropriately generate rendering parameters.

[0380] Figure 38 It shows Figure 37 A block diagram of a configuration example of the parameter estimator learning unit 213 in FIG.

[0381] The parameter estimator learning unit 213 is composed of a learning unit 251 , an estimator DB 252 , and an estimator providing unit 253 .

[0382] The learning unit 251 performs learning of the parameter estimator stored in the estimator DB 252 by using the rendering parameters generated by the metadata encoder 212 as learning data.

[0383] The estimator providing unit 253 controls the transmission control unit 116 ( Figure 22) to transmit the parameter estimator stored in the estimator DB 252 to the reproducing apparatus 1. The estimator supply unit 253 functions as a supply unit that supplies the parameter estimator to the reproducing apparatus 1.

[0384] Therefore, a parameter estimator for each location may be prepared on the content generating apparatus 101 side, and the parameter estimator may be provided to the reproducing apparatus 1 at a predetermined timing, for example, before starting to reproduce content.

[0385] The parameter estimator for each site is placed Figure 36 Although the content is generated in the content generating device 101 at each location in the example of FIG, it may be generated by a server connected to the Internet 201.

[0386] Figure 39 is a block diagram showing still another configuration example of the distribution system.

[0387] Figure 39 The management server 301 in receives rendering parameters sent from the content generating devices 101-1 to 101-3 placed at the sites #1-1 to #1-3, and learns parameter estimators for the respective sites. In other words, the management server 301 has Figure 38 The parameter estimator learning unit 213 in.

[0388] In the case where the reproduction apparatus 1 reproduces content of a music live performance at a predetermined venue, the management server 301 transmits a parameter estimator for the venue to the reproduction apparatus 1. The reproduction apparatus 1 appropriately uses the parameter estimator transmitted from the management server 301 to reproduce audio.

[0389] Therefore, the parameter estimator may be provided via the management server 301 connected to the Internet 201. Note that learning of the parameter estimator may be performed on the content generating apparatus 101 side, and the generated parameter estimator may be provided to the management server 301.

[0390] Note that the embodiments of the present technology are not limited to the above-described embodiments, and various changes can be made within the scope not departing from the gist of the present technology.

[0391] For example, the present technology may adopt a configuration of cloud computing in which one function is shared and collaboratively processed by a plurality of devices via a network.

[0392] In addition, each step described in the above flowchart may be executed by one device, or may be shared and executed by a plurality of devices.

[0393] Furthermore, when a plurality of processes are included in one step, the plurality of processes included in one step may be executed by one device, or may be shared and executed by a plurality of devices.

[0394] The effects described in the specification are merely examples and are not limiting, and other effects may be exerted.

[0395] -About the program

[0396] The above-described series of processing can be executed by hardware or can be executed by software. In the case where the series of processing is executed by software, the program constituting the software is installed in a computer incorporated into dedicated hardware, a general-purpose personal computer, or the like.

[0397] The program to be installed is recorded in Figure 7 The program is provided on a removable medium 22 shown, which is composed of an optical disk (a compact disk read-only memory (CD-ROM), a digital versatile disk (DVD, etc.), a semiconductor memory, etc.). In addition, the program can be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting. The program can also be pre-installed in the ROM 12 and the storage unit 19.

[0398] Note that the program executed by the computer may be one in which processing is performed in time series according to the order described in the specification, or may be one in which processing is performed in parallel or at necessary timing such as when a call is made.

[0399] -About the combination

[0400] The present technology can also adopt the following configurations.

[0401] (1) A reproduction device comprising:

[0402] an acquisition unit that acquires content including sound data of each of the audio objects and rendering parameters of the sound data for each of a plurality of assumed listening positions; and

[0403] A rendering unit renders the sound data based on rendering parameters for the selected predetermined assumed listening position and outputs a sound signal.

[0404] (2) The reproduction device according to (1), wherein the content further includes information on a pre-set assumed listening position, and

[0405] The reproducing apparatus further includes a display control unit that causes a screen for selecting an assumed listening position to be displayed based on the information about the assumed listening position.

[0406] (3) The reproduction apparatus according to (1) or (2), wherein the rendering parameter for each of the hypothetical listening positions includes positioning information representing a position at which the audio object is positioned and gain information as a parameter for gain adjustment of the sound data.

[0407] (4) The reproduction apparatus according to any one of (1) to (3), wherein the rendering unit renders sound data of an audio object selected as an audio object whose sound source position is fixed, based on a rendering parameter different from the rendering parameter for the selected hypothetical listening position.

[0408] (5) The reproduction apparatus according to any one of (1) to (4), wherein the rendering unit does not render sound data of a predetermined audio object among a plurality of audio objects constituting sound of the content.

[0409] (6) The reproduction apparatus according to any one of (1) to (5), further comprising a generation unit that generates a rendering parameter for each of the audio objects for a hypothetical listening position for which a rendering parameter is not prepared, based on a rendering parameter for a hypothetical listening position,

[0410] wherein the rendering unit renders sound data of each of the audio objects by using the rendering parameter generated by the generation unit.

[0411] (7) The reproduction apparatus according to (6), wherein the generation unit generates the rendering parameter for a hypothetical listening position for which the rendering parameter is not prepared, based on a rendering parameter for a plurality of nearby hypothetical listening positions for which the rendering parameter is prepared.

[0412] (8) The reproduction apparatus according to (6), wherein the generation unit generates the rendering parameter for a hypothetical listening position for which a rendering parameter is not prepared, based on a rendering parameter included in the content acquired in the past.

[0413] (9) The reproduction apparatus according to (6), wherein the generation unit generates the rendering parameter for a hypothetical listening position for which a rendering parameter is not prepared, by using an estimator.

[0414] (10) The reproduction apparatus according to (9), wherein the acquisition unit acquires the estimator corresponding to a place where the content is recorded, and

[0415] the generation unit generates the rendering parameter by using the estimator acquired by the acquisition unit.

[0416] (11) The reproduction apparatus according to (9) or (10), wherein the estimator is constituted by learning using at least a rendering parameter included in the content acquired in the past.

[0417] (12) The reproduction apparatus according to any one of (1) to (11), wherein the content further includes video data for displaying an image from a supposed listening position that is a viewpoint position, and

[0418] the reproduction apparatus further includes a video reproduction unit that reproduces the video data and causes an image from a selected predetermined supposed listening position that is a viewpoint position to be displayed.

[0419] (13) A reproduction method comprising the steps of:

[0420] acquiring content including sound data of each of audio objects and a rendering parameter of the sound data for each of a plurality of supposed listening positions; and

[0421] rendering the sound data based on the rendering parameter for a selected predetermined supposed listening position, and outputting a sound signal.

[0422] (14) A program causing a computer to execute a process comprising the steps of:

[0423] acquiring content including sound data of each of audio objects and a rendering parameter of the sound data for each of a plurality of supposed listening positions; and

[0424] rendering the sound data based on the rendering parameter for a selected predetermined supposed listening position, and outputting a sound signal.

[0425] (15) An information processing apparatus comprising:

[0426] a parameter generation unit that generates a rendering parameter of sound data of each of audio objects for each of a plurality of supposed listening positions; and

[0427] a content generation unit that generates content including the sound data of each of the audio objects and the generated rendering parameter.

[0428] (16) The information processing apparatus according to (15), wherein the parameter generation unit further generates information on a supposed listening position set in advance, and

[0429] the content generation unit generates the content further including the information on a supposed listening position.

[0430] (17) The information processing device according to (15) or (16), further including a video generating unit that generates video data for displaying an image from the assumed listening position as a viewpoint position,

[0431] The content generating unit generates the content further including the video data.

[0432] (18) The information processing device according to any one of (15) to (17) further includes a learning unit that generates an estimator, which is used to generate the rendering parameters when a position other than the multiple assumed listening positions for which the rendering parameters are generated is set as the listening position.

[0433] (19) The information processing device according to (18), further including a providing unit that provides the estimator to a reproduction device that reproduces the content.

[0434] (20) An information processing method comprising the following steps:

[0435] generating rendering parameters for sound data of each of the audio objects for each of a plurality of assumed listening positions; and

[0436] Content including the sound data of each of the audio objects and the generated rendering parameters is generated.

[0437] Reference Signs List

[0438] 1 Reproduction device

[0439] 33 Audio reproduction unit

[0440] 51 Rendering parameter selection unit

[0441] 52 Object Data Storage Unit

[0442] 53 Viewpoint information display unit

[0443] 54 rendering units

Claims

1. A reproduction device comprising: A processing unit configured to: Retrieve content including sound data for each of the audio objects and rendering parameters for the sound data for each of a plurality of assumed listening positions; as well as rendering the sound data based on rendering parameters for the selected predetermined assumed listening position and outputting a sound signal, The processing unit renders sound data of an audio object whose sound source position is fixed regardless of the selected assumed listening position, specified by a user or a creator, based on rendering parameters different from rendering parameters for the selected assumed listening position.

2. The reproduction device according to claim 1, wherein The content also includes information about pre-set assumed listening positions, and The processing unit is further configured to cause a screen for selecting the assumed listening position to be displayed based on the information about the assumed listening position.

3. The reproduction device according to claim 1, wherein The rendering parameters for each of the assumed listening positions include localization information indicating a position at which the audio object is localized and gain information which is a parameter for gain adjustment of the sound data.

4. The reproduction device according to claim 1, wherein The processing unit does not render sound data of a predetermined audio object among a plurality of audio objects constituting the sound of the content.

5. The reproduction device according to claim 1, wherein The processing unit is further configured to generate rendering parameters for each of the audio objects for an assumed listening position for which no rendering parameters are prepared, based on the rendering parameters for the assumed listening position, The processing unit renders the sound data of each of the audio objects by using the generated rendering parameters.

6. The reproduction device according to claim 1, wherein The content also includes video data for displaying an image from an assumed listening position as a viewpoint position, and The processing unit is further configured to reproduce the video data and cause an image from a selected predetermined assumed listening position as a viewpoint position to be displayed.

7. A reproduction method comprising the following steps: Retrieve content including sound data for each of the audio objects and rendering parameters for the sound data for each of a plurality of assumed listening positions; as well as rendering the sound data based on rendering parameters for the selected predetermined assumed listening position and outputting a sound signal, Here, sound data of an audio object whose sound source position is fixed regardless of the selected assumed listening position, specified by a user or a creator, is rendered based on rendering parameters different from rendering parameters for the selected assumed listening position.

8. A non-transitory computer-readable medium having a program embodied thereon, the program causing a computer to perform a process comprising the steps of: Retrieve content including sound data for each of the audio objects and rendering parameters for the sound data for each of a plurality of assumed listening positions; as well as rendering the sound data based on rendering parameters for the selected predetermined assumed listening position and outputting a sound signal, Here, sound data of an audio object whose sound source position is fixed regardless of the selected assumed listening position, specified by a user or a creator, is rendered based on rendering parameters different from rendering parameters for the selected assumed listening position.

9. An information processing device comprising: A processing unit configured to: generating rendering parameters for sound data of each of the audio objects for each of a plurality of assumed listening positions; as well as generating content including the sound data of each of the audio objects and the generated rendering parameters, Here, the rendering parameters of the sound data of the audio object whose sound source position is fixed regardless of the selected assumed listening position, specified by the user or the creator, are different from the rendering parameters for the selected assumed listening position.

10. An information processing method comprising the following steps: generating rendering parameters for sound data of each of the audio objects for each of a plurality of assumed listening positions; as well as generating content including the sound data of each of the audio objects and the generated rendering parameters, Here, the rendering parameters of the sound data of the audio object whose sound source position is fixed regardless of the selected assumed listening position, specified by the user or the creator, are different from the rendering parameters for the selected assumed listening position.

Citation Information

Patent Citations

  • Automatic multi-channel music mix from multiple audio stems

    CN105075117A

  • Reproduction methods, apparatus and media, information processing methods and apparatus

    CN109983786B

  • Virtual reality television system

    US5714997A

  • Virtual audio arena effect for live TV presentations: system, methods and program products

    US7526790B1