Information processing device and method, and program
By using polar coordinates to place audio objects based on musicality, the method addresses the challenge of aligning video and audio in free-viewpoint content, enhancing the musicality and artistic intent in 3DoF and 6DoF content creation.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- SONY GROUP CORP
- Filing Date
- 2022-10-31
- Publication Date
- 2026-07-22
AI Technical Summary
Existing methods for creating free-viewpoint content struggle to align with the artistic intent of music producers, as they prioritize realism over musicality, leading to mismatched video and audio in 3DoF and 6DoF content.
An information processing device and method that generates metadata sets for multiple objects using polar coordinates, allowing for the placement of audio objects based on musicality rather than physical positions, and includes control viewpoint information to create free-viewpoint content that aligns with the creator's intentions.
Enables the creation of free-viewpoint content that enhances musicality by placing audio objects in polar coordinate spaces, ensuring alignment with the creator's artistic intent and improving the audio-visual experience.
Smart Images

Figure 0007893260000012 
Figure 0007893260000013 
Figure 0007893260000014
Abstract
Description
[Technical Field]
[0001] This technology relates to an information processing device and method, as well as a program, and more particularly to an information processing device and method, as well as a program that enables content playback based on the intentions of the content creator. [Background technology]
[0002] Traditionally, free-viewpoint audio has been primarily used in games, where the positional relationship on the game's displayed image, i.e., the matching of image and sound, is crucial. Therefore, object audio using an absolute coordinate system is employed (see, for example, Patent Document 1).
[0003] On the other hand, in the world of music content, unlike games, the perceived balance of sound and picture is prioritized over picture-to-picture synchronization in order to enhance musicality. Therefore, picture-to-picture synchronization is not achieved in 2-channel stereo content, nor in 5.1-channel multi-channel content.
[0004] Furthermore, in commercial 3DoF (Degree of Freedom) services, musicality takes precedence, resulting in a world where content relies solely on sound, and many instances exist where the video and audio do not match. [Prior art documents] [Patent Documents]
[0005] [Patent Document 1] International Publication No. 2019 / 198540 [Overview of the Initiative] [Problems that the invention aims to solve]
[0006] Incidentally, while the method of representing the position of objects using absolute coordinates as described above can achieve a high sense of realism, it was difficult to create free-viewpoint content that satisfies the musical intent of the music producer. In other words, it was difficult to achieve content playback based on the content creator's intentions in free-viewpoint content.
[0007] This technology was developed in light of these circumstances and aims to enable content playback that is in line with the intentions of the content creator. [Means for solving the problem]
[0008] The first aspect of this technology is an information processing device which generates a plurality of metadata sets consisting of metadata for a plurality of objects, each containing object position information indicating the position of an object as seen from a control viewpoint when the direction from the control viewpoint toward the target position in space is defined as the direction of the median plane; for each of the plurality of control viewpoints, it generates control viewpoint information which includes control viewpoint position information indicating the position of the control viewpoint in space and information indicating the metadata set associated with the control viewpoint from among the plurality of metadata sets; and it includes a control unit which generates content data which includes a plurality of different metadata sets and configuration information which includes the control viewpoint information for the plurality of control viewpoints.
[0009] The first aspect of the information processing method or program of this technology includes the steps of generating a plurality of metadata sets consisting of metadata for a plurality of objects, each containing object position information indicating the position of an object as seen from a control viewpoint when the direction from the control viewpoint toward a target position in space is defined as the direction of the median plane; generating control viewpoint information for each of the plurality of control viewpoints, which includes control viewpoint position information indicating the position of the control viewpoint in space and information indicating the metadata set associated with the control viewpoint from among the plurality of metadata sets; and generating content data which includes a plurality of different metadata sets and configuration information including the control viewpoint information for the plurality of control viewpoints.
[0010] In the first aspect of this technology, multiple metadata sets are generated, each containing metadata for multiple objects, including object position information indicating the position of an object as seen from a control viewpoint when the direction from the control viewpoint toward the target position in space is defined as the direction of the median plane. For each of the multiple control viewpoints, control viewpoint information is generated, which includes control viewpoint position information indicating the position of the control viewpoint in space, and information indicating the metadata set associated with the control viewpoint from among the multiple metadata sets. Content data is then generated, which includes multiple distinct metadata sets and configuration information including the control viewpoint information for the multiple control viewpoints.
[0011] The information processing device for the second aspect of this technology includes an acquisition unit that acquires object position information indicating the position of an object as seen from a control viewpoint when the direction from the control viewpoint toward the target position in space is defined as the direction of the midline plane, and control viewpoint position information indicating the position of the control viewpoint in space; a listener position information acquisition unit that acquires listener position information indicating the listening position in space; and a position calculation unit that calculates listener-referenced object position information indicating the position of the object as seen from the listening position based on the listener position information, the control viewpoint position information of a plurality of control viewpoints, and the object position information of a plurality of control viewpoints.
[0012] The second aspect of the present technology includes an information processing method or program which includes the steps of obtaining object position information indicating the position of an object as seen from a control viewpoint when the direction from the control viewpoint toward the target position in space is defined as the direction of the median plane, and control viewpoint position information indicating the position of the control viewpoint in space, obtaining listener position information indicating the listening position in space, and calculating listener-referenced object position information indicating the position of the object as seen from the listening position based on the listener position information, the control viewpoint position information of a plurality of control viewpoints, and the object position information of a plurality of control viewpoints.
[0013] In a second aspect of the present technology, object position information indicating the position of an object as seen from a control viewpoint when the direction from the control viewpoint in space toward the target position is the direction of the center plane, control viewpoint position information indicating the position of the control viewpoint in the space, and listener position information indicating the listening position in the space are acquired, and listener reference object position information indicating the position of the object as seen from the listening position is calculated based on the listener position information, the control viewpoint position information of a plurality of the control viewpoints, and the object position information of the plurality of the control viewpoints.
Brief Description of the Drawings
[0014] [Figure 1] It is a diagram for explaining 2D audio and 3D audio. [Figure 2] It is a diagram for explaining object arrangement. [Figure 3] It is a diagram for explaining the CVP and the target position TP. [Figure 4] It is a diagram for explaining the positional relationship between the target position TP and the CVP. [Figure 5] It is a diagram for explaining object arrangement. [Figure 6] It is a diagram for explaining the CVP and the object position pattern. [Figure 7] It is a diagram showing an example of the format of configuration information. [Figure 8] It is a diagram showing an example of the frame length index. [Figure 9] It is a diagram showing an example of the format of CVP information. [Figure 10] It is a diagram showing an example of the format of the object metadata set. [Figure 11] It is a diagram showing an example of the arrangement of the CVP in the free viewpoint space. [Figure 12] It is a diagram showing an example of the arrangement of the reverberation object. [Figure 13] It is a diagram showing an example of the configuration of the information processing apparatus. [Figure 14]It is a flowchart for explaining content production processing. [Figure 15] It is a diagram showing an example of the server configuration. [Figure 16] It is a flowchart for explaining distribution processing. [Figure 17] It is a diagram showing an example of the client configuration. [Figure 18] It is a flowchart for explaining the generation processing of reproduced audio data. [Figure 19] It is a diagram for explaining the selection of CVP used for interpolation processing. [Figure 20] It is a diagram for explaining the object three-dimensional position vector. [Figure 21] It is a diagram for explaining the synthesis of object three-dimensional position vectors. [Figure 22] It is a diagram for explaining vector synthesis. [Figure 23] It is a diagram for explaining vector synthesis. [Figure 24] It is a diagram for explaining the contribution rate of each vector during vector synthesis. [Figure 25] It is a diagram for explaining the listener reference object position information according to the orientation of the listener's face. [Figure 26] It is a diagram for explaining the listener reference object position information according to the orientation of the listener's face. [Figure 27] It is a diagram for explaining the grouping of CVP. [Figure 28] It is a diagram for explaining the CVP group and interpolation processing. [Figure 29] It is a diagram for explaining the CVP group and interpolation processing. [Figure 30] It is a diagram for explaining the CVP group and interpolation processing. [Figure 31] It is a diagram for explaining the CVP group and interpolation processing. [Figure 32] It is a diagram showing an example of the format of configuration information. [Figure 33]This figure shows an example of the format for CVP group information. [Figure 34] This is a flowchart explaining the process of generating playback audio data. [Figure 35] This figure shows examples of CVP placement patterns. [Figure 36] This figure shows an example of the placement of CVPs, etc., in a common absolute coordinate system. [Figure 37] This figure shows examples of listener-referenced object position information and listener-referenced gain. [Figure 38] This figure shows examples of CVP placement patterns. [Figure 39] This figure shows an example of the placement of CVPs, etc., in a common absolute coordinate system. [Figure 40] This figure shows examples of listener-referenced object position information and listener-referenced gain. [Figure 41] This figure shows an example of CVP selection used in interpolation processing. [Figure 42] This figure shows examples of listener-referenced object position information and listener-referenced gain. [Figure 43] This figure shows examples of listener-referenced object position information and listener-referenced gain. [Figure 44] This figure shows an example of configuration information. [Figure 45] This is a flowchart explaining the process for calculating the contribution coefficient. [Figure 46] This is a flowchart explaining the process for calculating the normalized contribution coefficient. [Figure 47] This is a flowchart explaining the process for calculating the normalized contribution coefficient. [Figure 48] This diagram explains the selection of CVP on the playback side. [Figure 49] This figure shows an example of the CVP selection screen. [Figure 50] This figure shows an example of configuration information. [Figure 51] This figure shows an example of a client configuration. [Figure 52] This is a flowchart explaining selective interpolation. [Figure 53] This is a diagram showing an example of a computer configuration. [Modes for carrying out the invention]
[0015] The following describes embodiments to which this technology is applied, with reference to the drawings.
[0016] <First Embodiment> <About this technology> This technology enables the creation of free-viewpoint content with artistic merit.
[0017] First, please refer to Figure 1 to explain 2D audio and 3D audio.
[0018] For example, as shown on the left side of Figure 1, in 2D audio, the sound source can only be placed at the height of the listener's ears. 2D audio can represent the movement of the sound source in the forward / backward and left / right directions.
[0019] In contrast, with 3D audio, as shown on the right side of the diagram, the sound source object can be placed above or below the listener's ear level, thereby allowing the vertical movement of the sound source (object) to be represented.
[0020] Furthermore, 3D audio content includes both 3DoF content and 6DoF content.
[0021] For example, with 3DoF content, users can rotate their heads up, down, left, right, and diagonally within the virtual space to view the content. This type of 3DoF content is also known as fixed-viewpoint content.
[0022] In contrast, with 6DoF content, users can rotate their heads up, down, left, right, and diagonally within the space, and can also move to any position within the space to view the content. Such 6DoF content is also called free-viewpoint content.
[0023] The content described below may consist of audio only, or it may consist of video and accompanying audio. However, without making any particular distinction, these types of content will simply be referred to as "content." In particular, the following will describe an example of creating 6DoF content, that is, free-viewpoint content. Also, in the following, audio objects will simply be referred to as "objects."
[0024] In the content creation process, in order to enhance the artistic (musical) aspect, objects are sometimes intentionally placed in locations different from their visible positions, rather than being constrained by their physical arrangement.
[0025] Such object placement can be easily represented using polar coordinate object placement techniques. Figure 2 shows the difference between absolute coordinate and polar coordinate object placement using an example of an actual band performance.
[0026] For example, as shown in Figure 2, in the case of a band performance, instead of physically placing objects (audio objects) corresponding to the vocals and guitar at the positions of the vocalist and guitarist, the placement is based on musicality.
[0027] In Figure 2, the left side of the figure shows the physical positions of the band members performing the song in three-dimensional space, i.e., their positions in physical space (absolute coordinate space). Specifically, object OV11 is the vocalist, object OD11 is the drums, object OG11 is the guitar, and object OB11 is the bass.
[0028] In particular, in this example, object OV11 (vocals) is positioned to the right of the user's front (midline), while object OG11 (guitar) is positioned near the right edge from the user's perspective.
[0029] When creating content, objects (audio objects) corresponding to each band member (instrument) are placed in accordance with the musicality, as shown on the right side of the diagram. Considering musicality means making the music easy to listen to.
[0030] In the diagram, the right side shows the arrangement of audio objects in polar coordinate space.
[0031] Specifically, object OV21 indicates the localization position of the audio object corresponding to object OV11, i.e., the vocal voice (sound).
[0032] Object OV11 is positioned slightly to the right of the midline, but the creators considered the vocals to be the central element of the content, and therefore, object OV21, which corresponds to the vocals, is positioned higher on the midline, that is, high in the center from the user's perspective, using polar coordinates to make it more prominent.
[0033] Objects OG21-1 and OG21-2 are audio objects of guitar chord accompaniment sounds, corresponding to Object OG11. Hereafter, unless there is a need to distinguish between Object OG21-1 and Object OG21-2, they will simply be referred to as Object OG21.
[0034] Here, the monaural object OG21 is not simply placed at the physical position of the guitarist in 3D space, i.e., the position of object OG11. Instead, it is placed at two positions, left and right, in front of the user, based on musical insights. In other words, by placing object OG21 with some width at each of the left and right positions in front of the user, it is possible to achieve an audio expression that envelops the listener (user). In other words, it is possible to express a sense of spaciousness (a sense of being enveloped).
[0035] Objects OD21-1 and OD21-2 are audio objects corresponding to object OD11 (drums), and objects OB21-1 and OB21-2 are audio objects corresponding to object OB11 (bass).
[0036] In the following, when it is not necessary to distinguish between object OD21-1 and object OD21-2, they will simply be referred to as object OD21, and when it is not necessary to distinguish between object OB21-1 and object OB21-2, they will simply be referred to as object OB21.
[0037] Object OD21 is positioned with width at the lower left and right from the user's perspective for the purpose of stability, while Object OB21 is positioned closer to the center and slightly higher than the drum (Object OD21) for stability.
[0038] In this way, the creator places each object in polar coordinate space, taking musicality into consideration, and produces free-viewpoint content.
[0039] Thus, object placement in a polar coordinate system is more suitable for creating free-viewpoint content that incorporates the artistry (musicality) intended by the creator than object placement in an absolute coordinate system, where the physical position of an object is uniquely determined. This technology is a free-viewpoint audio technology that is realized by using multiple polar coordinate system object placement patterns based on the above-mentioned polar coordinate system object placement method.
[0040] By the way, when creating 3DoF content using polar coordinate system object placement, the creator assumes a single listening position in space and places each object using a polar coordinate system centered on the listener (listening position).
[0041] In this case, the metadata for each object mainly consists of three elements: Azimuth, Elevation, and Gain.
[0042] Here, Azimuth is the horizontal angle indicating the object's position as seen by the listener, Elevation is the vertical angle indicating the object's position as seen by the listener, and Gain is the gain of the object's audio data.
[0043] The production tool outputs the metadata described above for each object, along with audio data (object audio data) for playing the sound of the object corresponding to that metadata, as the final output.
[0044] Here, we consider the development of 3DoF content creation methods into free-viewpoint (6DoF) content, that is, into free-viewpoint content.
[0045] For example, as shown in Figure 3, suppose a creator of free-viewpoint content defines a control viewpoint (CVP) as the position of the viewpoint they want to represent within the free-viewpoint space (3D space), and defines multiple CVPs.
[0046] For example, CVP is the desired listening position when playing content. Below, the i-th CVP will also be referred to as CVPi.
[0047] In the example shown in Figure 3, three control viewpoints (CVPs), CVP1 through CVP3, are defined within the free-viewpoint space from which the listener receives the content.
[0048] If we define the coordinate system of absolute coordinates that indicates an absolute position within the free viewpoint space as the common absolute coordinate system, then, as shown in the center of the figure, the common absolute coordinate system is a Cartesian coordinate system with a predetermined position in the free viewpoint space as the origin O, and mutually orthogonal X, Y, and Z axes as its axes.
[0049] In this example, the X-axis represents the horizontal axis, the Y-axis represents the depth axis, and the Z-axis represents the vertical axis. The position of the origin O in the free-viewpoint space is left to the content creator's discretion, but it may be set to the center of the venue, which is the intended free-viewpoint space.
[0050] Here, the coordinates indicating the positions of CVP1 to CVP3 in the common absolute coordinate system, i.e., the absolute coordinate positions of each CVP, are (X1, Y1, Z1), (X2, Y2, Z2), and (X3, Y3, Z3).
[0051] Furthermore, the content creator defines a single position (a single point) in the free-viewpoint space as the target position TP, which is assumed to be the position as seen from all CVPs. The target position TP is the reference position for interpolating the positional information of objects, and in particular, all virtual listeners in each CVP are assumed to be facing the direction of the target position TP.
[0052] In this example, the coordinates (absolute coordinate position) indicating the target position TP in the common absolute coordinate system are (x tp ,y tp ,z tp )
[0053] Furthermore, at each CVP, a polar coordinate space (hereinafter also referred to as the CVP polar coordinate space) is formed centered on the location of the CVP.
[0054] A position within the CVP polar coordinate space is expressed by the coordinates (polar coordinates) of a polar coordinate system (hereinafter also referred to as the CVP polar coordinate system) consisting of mutually orthogonal x, y, and z axes, with the CVP's absolute coordinate position, O', as the origin.
[0055] Specifically, the positive y-axis represents the direction from the CVP's position to the target position TP, the x-axis represents the left-right direction as seen from the virtual listener at the CVP, and the z-axis represents the up-down direction as seen from the virtual listener at the CVP.
[0056] Once the content creator sets (specifies) the CVP and target location TP, the relationship between the horizontal angle (Yaw) and the vertical angle (Pitch) is determined as information indicating the positional relationship between the CVP and the target location TP.
[0057] The horizontal angle "Yaw" is the horizontal angle between the Y-axis of the Common Absolute Coordinate System and the Y-axis of the CVP polar coordinate system. In other words, the angle "Yaw" is the horizontal angle relative to the Y-axis in the Common Absolute Coordinate System, indicating the orientation of the face of a hypothetical listener located in the CVP and looking at the target position TP.
[0058] Furthermore, the vertical angle "Pitch" is the angle between the y-axis of the CVP polar coordinate system and the XY plane, which contains the X and Y axes of the common absolute coordinate system. In other words, the angle "Pitch" is the vertical angle in the common absolute coordinate system that indicates the orientation of the face of a hypothetical listener who is in the CVP and looking at the target position TP, relative to the XY plane.
[0059] Specifically, the coordinates indicating the absolute position of the target location TP are (x tp ,y tp ,z tp ) and the coordinates indicating the absolute coordinate position of a given CVP, i.e., the coordinates in the common absolute coordinate system, are (x cvp ,ycvp , z cvp is assumed to be).
[0060] In such a case, the horizontal angle "Yaw" and the vertical angle "Pitch" of the CVP are obtained by the following formula (1). In other words, the relationship of the following formula (1) holds.
[0061] [Number]
[0062] In FIG. 3, for example, the angle "Yaw2" formed by the straight line obtained by projecting the y-axis of the CVP polar coordinate system of CVP2, which is represented by a dotted line, onto the XY plane and the straight line parallel to the Y-axis of the common absolute coordinate system, which is represented by a dotted line, is the horizontal angle "Yaw" for CVP2 obtained by the formula (1). Similarly, the angle "Pitch2" formed by the y-axis of the CVP polar coordinate system of CVP2 and the XY plane is the vertical angle "Pitch" for CVP2 obtained by the formula (1).
[0063] The transmission side (generation side) of the free viewpoint content transmits, as configuration information, CVP position information indicating the absolute coordinate position of the CVP in the common absolute coordinate system and CVP orientation information including the Yaw and Pitch of the CVP obtained by the formula (1) to the reception side (reproduction side).
[0064] [[ID=2�]] As an alternative means for transmitting the CVP orientation information, coordinates (absolute coordinate values) indicating the target position TP in the common absolute coordinate system may be transmitted from the transmission side to the reception side. That is, instead of the CVP orientation information, target position information indicating the target position TP in the common absolute coordinate system (free viewpoint space) may be stored in the configuration information. In such a case, on the reception side (reproduction side), Yaw and Pitch are calculated for each CVP by the above formula (1) based on the coordinates indicating the received target position TP.
[0065] Furthermore, the CVP orientation information may include not only the CVP Yaw and Pitch, but also the rotation angle (Roll) of the CVP polar coordinate system relative to the common absolute coordinate system, with the y-axis of the CVP polar coordinate system as the axis of rotation. In the following, the Yaw, Pitch, and Roll included in the CVP orientation information will be referred to as CVP Yaw information, CVP Pitch information, and CVP Roll information, respectively.
[0066] Now, focusing on CVP3 shown in Figure 3, we will further explain the CVP orientation information.
[0067] Figure 4 shows the positional relationship between the target position TP and CVP3 shown in Figure 3.
[0068] When CVP3 is set (specified), a CVP polar coordinate system, or a single polar coordinate space, is defined centered on CVP3. In this polar coordinate space, the direction of the target position TP as seen from CVP3 is the direction of the median plane (the direction where the horizontal angle Azimuth=0 and the vertical angle Elevation=0). In other words, the direction from CVP3 towards the target position TP is the positive direction of the y-axis in the CVP polar coordinate system centered on CVP3.
[0069] Here, if we define the X'Y' plane as the plane parallel to the XY plane containing the origin O' of the CVP polar coordinate system of CVP3 and the X and Y axes of the common absolute coordinate system, then line LN11 is the line obtained by projecting the y axis of the CVP polar coordinate system of CVP3 onto the X'Y' plane. Also, line LN12 is the line parallel to the Y axis of the common absolute coordinate system on the X'Y' plane.
[0070] At this time, the angle "Yaw3" between the line LN11 and the line LN12 becomes the CVP Yaw information for CVP3 obtained by equation (1), and the angle "Pitch3" between the y-axis and the line LN11 becomes the CVP Pitch information for CVP3 obtained by equation (1).
[0071] The CVP orientation information, consisting of CVP Yaw and CVP Pitch information obtained in this way, indicates the direction that the virtual listener at CVP3 is facing, that is, the direction from CVP3 to the target position TP in free viewpoint space. In other words, CVP orientation information can be said to be information that shows the relative relationship between the orientation (direction) of the common absolute coordinate system and the CVP polar coordinate system.
[0072] Furthermore, once the polar coordinate system of each CVP is determined, a relative positional relationship is established between the CVP and each object (audio object), which is expressed using the horizontal angle Azimuth and the vertical angle Elevation, indicating the position (direction) of the object as seen from the CVP.
[0073] Refer to Figure 5 to explain the relative position of a given object as seen from CVP3. Note that in Figure 5, the target position TP and CVP3 are shown in the same positional relationship as in Figure 3.
[0074] In this example, four audio objects, including object obj1, are placed within the free viewpoint space.
[0075] In the diagram, the left side shows the entire free viewpoint space, and in particular, object obj1 is placed near the target position TP.
[0076] Content creators can specify the placement position of an object for each CVP, so that even for the same object, its absolute position in the free-view space will differ for each CVP.
[0077] For example, the creator can specify the position of object obj1 in the free viewpoint space as seen from CVP1 and the position of object obj1 in the free viewpoint space as seen from CVP3 separately, and these positions are not necessarily the same.
[0078] In the diagram, the right side shows the view of the target position TP and object obj1 from CVP3's polar coordinate space, indicating that object obj1 is positioned to the left and slightly in front of CVP3.
[0079] In this case, a relative positional relationship exists between CVP3 and object obj1, determined by the horizontal angle Azimuth_obj1 and the vertical angle Elevation_obj1. In other words, the relative position of object obj1 as seen from CVP3 can be expressed using the coordinates (polar coordinates) of the CVP polar coordinate system, which consists of the horizontal angle Azimuth_obj1 and the vertical angle Elevation_obj1.
[0080] This method of object placement using polar coordinates is the same as the placement technique used in 3DoF content creation. In other words, this technology makes it possible to place objects in 6DoF content using the same polar coordinate representation as in 3DoF content.
[0081] As described above, this technology allows objects to be placed in 3D space for each of the multiple CVPs set within the free viewpoint space using the same method as in the case of 3DoF content. This generates object position patterns that correspond to multiple CVPs.
[0082] When creating free-viewpoint content, the creator will specify the placement of all objects for each CVP (Common Viewpoint) they have set up.
[0083] Furthermore, the object placement patterns in each CVP are not limited to corresponding to a single individual CVP; the same placement pattern may be assigned to multiple CVPs. This allows for the specification of object positions for multiple CVPs located over a wider area within the free-viewpoint space while efficiently reducing production costs.
[0084] Figure 6 shows an example of the association between configuration information (CVP set) that manages CVPs and object location patterns (Object Set).
[0085] In the diagram, the left side shows (N+2) individual CVPs, and the configuration information stores information about these (N+2) CVPs. Specifically, for example, the configuration information includes the CVP location information and CVP orientation information for each CVP.
[0086] In contrast, the right side of the diagram shows N different object position patterns, i.e., object arrangement patterns.
[0087] For example, the object position pattern information indicated by the text "OBJ Positions 1" shows the position of all objects in the CVP polar coordinate system when objects are placed according to a specific placement pattern defined by the creator or other relevant parties.
[0088] Therefore, for example, the object placement pattern shown by "OBJ Positions 1" and the object placement pattern shown by "OBJ Positions 2" are different placement patterns.
[0089] Furthermore, the arrows in the diagram, pointing from each CVP shown on the left to the object position patterns shown on the right, represent the link relationship between the CVP and the object position patterns. In this technology, the combination patterns of information about each CVP and the object position information exist independently, and the relationship between the two is managed by linking them together, as shown in Figure 6.
[0090] Specifically, in this technology, for example, an object metadata set is prepared for each object placement pattern.
[0091] For example, an object metadata set consists of object metadata for each object, and each object metadata contains object position information for the object according to its placement pattern.
[0092] This object position information is such that it represents the polar coordinates indicating the position of the object in the CVP polar coordinate system when the object is placed according to the corresponding placement pattern. More specifically, for example, the object position information is coordinate information that represents the position of the object as seen from the CVP in free viewpoint space, where the direction from the CVP towards the target position TP is the direction of the median plane, expressed in polar coordinate coordinates (polar coordinates) similar to the CVP polar coordinate system.
[0093] Furthermore, the configuration information includes a metadata set index for each CVP, which indicates the object metadata set corresponding to the object position pattern (object placement pattern) set by the creator for that CVP. The receiving side (playback side) then obtains the object metadata set for the corresponding object position pattern based on the metadata set index included in the configuration information.
[0094] Linking CVPs and object location patterns (object metadata sets) using such metadata set indexes means that mapping information exists between CVPs and object location patterns. This improves the readability of the format in terms of data management and implementation, while also reducing the amount of memory used.
[0095] For example, by preparing an object metadata set for each object position pattern, content creators can share combination patterns of object position information for multiple existing objects across multiple CVPs (Content Management Platforms) within their production tools.
[0096] Specifically, for example, in the diagram, the object position pattern shown by "OBJ Positions 2" on the right side of the diagram is specified for CVP2 shown on the left side of the diagram. However, the same object position pattern shown by "OBJ Positions 2" can also be specified for CVPN+1 as in the case of CVP2.
[0097] In this case, the object position pattern referenced by CVP2 (associated with CVP2) and the object position pattern referenced by CVPN+1 are both the same "OBJ Positions 2". Therefore, the relative position of the object indicated by "OBJ Positions 2" as seen from CVP2 is the same as the relative position of the object indicated by "OBJ Positions 2" as seen from CVPN+1.
[0098] However, the position of an object in the free viewpoint space indicated by "OBJ Positions 2" in CVP2 is different from the position of an object in the free viewpoint space indicated by "OBJ Positions 2" in CVPN+1. This is because, for example, the object position information in the object position pattern "OBJ Positions 2" is expressed in polar coordinates of a polar coordinate system, but the position of the origin and the direction of the y-axis (direction of the median plane) in the free viewpoint space of that polar coordinate system are different when CVP2 refers to "OBJ Positions 2" and when CVPN+1 refers to "OBJ Positions 2". In other words, the position of the origin and the direction of the y-axis of the CVP polar coordinate system are different between CVP2 and CVPN+1.
[0099] Furthermore, although N object position patterns are provided here, even if new object position patterns are generated later, the management of CVPs and object metadata does not become complicated, and processing can be carried out systematically. This improves the visibility of data in the software and simplifies implementation.
[0100] As described above, this technology allows for the creation of 6DoF content (free-viewpoint content) simply by using the 3DoF method and placing objects on the polar coordinates associated with each CVP.
[0101] The audio data sets corresponding to the objects used in each CVP are selected by the content creator, but these audio data can be used in common across multiple CVPs. Additionally, audio data corresponding to objects used only in specific CVPs may be added.
[0102] By using common audio data across each CVP and controlling the object's position and gain for each CVP, low-redundancy transmission becomes possible.
[0103] A production tool that generates 6DoF content (free-viewpoint content) in response to the creator's operations outputs two data structures, configuration information and an object metadata set, as files or binary data.
[0104] Figure 7 shows an example of the format (syntax) of the configuration information.
[0105] In the example shown in Figure 7, the configuration information includes the frame length index "FrameLengthIndex", the number of objects "NumOfObjects", the number of CVPs "NumOfControlViewpoints", and the number of metadata sets "NumOfObjectMetaSets".
[0106] The FrameLengthIndex is an index that indicates the length of one frame of audio data used to play the sound of an object, that is, how many samples make up one frame.
[0107] For example, the correspondence between each value of the frame length index "FrameLengthIndex" and the frame length indicated by the frame length index is as shown in Figure 8.
[0108] In this example, if the frame length index value is "5", the frame length is set to "1024". That is, one frame consists of 1024 samples.
[0109] Returning to the explanation of Figure 7, the object count information "NumOfObjects" indicates the number of audio data that make up the content, i.e., the number of objects (audio objects). The CVP count information "NumOfControlViewpoints" indicates the number of CVPs (Control Viewpoints) set by the creator. The metadata set count information "NumOfObjectMetaSets" indicates the number of object metadata sets.
[0110] Furthermore, the configuration information includes CVP information "ControlViewpointInfo(i)" for each CVP, as indicated by the CVP count information "NumOfControlViewpoints". In other words, CVP information is stored for each CVP set by the developer.
[0111] Furthermore, the configuration information includes, for each CVP, coordinate mode information "CoordinateMode[i][j]", which is flag information indicating how the object position information included in the object metadata is described for each object.
[0112] For example, a coordinate mode value of "0" indicates that the object's position information is described using absolute coordinates in the Common Absolute Coordinate System. In contrast, a coordinate mode value of "1" indicates that the object's position information is described using polar coordinates in the CVP Polar Coordinate System. The following explanation will assume that the coordinate mode value is "1".
[0113] Furthermore, Figure 9 shows an example of the format (syntax) of the CVP information "ControlViewpointInfo(i)" included in the configuration information.
[0114] In this example, the CVP information includes the CVP index "ControlViewpointIndex[i]" and the metadata set index "AssociatedObjectMetaSetIndex[i]".
[0115] The CVP index "ControlViewpointIndex[i]" is index information used to identify the CVP corresponding to the CVP information.
[0116] The metadata set index "AssociatedObjectMetaSetIndex[i]" is index information (specification information) that indicates the object metadata set specified by the creator for the CVP indicated by the CVP index. In other words, the metadata set index is information that indicates the object metadata set associated with the CVP.
[0117] Furthermore, CVP information includes CVP location information and CVP orientation information.
[0118] In other words, the CVP position information stores the X coordinate "CVPosX[i]", Y coordinate "CVPosY[i]", and Z coordinate "CVPosZ[i]", which indicate the position of the CVP in the Common Absolute Coordinate System.
[0119] Additionally, CVP-oriented information such as CVP Yaw information "CVYaw[i]", CVP Pitch information "CVPitch[i]", and CVP Roll information "CVRoll[i]" is stored.
[0120] Figure 10 shows an example of the format (syntax) of an object metadata set, or more specifically, an example of an object metadata set.
[0121] In this example, "NumOfObjectMetaSets" indicates the number of object metadata sets stored. This number of object metadata sets can be obtained from the metadata set count information included in the configuration information. "ObjectMetaSetIndex[i]" indicates the index of the object metadata set, and "NumOfObjects" indicates the number of objects.
[0122] "NumOfChangePoints" indicates the number of change points, which are the times when the contents of the object metadata set change.
[0123] In this example, the object metadata sets between change points are not stored. Furthermore, for each object metadata set, for each change point for an object, the following are stored: a frame index "frame_index[i][j][k]" to identify the change point, "PosA[i][j][k]", "PosB[i][j][k]", and "PosC[i][j][k]" indicating the object's position, and the object's gain "Gain[i][j][k]". This gain "Gain[i][j][k]" is the gain of the object (audio data) as seen from the CVP.
[0124] The frame index "frame_index[i][j][k]" is an index that indicates the frame of the audio data of the object that marks a change point. On the receiving (playback) side, the sample position of the audio data that marks a change point is identified based on this frame index "frame_index[i][j][k]" and the frame length index "FrameLengthIndex" included in the configuration information.
[0125] The "PosA[i][j][k]", "PosB[i][j][k]", and "PosC[i][j][k]" fields, which indicate the object's position, represent the horizontal angle Azimuth, vertical angle Elevation, and radius Radius of the object in the CVP polar coordinate system. In other words, the information consisting of "PosA[i][j][k]", "PosB[i][j][k]", and "PosC[i][j][k]" constitutes the object's position information.
[0126] However, here we assume that the value of the coordinate mode information is "1". If the value of the coordinate mode information is "0", then "PosA[i][j][k]", "PosB[i][j][k]", and "PosC[i][j][k]" are considered to be the X, Y, and Z coordinates indicating the object's position in the common absolute coordinate system.
[0127] As described above, in the format shown in Figure 10, for each object metadata set, the frame index for each object, along with the object position information and gain for each object, are stored for each change point.
[0128] While an object's position may always be fixed, its position may also change dynamically over time.
[0129] In the format example shown in Figure 10, the position at the time when the object's position changes, i.e., the position of the change point mentioned above, is recorded as the frame index "frame_index[i][j][k]". Furthermore, object position information and gain between change points can be obtained, for example, on the receiving (playback) side through linear interpolation based on the object position information and gain at the change point.
[0130] Thus, by adopting the format shown in Figure 10, it is possible to handle cases where the position of an object changes dynamically, and since it is not necessary to retain data for all time points (frames), the file size can be reduced.
[0131] Figure 11 shows an example of the placement of CVP in a free-viewpoint space when creating actual live content as free-viewpoint content.
[0132] In this example, the entire space, including the live venue, is a free-viewpoint space, where artists, as objects, perform music and other acts on stage ST11. The audience seating is arranged around stage ST11 within the live venue.
[0133] The location of the origin O in the free viewpoint space, or common absolute coordinate system, is left to the content creator's discretion, but in this example, the center of the live venue is set as the location of the origin O.
[0134] In this example, the producer has set a target location TP on stage ST11, and seven CVP1 through CVP7 locations are set within the live venue.
[0135] As mentioned above, in each CVP (CVP polar coordinate system), the direction from the CVP to the target position TP is considered to be the direction of the median plane.
[0136] Therefore, for example, if CVP1 is the viewpoint (listening position) and the user is facing the direction of the target position TP, the content video will be presented as if the user is looking directly at the stage ST11, as shown by arrow Q11.
[0137] Similarly, for example, a user whose viewpoint is CVP4 and who is facing the direction of the target position TP will be presented with content video that makes it appear as if they are viewing the stage ST11 from diagonally in front, as shown by arrow Q12. Furthermore, for example, a user whose viewpoint is CVP5 and who is facing the direction of the target position TP will be presented with content video that makes it appear as if they are viewing the stage ST11 from diagonally behind, as shown by arrow Q13.
[0138] When creating this type of free-viewpoint content, the process of placing objects for a single CVP (Common Viewpoint) by the creator becomes equivalent to the process of creating 3DoF (3 degrees of freedom) content.
[0139] In the example shown in Figure 11, the creator of the free-viewpoint content only needs to perform object placement for one CVP, and then set up six more CVPs to support 6DoF, and perform object placement for those CVPs as well. In this way, this technology allows for the creation of free-viewpoint content using the same process as for 3DoF content.
[0140] By the way, reverberation components in a space are inherently generated by propagation and reflection within the physical space.
[0141] Therefore, if we consider the physical reverberation components that arrive through reflection and propagation within the concert venue as a free-viewpoint space as objects, the arrangement of reverberation objects in the CVP polar coordinate system (CVP polar coordinate space) will be as shown in Figure 12, for example. Note that in Figure 12, the target position TP and CVP3 shown in Figure 3 are shown in the same positional relationship.
[0142] In Figure 12, the reverberation sound emitted from a sound source (performer) near the target position TP and traveling towards CVP3 is represented by dotted arrows. In other words, the arrows in the figure represent physical reverberation paths. Specifically, four reverberation paths are depicted here.
[0143] Therefore, if those reverberations are placed directly into the CVP polar coordinate space of CVP3 as reverberation objects, each reverberation object will be placed at positions P11 through P14.
[0144] However, if a signal with strong reverberation components in such a narrow region (free viewpoint space) is used as is, while a high sense of presence can be obtained, the original musical signal is often drowned out by the reverberation, resulting in a lower musical quality.
[0145] By using the CVP polar coordinate system for object placement in the production tools of this technology, content creators can treat reverberation components as objects (reverberation objects), determine the optimal direction of arrival from a musical perspective, and place objects according to that determination. This allows for the creation of reverberation effects within a space.
[0146] Specifically, for example, if CVP3 is the listening position, placing reverberation objects at positions P11 to P14 would concentrate the reverberant sound in a narrow area at the front of the concert hall. Therefore, the producer places these reverberation objects at positions P'11 to P'14, which are behind the listener at CVP3. In other words, taking reverberation paths and other factors into consideration, the placement of reverberation objects is moved from positions P11 to P14 to positions P'11 to P'14 in order to add musicality.
[0147] By intentionally positioning reverberation objects behind the listener in this way, it is possible to avoid the music being drowned out by reverberation, that is, making the music itself difficult to hear.
[0148] <Example of information processing device configuration> Next, we will describe the information processing device that realizes the production tools and produces the free-viewpoint content described above.
[0149] Such an information processing device consists of, for example, a personal computer and is configured as shown in Figure 13.
[0150] The information processing device 11 shown in Figure 13 includes an input unit 21, a display unit 22, a recording unit 23, a communication unit 24, an acoustic output unit 25, and a control unit 26.
[0151] The input unit 21 consists of, for example, a mouse, keyboard, touch panel, buttons, switches, etc., and supplies signals to the control unit 26 according to the operations of the content creator. The display unit 22 displays any image, such as the display screen of a content creation tool, in accordance with the control of the control unit 26.
[0152] The recording unit 23 records various data such as audio data for each object for content creation, configuration information and object metadata sets supplied from the control unit 26, and supplies the recorded data to the control unit 26 as appropriate.
[0153] The communication unit 24 communicates with external devices such as servers. For example, the communication unit 24 transmits data supplied from the control unit 26 to the server, etc., and receives data transmitted from the server, etc., and supplies it to the control unit 26.
[0154] The audio output unit 25 consists of, for example, a speaker, and outputs sound based on audio data supplied from the control unit 26.
[0155] The control unit 26 controls the operation of the entire information processing device 11. For example, the control unit 26 generates (creates) free-viewpoint content based on signals supplied from the input unit 21 in response to the creator's operation.
[0156] <Explanation of content creation process> Next, the operation of the information processing device 11 will be explained. Specifically, the content creation process by the information processing device 11 will be explained below with reference to the flowchart in Figure 14.
[0157] For example, when the control unit 26 reads and executes a program recorded in the recording unit 23, a production tool for creating free-viewpoint content is realized.
[0158] When the production tool is started, the control unit 26 supplies predetermined image data to the display unit 22 to display the production tool's display screen. This display screen shows, for example, an image of a free viewpoint space.
[0159] Furthermore, for example, the control unit 26, in response to the creator's operation, reads audio data from the recording unit 23 for each object that constitutes the free-viewpoint content to be created and supplies it to the sound output unit 25 to play the sound of the object.
[0160] The creator performs operations for content creation by operating the input unit 21 while listening to the sounds of objects and checking images of the free viewpoint space on the display screen as needed.
[0161] In step S11, the control unit 26 sets the target position TP.
[0162] For example, the creator can specify any position (point) in the free viewpoint space as the target position TP by operating the input unit 21.
[0163] When the creator specifies a target position TP, the input unit 21 supplies a signal to the control unit 26 corresponding to the creator's operation. Based on the signal supplied from the input unit 21, the control unit 26 sets the position specified by the creator in the free viewpoint space as the target position TP. In other words, the control unit 26 sets the target position TP.
[0164] Furthermore, the control unit 26 may be configured to set the position of the origin O in the free viewpoint space (common absolute coordinate system) according to the creator's operation.
[0165] In step S12, the control unit 26 sets the number of CVPs and the number of object metadata sets to 0. That is, the control unit 26 sets the value of the CVP count information "NumOfControlViewpoints" to 0 and the value of the metadata set count information "NumOfObjectMetaSets" to 0.
[0166] In step S13, the control unit 26 determines, based on the signal from the input unit 21, whether the editing mode selected by the creator's operation is the CVP editing mode.
[0167] Here, we assume that there are three editing modes: CVP editing mode for editing CVPs, object metadata set editing mode for editing object metadata sets, and association editing mode for associating (linking) CVPs with object metadata sets.
[0168] If it is determined in step S13 that the CVP editing mode is active, in step S14 the control unit 26 determines whether or not to change the CVP configuration.
[0169] For example, if a creator performs an operation in CVP editing mode to add (configure) a new CVP or delete an existing CVP, it is determined that the CVP configuration has been changed.
[0170] If it is determined in step S14 that the CVP configuration will not be changed, the process returns to step S13, and the process described above is repeated.
[0171] If, in step S14, it is determined that the CVP configuration should be changed, the process then proceeds to step S15.
[0172] In step S15, the control unit 26 updates the CVP count based on the signal supplied from the input unit 21 in response to the operator's operation.
[0173] For example, if the creator adds (sets) a new CVP, the control unit 26 updates the CVP count by adding 1 to the value of the CVP count information it holds, "NumOfControlViewpoints". Conversely, if the creator deletes an existing CVP, the control unit 26 updates the CVP count by subtracting 1 from the value of the CVP count information it holds, "NumOfControlViewpoints".
[0174] In step S16, the control unit 26 edits the CVP according to the creator's operation.
[0175] For example, when the creator performs an operation to specify (add) a CVP, the control unit 26 sets the location specified by the creator in the free viewpoint space as the location of the newly added CVP, based on the signal supplied from the input unit 21. In other words, the control unit 26 sets a new CVP. Also, for example, when the creator performs an operation to delete a CVP, the control unit 26 deletes the CVP specified by the creator in the free viewpoint space, based on the signal supplied from the input unit 21.
[0176] Once the CVP is edited, the process returns to step S14, and the process described above is repeated. That is, a new CVP is edited.
[0177] Furthermore, if it is determined in step S13 that the system is not in CVP editing mode, the control unit 26 determines in step S17 whether or not it is in object metadata set editing mode.
[0178] If it is determined in step S17 that the object metadata set is in editing mode, the control unit 26 determines in step S18 whether or not to modify the object metadata set.
[0179] For example, if a creator performs an operation in object metadata set editing mode to add (set) a new object metadata set, or to delete an existing object metadata set, it is determined that the object metadata set has been modified.
[0180] If it is determined in step S18 that the object metadata set will not be changed, the process returns to step S13, and the process described above is repeated.
[0181] If, in step S18, it is determined that the object metadata set should be modified, the process then proceeds to step S19.
[0182] In step S19, the control unit 26 updates the number of object metadata sets based on the signal supplied from the input unit 21 in response to the creator's operation.
[0183] For example, if a creator adds (sets) a new object metadata set, the control unit 26 updates the number of object metadata sets by adding 1 to the value of the metadata set count information it holds, "NumOfObjectMetaSets". Conversely, if a creator deletes an existing object metadata set, the control unit 26 updates the number of object metadata sets by subtracting 1 from the value of the metadata set count information it holds, "NumOfObjectMetaSets".
[0184] In step S20, the control unit 26 edits the object metadata set according to the creator's operation.
[0185] For example, when the creator performs an operation to set (add) a new object metadata set, the control unit 26 generates a new object metadata set based on the signal supplied from the input unit 21.
[0186] At this time, for example, the control unit 26 displays an image in CVP polar coordinate space on the display unit 22 as appropriate, and the creator specifies the position (point) on the image in CVP polar coordinate space as the placement position of the object for the new object metadata set.
[0187] When the creator specifies the placement position of one or more objects, the control unit 26 generates a new object metadata set by setting the positions specified by the creator in the CVP polar coordinate space as the placement positions of the objects.
[0188] Furthermore, for example, when the creator performs an operation to delete an object metadata set, the control unit 26 deletes the object metadata set specified by the creator based on the signal supplied from the input unit 21.
[0189] Once the object metadata set is edited, the process returns to step S18, and the process described above is repeated. That is, a new object metadata set is edited. Note that it is also possible to modify an existing object metadata set as part of editing the object metadata set.
[0190] Furthermore, if it is determined in step S17 that the object metadata set editing mode is not in progress, the control unit 26 determines in step S21 whether or not the system is in linking editing mode.
[0191] If it is determined in step S21 that the mode is linked editing, in step S22 the control unit 26 associates the CVP with the object metadata set according to the creator's operation.
[0192] Specifically, for example, the control unit 26 generates a metadata set index "AssociatedObjectMetaSetIndex[i]" that indicates the object metadata set specified by the creator for the CVP specified by the creator, based on the signal supplied from the input unit 21. This associates the CVP with the object metadata set.
[0193] Once the process in step S22 is completed, the process returns to step S13, and the process described above is repeated.
[0194] In contrast, if it is determined in step S21 that the system is not in linked editing mode, that is, if the production of free-viewpoint content is instructed to end, the process proceeds to step S23.
[0195] In step S23, the control unit 26 outputs content data via the communication unit 24.
[0196] For example, the control unit 26 generates configuration information based on the target location TP, CVP, and object metadata set settings, as well as the results of linking the CVP and object metadata set.
[0197] Specifically, for example, the control unit 26 generates configuration information including frame length index, object count information, CVP count information, metadata set count information, CVP information, and coordinate mode information, as explained with reference to Figures 7 and 9. At this time, the control unit 26 calculates CVP orientation information by performing calculations similar to those in equation (1) above as appropriate, and generates CVP information including CVP index, metadata set index, CVP position information, and CVP orientation information.
[0198] Furthermore, in the control unit 26, as explained with reference to Figure 10, the processing in step S20 generates multiple object metadata sets for each change point, each containing a frame index for each object, object position information for each object, and gain.
[0199] As a result, content data is generated for each free-viewpoint content, consisting of audio data for each object, configuration information, and multiple sets of object metadata that are different from each other. The control unit 26 supplies the generated content data to the recording unit 23 for recording as appropriate, and also supplies it to the communication unit 24.
[0200] The communication unit 24 outputs content data supplied from the control unit 26. That is, the communication unit 24 transmits content data to the server via the network at any time. The content data may also be supplied to a recording medium and provided to the server via the recording medium.
[0201] Once the content data is output, the content creation process is complete.
[0202] As described above, the information processing device 11 performs actions such as setting the target position TP, setting the CVP, and setting the object metadata set in response to the creator's operation, and generates content data consisting of audio data, configuration information, and the object metadata set.
[0203] In this way, the playback device can play back free-viewpoint content based on the object placement specified by the creator. Therefore, it is possible to achieve musical content playback that is in line with the creator's intentions.
[0204] <Example Server Configuration> Next, we will describe a server that receives content data for free-viewpoint content from the information processing device 11 and distributes that content data to the client.
[0205] Such a server is configured as shown in Figure 15, for example.
[0206] The server 51 shown in Figure 15 consists of an information processing device such as a computer. The server 51 has a communication unit 61, a control unit 62, and a recording unit 63.
[0207] The communication unit 61 communicates with the information processing device 11 and the client according to the control of the control unit 62. For example, the communication unit 61 receives content data of free-viewpoint content transmitted from the information processing device 11 and supplies it to the control unit 62, or transmits the encoded bitstream supplied by the control unit 62 to the client.
[0208] The control unit 62 controls the operation of the entire server 51. For example, the control unit 62 has an encoding unit 71, which generates an encoded bitstream by encoding the content data of the free-viewpoint content.
[0209] The recording unit 63 records various types of data, such as content data of free-viewpoint content supplied from the control unit 62, and supplies the recorded data to the control unit 62 as needed. In the following, it is assumed that the content data of free-viewpoint content received from the information processing device 11 is recorded in the recording unit 63.
[0210] <Explanation of distribution process> When server 51 receives a request for free-viewpoint content from a client connected via the network, it performs a distribution process to deliver the free-viewpoint content in response to the request. The distribution process by server 51 will be explained below with reference to the flowchart in Figure 16.
[0211] In step S51, the control unit 62 generates an encoded bitstream.
[0212] Specifically, the control unit 62 reads the content data of the free-viewpoint content from the recording unit 63. The encoding unit 71 of the control unit 62 then encodes the audio data, configuration information, and multiple object metadata sets of each object that make up the read content data to generate an encoded bitstream. The control unit 62 then supplies the obtained encoded bitstream to the communication unit 61.
[0213] In this case, for example, the encoding unit 71 encodes audio data, configuration information, and object metadata sets according to the encoding scheme used in MPEG (Moving Picture Experts Group)-I or MPEG-H. This reduces the amount of data transmitted. Furthermore, since the audio data for objects is common to all CVPs, it is sufficient to store one audio data file per object, regardless of the number of CVPs.
[0214] In step S52, the communication unit 61 transmits the encoded bitstream supplied by the control unit 62 to the client, and the distribution process ends.
[0215] In this example, encoded audio data, configuration information, and object metadata sets are multiplexed to generate a single encoded bitstream. However, the configuration information and object metadata sets may be sent to the client at different times than the audio data. For example, the configuration information and object metadata sets may be sent to the client first, followed by the audio data.
[0216] In this manner, server 51 generates an encoded bitstream containing audio data, configuration information, and an object metadata set, and sends it to the client. This allows the client to achieve musical content playback based on the content creator's intentions.
[0217] <Example of client structure> Furthermore, a client that receives an encoded bitstream from server 51 and generates playback audio data for playing free-viewpoint content is configured, for example, as shown in Figure 17.
[0218] The client 101 shown in Figure 17 consists of an information processing device such as a personal computer or a smartphone. The client 101 has a listener location information acquisition unit 111, a communication unit 112, a decoding unit 113, a location calculation unit 114, and a rendering processing unit 115.
[0219] The listener location information acquisition unit 111 acquires listener location information, which indicates the absolute position of the listener within the free viewpoint space, i.e., the listening position, input by the user who will be the listener, and supplies it to the position calculation unit 114.
[0220] For example, listener position information may be defined as absolute coordinates indicating the listener's position in a free viewpoint space, i.e., a common absolute coordinate system.
[0221] Furthermore, the listener position information acquisition unit 111 may also acquire listener orientation information indicating the direction of the listener's face in the free viewpoint space (common absolute coordinate system) and supply it to the position calculation unit 114.
[0222] The communication unit 112 receives the encoded bitstream transmitted from the server 51 and supplies it to the decoding unit 113. In other words, the communication unit 112 functions as an acquisition unit that obtains the audio data, configuration information, and object metadata set of each encoded object contained in the encoded bitstream.
[0223] The decoding unit 113 decodes the encoded bitstream supplied from the communication unit 112, i.e., the encoded audio data, configuration information, and object metadata set of each object. The decoding unit 113 supplies the audio data of each object obtained by decoding to the rendering processing unit 115, and supplies the configuration information and object metadata set obtained by decoding to the position calculation unit 114.
[0224] The position calculation unit 114 calculates listener-referenced object position information, which indicates the position of each object as seen from the listener (listening position), based on the listener position information supplied from the listener position information acquisition unit 111 and the configuration information and object metadata set supplied from the decoding unit 113.
[0225] The object's position, as indicated by the listener-referenced object position information, is expressed using polar coordinates (polar coordinates) in a polar coordinate system with the listening position as the origin (reference), and represents the relative position of the object as seen from the listener (listening position).
[0226] For example, listener-referenced object position information is calculated by interpolation based on the CVP position information and object position information of all or some CVPs, and listener position information. The interpolation process can be any method, such as vector synthesis. In addition, CVP orientation information and listener orientation information may also be used in calculating listener-referenced object position information.
[0227] Furthermore, the position calculation unit 114 calculates the listener-referenced gain of each object at the listening position indicated by the listener position information by interpolation processing, based on the gain of each object included in the object metadata set supplied from the decoding unit 113. The listener-referenced gain is the gain of the object as seen from the listening position.
[0228] The position calculation unit 114 supplies the listener reference gain and listener reference object position information for each object at the listening position to the rendering processing unit 115.
[0229] The rendering processing unit 115 performs rendering processing based on the audio data of each object supplied from the decoding unit 113 and the listener reference gain and listener reference object position information supplied from the position calculation unit 114, and generates playback audio data.
[0230] In the rendering processing unit 115, rendering is performed in a polar coordinate system as defined by MPEG-H, such as VBAP (Vector Based Amplitude Panning), to generate playback audio data. This playback audio data is for playing the sound of the free-viewpoint content, including the sounds of all objects.
[0231] <Explanation of the playback audio data generation process> Next, we will describe the operation of client 101. Specifically, we will describe the playback audio data generation process by client 101, referring to the flowchart in Figure 18 below.
[0232] In step S81, the communication unit 112 receives the encoded bitstream transmitted from the server 51 and supplies it to the decoding unit 113.
[0233] In step S82, the decoding unit 113 decodes the encoded bitstream supplied from the communication unit 112.
[0234] The decoding unit 113 supplies the audio data of each object obtained by decoding to the rendering processing unit 115, and also supplies the configuration information and object metadata set obtained by decoding to the position calculation unit 114.
[0235] Furthermore, the configuration information and object metadata set may be received at a different time than the audio data.
[0236] In step S83, the listener location information acquisition unit 111 acquires the listener's location information at the current time and supplies it to the location calculation unit 114. The listener location information acquisition unit 111 may also acquire listener orientation information and supply it to the location calculation unit 114.
[0237] In step S84, the position calculation unit 114 performs interpolation processing based on the listener position information supplied from the listener position information acquisition unit 111 and the configuration information and object metadata set supplied from the decoding unit 113.
[0238] Specifically, for example, the position calculation unit 114 calculates listener-reference object position information by performing vector synthesis as an interpolation process, and also calculates listener-reference gain through the interpolation process, and supplies this listener-reference object position information and listener-reference gain to the rendering processing unit 115.
[0239] In addition, when performing interpolation processing, the current time (sample) may be the time between change points, and the object metadata may not contain object position information or gain for the current time. In such cases, the position calculation unit 114 calculates the object position information and gain at the current time in CVP by performing interpolation processing based on the object position information and gain at multiple change points close to the current time, such as immediately before and immediately after the current time.
[0240] In step S85, the rendering processing unit 115 performs rendering based on the audio data of each object supplied from the decoding unit 113 and the listener reference gain and listener reference object position information supplied from the position calculation unit 114.
[0241] For example, the rendering processing unit 115 performs gain correction on the audio data of each object based on the listener reference gain of each object.
[0242] The rendering processing unit 115 then performs rendering processing such as VBAP based on the audio data of each object after gain correction and the listener-referenced object position information to generate playback audio data.
[0243] The rendering processing unit 115 outputs the generated playback audio data to subsequent blocks such as speakers.
[0244] This approach allows for the playback of multi-viewpoint free-view content (6DoF content), where any position within the free-viewpoint space is used as the listening position.
[0245] In step S86, client 101 determines whether or not to terminate the process. For example, in step S86, if the encoded bitstream has been received for all frames of the free-viewpoint content and playback audio data has been generated, it is determined to terminate the process.
[0246] In step S86, if it is determined that the process has not yet ended, the process returns to step S81, and the above-described process is repeated.
[0247] On the other hand, if it is determined in step S86 that the process ends, the client 101 terminates the operations of each part, and the reproduced audio data generation process ends.
[0248] As described above, the client 101 performs interpolation processing based on the listener position information, configuration information, and object metadata set, and calculates the listener reference gain and listener reference object position information at the listening position.
[0249] By doing so, it is possible to realize music content reproduction based on the intention of the content producer according to the listening position, rather than just the physical relationship between the listener and the object, and to sufficiently convey the interestingness of the content to the listener.
[0250] 〈Regarding Interpolation Processing〉 Here, a specific example of the interpolation processing performed in step S84 of FIG. 18 will be described. In particular, the case where polar coordinate vector synthesis is performed will be described here.
[0251] For example, as shown in FIG. 19, assume that an arbitrary viewpoint position indicated by the listener position information in the free viewpoint space (common absolute coordinate system), that is, the position of the listener at the current time, is the listening position LP11. Note that FIG. 19 shows a bird's-eye view of the free viewpoint space from above.
[0252] And For example, if a predetermined object is set as the target object, in order to generate the reproduced audio data at the listening position LP11 by rendering processing, listener reference object position information indicating the position PosF of the target object in the polar coordinate system with the listening position LP11 as the origin is required.
[0253] Therefore, the position calculation unit 114 selects, for example, several CVPs around the listening position LP11 as CVPs to be used for interpolation processing. In this example, three CVPs, CVP0 to CVP2, are selected as CVPs to be used for interpolation processing.
[0254] For example, the selection of CVPs can be carried out in any way, such as selecting three or more predetermined CVPs that are located around the listening position LP11 and have the shortest distance from the listening position LP11. Alternatively, all CVPs may be used for interpolation processing. In this case, the position calculation unit 114 can determine the position of each CVP in the common absolute coordinate system by referring to the CVP position information included in the configuration information.
[0255] When CVP0 to CVP2 are selected as the CVP to be used for interpolation, the position calculation unit 114 obtains, for example, the object's 3D position vector shown in Figure 20.
[0256] In Figure 20, the position of the object of interest in the polar coordinate space of CVP0 is shown on the left side of the figure. In this example, position Pos0 is the position of the object of interest as seen from CVP0, and the position calculation unit 114 calculates a vector V11 as the object's 3D position vector with respect to CVP0, starting from the origin O' of the CVP polar coordinate system of CVP0 and ending at position Pos0.
[0257] Furthermore, the position of the object of interest in the polar coordinate space of CVP1 is shown in the center of the figure. In this example, position Pos1 is the position of the object of interest as seen from CVP1, and the position calculation unit 114 calculates a vector V12 as the 3D position vector of the object with respect to CVP1, starting from the origin O' of the CVP polar coordinate system of CVP1 and ending at position Pos1.
[0258] Similarly, the right side of the figure shows the position of the object of interest in the polar coordinate space of CVP2. In this example, position Pos2 is the position of the object of interest as seen from CVP2, and the position calculation unit 114 calculates a vector V13 as the 3D position vector of the object with respect to CVP2, starting from the origin O' of the CVP polar coordinate system of CVP2 and ending at position Pos2.
[0259] Here, we will explain the specific method for calculating the object's 3D position vector.
[0260] For example, if we define the absolute coordinate system (orthogonal coordinate system) as having the origin O' of the CVP polar coordinate system of CVPi as its origin, and using the x, y, and z axes of that CVP polar coordinate system as the x, y, and z axes, then we will call it the CVP absolute coordinate system (CVP absolute coordinate space).
[0261] The object's 3D position vector is a vector expressed using coordinates in the CVP absolute coordinate system.
[0262] For example, suppose the polar coordinates representing the position Posi of the object of interest in the CVP polar coordinate system of CVPi are (Azi[i], Ele[i], rad[i]). These Azi[i], Ele[i], and rad[i] correspond to PosA[i][j][k], PosB[i][j][k], and PosC[i][j][k] as explained with reference to Figure 10. Also, suppose the gain of the object of interest as seen from CVPi is denoted as g[i]. This gain g[i] corresponds to Gain[i][j][k] as explained with reference to Figure 10.
[0263] Furthermore, let's assume that the absolute coordinates representing the position (Posi) of the object of interest in the CVP absolute coordinate system of CVPi are (vx[i], vy[i], vz[i]).
[0264] In this case, the 3D position vector of the object for CVPi is (vx[i], vy[i], vz[i]), and this 3D position vector of the object can be obtained by the following equation (2).
[0265]
number
[0266] The position calculation unit 114 reads a metadata set index from the CVP information of the CVPi included in the configuration information, which indicates the object metadata set referenced by the CVPi. The position calculation unit 114 also reads the object position information and gain of the object of interest in the CVPi from the object metadata of the object of interest that constitutes the object metadata set indicated by the read metadata set index.
[0267] Then, the position calculation unit 114 calculates equation (2) based on the object position information of the object of interest in CVPi and obtains the object's 3D position vector (vx[i], vy[i], vz[i]). This calculation of equation (2) is a conversion from polar coordinates to absolute coordinates.
[0268] The position calculation unit 114 obtains vectors V11 to V13, which are the three-dimensional position vectors of the object shown in Figure 20, using equation (2), and then calculates the vector sum of these vectors V11 to V13, for example, as shown in Figure 21. In Figure 21, the parts corresponding to those in Figure 20 are denoted by the same reference numerals, and their explanations are omitted as appropriate.
[0269] In this example, the vector sum of vectors V11 through V13 is calculated, and as a result, vector V21 is obtained. In other words, vector V21 is obtained by vector composition based on vectors V11 through V13.
[0270] More specifically, the contribution rate of each CVP (object 3D position vector) in the calculation of the vector V21 indicating the position PosF is weighted, and the vectors V11 to V13 are synthesized based on these weights, thereby obtaining the vector V21. In FIG. 21, for simplicity of explanation, the contribution rate of each CVP is set to 1.
[0271] This vector V21 is a vector indicating the position PosF of the target object in the absolute coordinate system when viewed from the listening position LP11. This absolute coordinate system has the listening position LP11 as the origin and the direction from the listening position LP11 to the target position TP as the positive direction of the y-axis.
[0272] For example, if the absolute coordinates representing the position PosF in the absolute coordinate system with the listening position LP11 as the origin are (vxF, vyF, vzF), the vector V21 is (vxF, vyF, vzF).
[0273] Also, if the gain of the target object when viewed from the listening position LP11 is gF and the contribution rates of CVP0 to CVP2 are dep[0] to dep[2], the vector (vxF, vyF, vzF), that is, the vector V21 and the gain gF can be obtained by the following equation (3).
[0274]
Equation
[0275] By converting the vector V21 thus obtained into polar coordinates indicating the position PosF of the target object in the polar coordinate system with the listening position LP11 as the origin, listener reference object position information can be obtained. Also, the gain gF obtained by equation (3) is the listener reference gain.
[0276] In this technology, by setting one target position TP common to all CVPs, the target listener reference object position information and the listener reference gain can be obtained by simple calculation.
[0277] Now, let's explain vector addition further.
[0278] For example, as shown on the left side of Figure 22, let's assume that position LP21 in the free viewpoint space is the listening position. Furthermore, let's assume that CVP1 to CVP5 are set in the free viewpoint space, and that these CVP1 to CVP5 are used to determine the listener reference object position information. Note that in Figure 22, for the sake of simplicity, the CVPs are shown as being arranged on a two-dimensional plane.
[0279] In this example, CVP1 through CVP5 are located around the target position TP. Furthermore, the positive y-axis direction of each CVP's polar coordinate system points from the CVP towards the target position TP.
[0280] Furthermore, the positions of the same object of interest as viewed from each of CVP1 through CVP5 correspond to positions OBP1 through OBP5. In other words, the polar coordinates of the CVP polar coordinate system that represent positions OBP1 through OBP5 are the object position information for CVP1 through CVP5.
[0281] At this point, assume that each CVP is rotated so that the y-axis of the CVP polar coordinate system is vertical, i.e., upward in the figure. Furthermore, if we rearrange the object of interest so that the origin O' of the CVP polar coordinate system of each rotated CVP becomes the same single origin of the CVP polar coordinate system, the relationship between the positions OBP1 to OBP5 of the object of interest in each CVP as seen from the origin of the CVP polar coordinate system will be as shown on the right side of the figure. In other words, the right side of the figure shows the object positions when each CVP is the origin and the median plane is the positive Y-axis direction.
[0282] In each CVP's polar coordinate system, there is a constraint that the direction of the median plane is toward the target position TP, so the positional relationship shown on the right side of the figure can be easily determined.
[0283] Furthermore, in the figure, on the right side, vectors starting from the position of the CVP, i.e., the origin, and ending at the positions OBP1 through OBP5 of the object of interest, are denoted as vectors V41 through V45. These vectors V41 through V45 correspond to vectors V11 through V13 shown in Figure 20.
[0284] Therefore, as shown in Figure 23, by using the contribution rate of each vector (CVP) as a weight and combining vectors V41 to V45, a vector V51 is obtained that indicates the position of the object of interest as seen from the listening position LP21. This vector V51 corresponds to vector V21 shown in Figure 21. Note that in Figure 23, the contribution rate of each CVP is set to 1 for simplicity of explanation.
[0285] Furthermore, the contribution rate of each CVP during vector synthesis may be determined, for example, by the ratio of the distance from the listening position to the CVP in the free viewpoint space (common absolute coordinate system).
[0286] Specifically, as shown in Figure 24, for example, the listening position indicated by the listener position information is position F, and the positions of the three CVPs used in the interpolation process are positions A to C. Furthermore, the absolute coordinates of position F in the common absolute coordinate system are (xf, yf, zf), and the absolute coordinates of positions A, B, and C in the common absolute coordinate system are (xa, ya, za), (xb, yb, zb), and (xc, yc, zc). The absolute coordinates indicating the position of each CVP in the common absolute coordinate system can be obtained from the CVP position information included in the configuration information.
[0287] At this time, the position calculation unit 114 calculates the ratio (distance ratio) of the distance AF from position F to position A, the distance BF from position F to position B, and the distance CF from position F to position C, and takes the reciprocal of this distance ratio as the ratio (dependence ratio) of the contribution rates of the CVPs at each position.
[0288] In other words, the position calculation unit 114 calculates the following equation (4) by setting AF:BF:CF=a:b:c and using dp(AF), dp(BF), and dp(CF) as the dependencies of each CVP located at positions A to C on the listening position (listener reference object position information).
[0289]
number
[0290] However, a, b, and c in equation (4) are as shown in equation (5).
[0291]
number
[0292] Furthermore, the position calculation unit 114 normalizes the dependencies dp(AF), dp(BF), and dp(CF) shown in equation (4) by calculating the following equation (6), and obtains the normalized dependencies ndp(AF), ndp(BF), and ndp(CF) as the final contribution rates. Note that a, b, and c in equation (6) are also obtained by equation (5).
[0293]
number
[0294] The contribution rates ndp(AF) to ndp(CF) obtained in this way correspond to the contribution rates dep[0] to dep[2] in equation (3), and the shorter the distance from the listening position to the CVP, the closer the contribution rate of that CVP will be to 1. Note that the contribution rate of each CVP is not limited to the example above and can be determined by any other method.
[0295] The position calculation unit 114 calculates the contribution rate of each CVP by determining the ratio of the distance from the listening position to the CVP based on the listener position information and the CVP position information.
[0296] To summarize the above, first, the position calculation unit 114 selects the CVP to be used for interpolation processing based on the listener position information and the CVP position information included in the configuration information. The CVP used for interpolation processing may be a portion of all CVPs located around the listening position, or all CVPs may be used for interpolation processing.
[0297] The position calculation unit 114 calculates the object's 3D position vector for each selected CVP based on the object's position information.
[0298] For example, if the 3D position vector of the j-th object as seen from the i-th CVPi is (Obj_vector_x[i][j],Obj_vector_y[i][j],Obj_vector_z[i][j]), then the 3D position vector of the object can be obtained by calculating the following equation (7).
[0299] Here, the polar coordinates represented by the object position information of the j-th object as seen from the i-th CVPi are assumed to be (Azi[i][j],Ele[i][j],rad[i][j]).
[0300]
number
[0301] Equation (7) is similar to equation (2) described above.
[0302] Next, the position calculation unit 114 performs calculations similar to those in equations (4) to (6) above, based on the listener position information and the CVP position information of each CVPi included in the configuration information, to determine the contribution rate dp(i) of each CVPi, which will be the weighting coefficient during interpolation. The contribution rate dp(i) is a weighting coefficient determined by the ratio of the distances from the listening position to the CVPi, or more specifically, the reciprocal ratio of the distances.
[0303] Furthermore, the position calculation unit 114 calculates the following equation (8) based on the object's 3D position vector obtained by the calculation of equation (7), the contribution rate dp(i) of each CVPi, and the gain Obj_gain[i][j] of the j-th object as seen from the CVPi. This provides the listener-referenced object position information (Intp_x(j), Intp_y(j), Intp_z(j)) and the listener-referenced gain Intp_gain(j) for the j-th object.
[0304]
number
[0305] Equation (8) calculates the weighted vector sum. That is, the sum of the 3D object position vectors of each CVPi multiplied by the contribution rate dp(i) is obtained as the listener-referenced object position information, and the sum of the gains of each CVPi multiplied by the contribution rate dp(i) is obtained as the listener-referenced gain. Equation (8) is the same as equation (3) described above.
[0306] Furthermore, the listener-referenced object position information obtained by equation (8) is in absolute coordinates of an absolute coordinate system where the listening position is the origin, and the direction from the listening position to the target position TP is the positive direction of the y-axis, i.e., the direction of the median plane.
[0307] However, since the rendering processing unit 115 performs rendering in a polar coordinate system, listener-referenced object position information in polar coordinate representation is required.
[0308] Therefore, the position calculation unit 114 converts the listener-referenced object position information (Intp_x(j),Intp_y(j),Intp_z(j)) in absolute coordinate representation obtained by equation (8) into listener-referenced object position information (Intp_azi(j),Intp_ele(j),Intp_rad(j)) in polar coordinate representation by calculating the following equation (9).
[0309]
number
[0310] The position calculation unit 114 outputs the listener-referenced object position information (Intp_azi(j), Intp_ele(j), Intp_rad(j)) obtained in this manner to the rendering processing unit 115 as the final listener-referenced object position information.
[0311] Furthermore, the listener-referenced object position information obtained by equation (9) is a polar coordinate system in which the listening position is the origin and the direction from the listening position to the target position TP is the positive direction of the y-axis, i.e., the direction of the median plane.
[0312] However, a listener actually at the listening position is not necessarily facing the direction of the target position TP. Therefore, if listener orientation information can be obtained by the listener position information acquisition unit 111, the listener reference object position information in polar coordinate representation obtained by equation (9) may be further subjected to a coordinate system rotation process to obtain the final listener reference object position information.
[0313] In this case, for example, the position calculation unit 114 rotates the position of the object as seen from the listening position by an angle determined by the position of the target position TP, which is known on the client side, the listener position information, and the listener orientation information. The rotation angle (correction amount) at this time is the angle between the direction from the listening position to the target position TP in the free viewpoint space and the direction of the listener's face indicated by the listener orientation information.
[0314] Furthermore, the target position TP in the common absolute coordinate system (free viewpoint space) can be calculated by the position calculation unit 114 from the CVP position information and CVP orientation information for multiple CVPs.
[0315] Through the above process, it is possible to obtain listener-referenced object position information in polar coordinates, which ultimately shows the object's position as seen from the listener's perspective.
[0316] Here, with reference to Figures 25 and 26, a specific example of calculating listener-referenced object position information according to the orientation of the listener's face will be explained. In Figures 25 and 26, corresponding parts are denoted by the same reference numerals, and their explanations will be omitted as appropriate.
[0317] For example, when viewing the XY plane of a free-viewpoint space (common absolute coordinate system), suppose the target position TP, each CVP, and the listening position LP41 are arranged as shown in Figure 25.
[0318] In this example, each circle without hatching (diagonal lines) represents a CVP, and the vertical angle indicated by the CVP Pitch information, which constitutes the CVP orientation information for each CVP, is assumed to be 0 degrees. In other words, the free viewpoint space is assumed to be essentially a two-dimensional plane. Also, here the target position TP is the position of the origin O in the common absolute coordinate system.
[0319] Furthermore, the line connecting the target position TP and the listening position LP41 is defined as line LN31, the line representing the direction of the listener's face as indicated by the listener orientation information is defined as line LN32, and the line passing through the listening position LP41 and parallel to the Y-axis of the common absolute coordinate system is defined as line LN33.
[0320] When the positive Y-axis direction is set to a horizontal angle of 0 degrees, the horizontal angle indicating the direction of the listener's face, that is, the angle between the line LN32 and the line LN33, is θcur_az. Also, when the positive Y-axis direction is set to a horizontal angle of 0 degrees, the horizontal angle indicating the direction of the target position TP as seen from an arbitrary listening position LP41, that is, the angle between the line LN31 and the line LN33, is θtp_az.
[0321] In this case, since the direction from an arbitrary listening position LP41 to the target position TP is defined as the direction of the midline, the angle between the straight line LN31 and the straight line LN32 is taken as the correction amount θcor_az, and the horizontal angle Intp_azi(j) of the listener reference object position information for each object should be corrected by the amount of this correction amount θcor_az. That is, the position calculation unit 114 adds the correction amount θcor_az to the horizontal angle Intp_azi(j) to obtain the final horizontal angle of the listener reference object position information.
[0322] The correction amount θcor_az can be obtained by calculating the following equation (10).
[0323]
number
[0324] Furthermore, for example, when viewing the free viewpoint space (common absolute coordinate system) from a direction parallel to the XY plane, suppose the target position TP and the listening position LP41 are arranged as shown in Figure 26.
[0325] Here, the line connecting the target position TP and the listening position LP41 is defined as line LN41, the line representing the direction of the listener's face as indicated by the listener orientation information is defined as line LN42, and the line passing through the listening position LP41 and parallel to the XY plane of the common absolute coordinate system is defined as line LN43.
[0326] Furthermore, the Z coordinate Rz represents the listening position LP41 in the common absolute coordinate system, which constitutes the listener position information, and the Z coordinate TPz represents the target position TP in the common absolute coordinate system.
[0327] In this case, the absolute value of the vertical angle (elevation angle) of the target position TP as seen from the listening position LP41 in the free viewpoint space is the angle θtp_el between the line LN41 and the line LN43.
[0328] Furthermore, the vertical angle (elevation angle) indicating the direction of the listener's face in the free viewpoint space is the angle θcur_el between the horizontal line LN43 and the line LN42 indicating the direction of the listener's face. In this case, when the listener is looking above the horizontal line, the angle θcur_el is a positive value, and when the listener is looking below the horizontal line, the angle θcur_el is a negative value.
[0329] In this example, the angle between the straight line LN41 and the straight line LN42 is used as the correction amount θcor_el, and the vertical angle Intp_ele(j) of the listener reference object position information for each object should be corrected by the amount of this correction amount θcor_el. That is, the position calculation unit 114 adds the correction amount θcor_el to the vertical angle Intp_ele(j) to obtain the final vertical angle of the listener reference object position information.
[0330] The correction amount θcor_el can be obtained by calculating the following equation (11).
[0331]
number
[0332] In the above, we have described an example using vector synthesis as an interpolation process. However, it is also possible to obtain listener-referenced object position information by using Ceva's theorem and interpolating with CVPs around the listening position.
[0333] For example, in interpolation using Ceva's theorem, a triangle is formed by three CVPs surrounding the listening position, and then Ceva's theorem is used to map this triangle to the object positions corresponding to those three CVPs, thereby realizing the interpolation process.
[0334] In this case, interpolation cannot be performed if the listening position lies outside the CVP triangle. However, the vector synthesis method described above can obtain listener-referenced object position information even if the listening position is outside the region enclosed by the CVP. Furthermore, the vector synthesis method allows for obtaining listener-referenced object position information easily with less processing effort.
[0335] <Second Embodiment> <About the CVP Group> Incidentally, this technology allows for the reproduction of the sound field from the producer's intended perspective for the listener within a closed space, such as a live music venue, while also allowing the listener to freely move their position. Furthermore, considering the case where the listener moves outside the live music venue, the difference in sound fields inside and outside the venue is significant, and much of the sound inside the venue should not be audible outside.
[0336] However, in the method of the first embodiment described above, even if sounds from outside the live venue are set, when reproducing the sound field at an arbitrary location outside the live venue, sounds from inside the live venue that should not be mixed in may be heard due to the influence of combination patterns of object position information inside the live venue.
[0337] Therefore, for example, three areas may be established: inside the live venue, outside the live venue, and the area transitioning from inside to outside the live venue, and the CVP used in each area may be separated. In such a case, the area in which the listener is currently located is selected according to the listener's position, and only the CVP belonging to that area is used to obtain the listener's reference object position information. The number of areas to be divided may be set arbitrarily by the creator, or it may be set according to each live venue.
[0338] By doing so, it is possible to avoid sound interference between inside and outside the live venue, while achieving audio playback of free-viewpoint content with appropriate listener-referenced object position information and listener-referenced gain.
[0339] There are several ways to define the regions that divide the free-viewpoint space, with concentric circles or polygons from a predetermined central coordinate being common examples. Alternatively, any number of small regions of various shapes may be created.
[0340] The following section provides a detailed explanation of an example in which a free viewpoint space is divided into multiple regions and a CVP (Coefficient of Motion) is selected for interpolation.
[0341] For example, as shown in Figure 27, suppose the free viewpoint space is divided into three group regions R11 to R13. In Figure 27, each small circle represents a CVP (Center of View).
[0342] Group region R11 is a circular region (space), group region R12 is an annular region surrounding the outside of group region R11, and group region R13 is an annular region surrounding the outside of group region R12.
[0343] In this example, group region R12 is defined as the transitional region between group region R11 and group region R13. Therefore, for example, the region (space) inside the live venue can be designated as group region R11, the region outside the live venue as group region R13, and the region between the inside and outside of the live venue as group region R12. Note that each group region is set up so that there are no overlapping parts (regions) between them.
[0344] In this example, the CVPs used for interpolation are grouped according to the listener's position. In other words, the creator specifies the range of the group area, and the CVPs are grouped according to that range.
[0345] For example, CVPs placed in the free viewpoint space are grouped so that they belong to at least one of the following: CVP group GP1 corresponding to group region R11, CVP group GP2 corresponding to group region R12, and CVP group GP3 corresponding to group region R13. In this case, a single CVP can belong to multiple different CVP groups.
[0346] Specifically, CVPs located within group area R11 belong to CVP group GP1, CVPs located within group area R12 belong to CVP group GP2, and CVPs located within group area R13 belong to CVP group GP3.
[0347] Therefore, for example, a CVP located at position P61 within group region R11 belongs to CVP group GP1, and a CVP located at position P62, which is the boundary position between group region R11 and group region R12, belongs to both CVP group GP1 and CVP group GP2.
[0348] Furthermore, the CVP located at position P63, which is the boundary between group region R12 and group region R13, belongs to CVP group GP2 and CVP group GP3, while the CVP located at position P64 within group region R13 belongs to CVP group GP3.
[0349] Applying this type of grouping to a live venue as a free-viewpoint space, as shown in Figure 11, results in something like Figure 28.
[0350] In this example, for instance, CVP1 through CVP7, represented by black circles in the diagram, are included in the group area (group space) corresponding to the inside of the live venue, while CVPs, represented by white circles in the diagram, are included in the group area corresponding to the outside of the live venue.
[0351] Furthermore, in the configuration information, for example, links can be established (associated) between CVPs within a specific live venue and CVPs outside the live venue. In such a case, for example, when a listener is between those two CVPs, the listener-referenced object position information can be obtained by vector synthesis using those two CVPs. In this case, the gain of a predetermined object to be muted may be set to 0.
[0352] The example in Figure 28 will be explained in more detail with reference to Figures 29 and 30. Note that in Figures 29 and 30, corresponding parts are denoted by the same reference numerals, and their explanations will be omitted as appropriate.
[0353] For example, as shown in Figure 29, let's assume that the area within the live venue is a circular region centered on the origin O of the common absolute coordinate system in the free viewpoint space.
[0354] In particular, the circular area R31, drawn with a dotted line, represents the area inside the live venue, while the area outside of area R31 represents the area outside the live venue.
[0355] Additionally, CVP1 through CVP15 are positioned inside the live venue, and CVP16 through CVP23 are positioned outside the venue.
[0356] In this case, for example, the region inside a circle with radius Area1_border centered at the origin O is considered one group region R41, and the group of CVPs consisting of CVP1 to CVP15 contained within that group region R41 is considered the CVP group GPI corresponding to group region R41. Group region R41 is the region within the live venue.
[0357] Furthermore, as shown on the right side of the figure, the region between the boundary of the circle with radius Area1_border centered at the origin O and the boundary of the circle with radius Area2_border centered at the origin O is defined as group region R42. This group region R42 is the transition region between the inside and outside of the live venue.
[0358] The group of CVPs consisting of CVP8 to CVP23 contained within group area R42 is considered the CVP group GPM corresponding to group area R42.
[0359] In this particular example, CVP8 through CVP15 are located at the boundary between group region R41 and group region R42, and therefore belong to both CVP group GPI and CVP group GPM.
[0360] Furthermore, as shown in Figure 30, the area outside the circle, including the boundary of a circle with radius Area2_border centered at the origin O, is defined as group region R43. Group region R43 is the area outside the live venue.
[0361] The group of CVPs consisting of CVP16 to CVP23 located within group area R43 is considered the CVP group GPO corresponding to group area R43. In this particular example, since CVP16 to CVP23 are located at the boundary between group area R42 and group area R43, these CVP16 to CVP23 belong to both the CVP group GPM and the CVP group GPO.
[0362] When the group area and CVP group are set as described above, the position calculation unit 114 performs interpolation processing as follows to obtain listener reference object position information and listener reference gain.
[0363] In other words, as shown on the left side of Figure 29, for example, when the listening position is within the group region R41, the position calculation unit 114 performs interpolation using some or all of the CVP1 to CVP15 belonging to the CVP group GPI.
[0364] Furthermore, as shown on the right side of Figure 29, for example, when the listening position is within the group region R42, the position calculation unit 114 performs interpolation processing using some or all of the CVP8 to CVP23 belonging to the CVP group GPM.
[0365] Furthermore, as shown in Figure 30, for example, when the listening position is within the group area R43, the position calculation unit 114 performs interpolation processing using some or all of the CVP16 to CVP23 belonging to the CVP group GPO.
[0366] In the above, we have described an example in which group regions are defined in a concentric circle. However, as shown in Figure 31, for example, circular regions R71 and R72 may be defined, each having a different center position and overlapping transition regions.
[0367] In this example, region R71 contains CVP1 through CVP7, and region R72 contains CVP5, CVP6, and CVP8 through CVP12. Additionally, the transition region where regions R71 and R72 overlap contains CVP5 and CVP6.
[0368] Here, let's assume that the region excluding the transition region within region R71, the region excluding the transition region within region R72, and the transition region are defined as group regions.
[0369] In such cases, for example, when the listening position is within a region R71 excluding the transition region, interpolation processing is performed using some or all of CVP1 to CVP7.
[0370] Furthermore, for example, when the listening position is within the transition region, interpolation processing is performed using CVP5 and CVP6. Additionally, when the listening position is within a region other than the transition region in region R72, interpolation processing is performed using some or all of CVP5, CVP6, and CVP8 to CVP12.
[0371] As described above, when the creator can specify a group area, i.e., a CVP group, the format of the configuration information will be as shown in Figure 32, for example.
[0372] In the example shown in Figure 32, the format is basically the same as in Figure 7, and the configuration information includes the frame length index "FrameLengthIndex", the number of objects "NumOfObjects", the number of CVPs "NumOfControlViewpoints", the number of metadata sets "NumOfObjectMetaSets", the CVP information "ControlViewpointInfo(i)", and the coordinate mode information "CoordinateMode[i][j]".
[0373] Furthermore, the configuration information shown in Figure 32 also includes the CVP group information existence flag "cvp_group_present".
[0374] The CVP group information existence flag "cvp_group_present" indicates whether or not the CVP group information "CvpGroupInfo2D()", which is information about CVP groups, is included in the configuration information.
[0375] For example, if the value of the CVP group information existence flag is "1", the configuration information contains the CVP group information "CvpGroupInfo2D()", and if the value of the CVP group information existence flag is "0", the configuration information does not contain the CVP group information "CvpGroupInfo2D()".
[0376] Furthermore, the format of the CVP group information "CvpGroupInfo2D()" included in the configuration information is as shown in Figure 33, for example. For simplicity, this explanation uses the example of a free-viewpoint space being a two-dimensional region (space), but it is of course possible to extend the CVP group information shown in Figure 33 to a three-dimensional region (space).
[0377] In this example, "numOfCVPGroup" indicates the number of CVP groups, i.e., the number of CVP groups. The CVP group information contains information about the CVP groups described below, corresponding to the number of CVP groups.
[0378] "vertex_idx" indicates the vertex index. The vertex index is index information that shows the number of vertices in the group region corresponding to the CVP group.
[0379] For example, if the vertex index value is between 0 and 5, the group region is defined as a polygonal region with a number of vertices equal to the vertex index value plus 3. Also, if the vertex index value is 255, the group region is defined as a circular region.
[0380] When the vertex index value is 255, that is, when the shape type of the group region is a circle, the CVP group information stores the normalized X coordinate "center_x[i]", the normalized Y coordinate "center_y[i]", and the normalized radius "radius[i]" as information to identify the circular group region (the boundary of the group region).
[0381] For example, the normalized X coordinate "center_x[i]" and normalized Y coordinate "center_y[i]" represent the X and Y coordinates of the center of the circle that constitutes the group region in the common absolute coordinate system (free viewpoint space), and the normalized radius "radius[i]" is the radius of the circle that constitutes the group region. This allows us to identify which regions constitute the group region in the free viewpoint space.
[0382] Furthermore, if the vertex index value is between 0 and 5, that is, if the group region is a polygonal region, the CVP group information stores the normalized X coordinate "border_pos_x[j]" and normalized Y coordinate "border_pos_y[j]" for each vertex of the group region.
[0383] For example, the normalized X coordinate "border_pos_x[j]" and normalized Y coordinate "border_pos_y[j]" represent the X and Y coordinates of the j-th vertex of a polygonal region, which is a group region in the common absolute coordinate system (free viewpoint space).
[0384] From the normalized X and Y coordinates of each of these vertices, it is possible to identify the polygonal region as a group area in free-viewpoint space.
[0385] Furthermore, the CVP group information includes "numOfCVP_ingroup[i]", which indicates the number of CVPs belonging to the CVP group, and also stores the number of "CvpIndex_ingroup[i][j]" within the group, as indicated by the number of CVPs within the group. The "CvpIndex_ingroup[i][j]" within the group is index information that identifies the j-th CVP belonging to the i-th CVP group.
[0386] For example, the value of the group-specific CVP index representing a given CVP can be the same as the value of the CVP index representing that given CVP included in the CVP information.
[0387] As described above, CVP group information includes the number of CVP groups, a vertex index indicating the shape type of the group region, information for identifying the group region, information on the number of CVPs within the group, and the CVP index within the group. In particular, the information for identifying the group region can be said to be information for identifying the boundaries of the group region.
[0388] Even when configuration information in the format shown in Figure 32 is generated, the information processing device 11 basically performs the content creation process described with reference to Figure 14.
[0389] However, in this case, the creator can perform operations to specify the group area or the CVPs belonging to the CVP group at any time, for example, in step S11 or step S16.
[0390] Then, the control unit 26 determines (sets) the CVPs belonging to the group area and CVP group in response to the creator's operation. In step S23, the control unit 26 generates configuration information shown in Figure 32, which appropriately includes the CVP group information shown in Figure 33, based on the setting results of the CVPs belonging to the group area and CVP group.
[0391] <Explanation of the playback audio data generation process> Furthermore, if the configuration information is in the format shown in Figure 32, the client 101 performs, for example, the playback audio data generation process shown in Figure 34.
[0392] The following describes the playback audio data generation process by client 101, referring to the flowchart in Figure 34.
[0393] Note that the processing in steps S121 to S123 is the same as the processing in steps S81 to S83 in Figure 18, so its explanation will be omitted.
[0394] In step S124, the position calculation unit 114 identifies the CVP group corresponding to the group area containing the listening position, based on the listener position information and configuration information.
[0395] For example, the position calculation unit 114 identifies the group region including the listening position (hereinafter also referred to as the target group region) based on the normalized X coordinates and normalized Y coordinates, which are information for identifying the regions that constitute each group region and are included in the CVP group information within the configuration information.
[0396] Furthermore, when the listening position is at the boundary of multiple group areas, those multiple group areas are considered the target group area.
[0397] Once the target group area is identified in this way, the CVP group corresponding to that target group area is also identified.
[0398] In step S125, the position calculation unit 114 selects each CVP belonging to the identified CVP group as a target CVP and obtains the object metadata set associated with those target CVPs.
[0399] For example, the position calculation unit 114 identifies the CVP belonging to the CVP group, i.e., the target CVP, by reading the in-group CVP index of the CVP group corresponding to the target group area from the CVP group information.
[0400] Furthermore, the position calculation unit 114 reads the metadata set index for the target CVP from the CVP information to identify the object metadata set associated with each target CVP, and then reads the identified object metadata set.
[0401] After the processing in step S125 is completed, the processing in steps S126 to S128 is performed, and the playback audio data generation process is completed. However, these processes are the same as the processing in steps S84 to S86 in Figure 18, so their explanation will be omitted.
[0402] However, in step S126, interpolation processing is performed using some or all of the target CVPs identified in steps S124 and S125. That is, the CVP position information and object position information of the target CVPs are used to calculate the listener-referenced object position information and listener-referenced gain.
[0403] This makes it possible to obtain playback audio data that reproduces an appropriate sound field according to the listener's location, such as whether the listener is inside or outside the live venue.
[0404] As described above, client 101 performs interpolation using an appropriate CVP based on the listener location information, configuration information, and object metadata set to calculate the listener reference gain and listener reference object location information at the listening location.
[0405] By doing so, it becomes possible to achieve musical content playback that is in line with the content creator's intentions, and to fully convey the appeal of the content to the listener.
[0406] <Third Embodiment> <Regarding object position information and gain interpolation processing> Incidentally, within the free viewpoint space, there are multiple CVPs (Common Viewpoints) that have been pre-set by the content creator. In the above, as a specific example, we explained an example in which the reciprocal ratio of the distance from the current arbitrary position of the listener (listening position) to each CVP is used in the interpolation process to determine the listener-referenced object position information and listener-referenced gain.
[0407] In such an example, suppose the object's gain is set to a large value for a CVP that is far from the listening position.
[0408] In this case, even though the actual distance is great, the gain of the object located far from the listening position can sometimes fail to minimize the listener-referenced gain, that is, the perceived auditory influence on the sound of the object heard by the listener. As a result, the sound image movement of the object presented to the listener becomes unnatural, and the sound quality of the content deteriorates.
[0409] In the following, we will also refer to cases where unnatural sound image movement occurs due to the relationship between the listening position and the CVP (Chip Pointer) position as Case A.
[0410] Furthermore, when a content creator sets the gain of an object in a particular CVP to 0, they may become unaware of the object's position, resulting in the object's position being left unaddressed. In other words, if the content creator sets the gain of an object to 0, the object's position may not be set, and as a result, the object's position information may be set to an inappropriate value.
[0411] However, such neglected object position information is also used in the interpolation process to determine listener-referenced object position information. As a result, the neglected and inappropriate object position information can sometimes cause the object's position as seen by the listener to be in a location unintended by the content creator.
[0412] In the following, we will also refer to the case where the object's position as seen by the listener, as indicated by the listener-referenced object position information, becomes an unintended position due to the influence of the left-behind object's position as Case B.
[0413] If we can suppress the occurrence of cases A and B described above, we can achieve higher quality content playback that is in line with the content creator's intentions.
[0414] Therefore, in the third embodiment, the occurrence of such cases A and B can be suppressed.
[0415] For example, in case A, where a high gain at a CVP far from the current listening position affects the listener's reference gain, a sensitivity coefficient is applied that adjusts the sensitivity using the Nth power of the distance across all CVPs.
[0416] This allows for weighting of the dependency (contribution rate) of each CVP in the interpolation process for determining the listener-referenced gain. Hereafter, the method that suppresses the occurrence of Case A by applying a sensitivity coefficient will be specifically referred to as Method SLA1.
[0417] By appropriately controlling the sensitivity coefficient, it becomes possible to further reduce the influence of CVPs located far from the current listening position, thereby suppressing the occurrence of Case A. This prevents unnatural gain fluctuations from occurring, for example, when the listener moves between CVPs.
[0418] The sensitivity coefficient value, i.e., the value of N, is a float value or the like. The content creator may choose to have the sensitivity coefficient value for each CVP described as a default value in the configuration information and transmitted to client 101, or the sensitivity coefficient value may be set on the listener side.
[0419] Furthermore, while a common sensitivity coefficient may be used for all objects for each CVP, the sensitivity coefficient may also be set individually for each object according to the content creator's intentions. In addition, for each group consisting of one or more CVPs, a common sensitivity coefficient or an individual sensitivity coefficient for each object may be set.
[0420] On the other hand, in case B, where including object position information of objects in a neglected state, such as those with a gain of 0, as an element in the vector sum during interpolation would result in listener-referenced object position information that does not conform to the content creator's intentions, the gain is added to the contributing material, or only objects with a gain greater than 0 are used. In other words, the occurrence of case B is suppressed by either method SLB1 or method SLB2 shown below.
[0421] In the SLB1 method, objects whose gain in CVP is below a predetermined threshold are considered objects with a gain of 0 (hereinafter also referred to as muted objects). For CVPs designated as muted objects, the object position information in that CVP is not used in the interpolation process. In other words, it is excluded from the CVPs used in the interpolation process.
[0422] The SLB2 method uses a Mute flag, which is specified by the content creator or others, to indicate whether an object's gain is 0 or not, i.e., whether it is a muted object.
[0423] Specifically, for CVPs that are muted by the Mute flag, the object position information in those CVPs will not be used in the interpolation process. In other words, CVPs corresponding to objects that are known in advance not to be used are excluded from the CVPs used in the interpolation process.
[0424] Methods like SLB1 and SLB2 allow for correct interpolation using only the CVP of objects whose gain is not considered zero, by excluding the CVP of objects whose gain is considered zero from the processing.
[0425] In particular, the SLB2 method can avoid the process of checking whether the gain can be considered 0 for each object in all CVPs, which is performed frame by frame in the SLB1 method, thereby further reducing the processing load.
[0426] <CVP placement pattern PTT1> Next, we will describe examples of actual CVP placement patterns in which the above-mentioned Cases A and B occur.
[0427] First, Figure 35 shows the first CVP placement pattern (hereinafter also referred to as CVP placement pattern PTT1). In this example, objects are placed in front of each CVP.
[0428] In Figure 35, each circle with a number represents a single CVP, and the number within the circle representing a CVP specifically indicates which CVP it is. Hereafter, the k-th CVP, where the number k (where k=1,2,...,6) is indicated, will also be referred to as CVPk.
[0429] In this example, we will focus on a single object, OBJ71, located in free viewpoint space.
[0430] For example, in CVP1 to CVP6, which are placed in a free viewpoint space, the object OBJ71 has its own defined position information and gain as seen from each CVP.
[0431] Now, for a predetermined listening position LP71, we consider performing interpolation processing using equations (7) to (11) described above, based on the object position information and gain at CVP1 to CVP6, to obtain listener-referenced object position information and listener-referenced gain.
[0432] In such cases, for example, when the gain of object OBJ71 is the same non-zero value in CVP1 through CVP6, neither Case A nor Case B described above occurs.
[0433] In contrast, if, for example, the gain of object OBJ71 in CVP1 to CVP3 is greater than the gain of object OBJ71 in CVP5 and CVP6, then case A may occur.
[0434] This is because the distance from the listening position LP71 to CVP1 through CVP3 is long (far), so the apportionment ratio of those CVPs is low. In other words, the contribution rate dp(i) obtained by calculations similar to those in equations (4) through (6) is small, but the original gain is large, so it greatly affects the listener's reference gain.
[0435] Furthermore, for example, in CVP6, object OBJ71 is designated as a mute object, but if the horizontal angle Azimuth of the object's position information is -180 degrees, which is significantly different from the angle Azimuth in other CVPs, then Case B occurs. This is because, when calculating the listener-referenced object position information, the influence of the object position information in CVP6, which is closest to the listening position LP71, becomes significant.
[0436] Here, Figure 36 shows an example of the arrangement of CVPs, etc., in a common absolute coordinate system when the free viewpoint space is essentially a two-dimensional plane and the positional relationship of the listener and CVP is as shown in Figure 35 (CVP arrangement pattern PTT1). Note that in Figure 36, the same reference numerals are used for parts corresponding to the case in Figure 35, and their explanations are omitted as appropriate.
[0437] In Figure 36, the horizontal and vertical axes represent the X and Y axes in the common absolute coordinate system. If we represent the position (coordinates) in the common absolute coordinate system as (x,y), then, for example, the listening position LP71 is represented by (0,-0.8).
[0438] Figure 37 shows examples of the object position and gain of object OBJ71 in each CVP when Case A and Case B occur, as well as the listener-referenced object position information and listener-referenced gain of object OBJ71, for such arrangements in a common absolute coordinate system. In the example in Figure 37, the contribution rate dp(i) is obtained by the same calculation as equations (4) to (6) described above.
[0439] In Figure 37, the column labeled "CaseA" shows an example of the object position and gain of object OBJ71 when Case A occurs.
[0440] Specifically, "azi(0)" represents the angle Azimuth as object position information for object OBJ71, and "Gain(0)" represents the gain of object OBJ71.
[0441] In this example, the gain (Gain(0)) at CVP1 through CVP3 is "1", while the gain (Gain(0)) at CVP5 and CVP6 is "0.2". In other words, the gain at CVP1 through CVP3, which are farther from the listening position LP71, is greater than the gain at CVP5 and CVP6, which are closer to the listening position LP71. Therefore, the listener-referenced gain (Gain(0)) at the listening position LP71 is "0.37501".
[0442] In this example, the listening position LP71 is located between CVP5 and CVP6, which have a gain of 0.2. Therefore, the listener reference gain should ideally be close to the gain of "0.2" at CVP5 and CVP6. However, in reality, due to the influence of the higher gains of CVP1 through CVP3, it becomes a value of "0.37501", which is greater than "0.2".
[0443] Additionally, the column labeled "CaseB" shows an example of the object position and gain of object OBJ71 when Case B occurs.
[0444] Specifically, "azi(1)" represents the angle Azimuth as object position information for object OBJ71, and "Gain(1)" represents the gain of object OBJ71.
[0445] In this example, the gain (Gain(1)) and angle Azimuth (azi(1)) in CVP1 through CVP5 are "1" and "0", respectively. In other words, object OBJ71 is not a muted object in CVP1 through CVP5.
[0446] In contrast, the gain (Gain(1)) and angle Azimuth (azi(1)) in CVP6 are "0" and "120," respectively. In other words, in CVP6, object OBJ71 is a muted object.
[0447] Furthermore, the angle Azimuth(azi(1)) as the listener-referenced object position information at the listening position LP71 is "67.87193".
[0448] In this example, the gain (Gain(1)) at CVP6 is "0", so the angle Azimuth (azi(1)) of "120" at CVP6 should be ignored. However, in reality, the angle Azimuth at CVP6 is used to calculate the listener-referenced object position information. As a result, the angle Azimuth (azi(1)) at listening position LP71 becomes "67.87193", which is significantly larger than "0".
[0449] <CVP placement pattern PTT2> Next, Figure 38 shows the second CVP placement pattern (hereinafter also referred to as CVP placement pattern PTT2). In this example, each CVP is placed to surround the object.
[0450] In Figure 38, as in Figure 35, each circle with a numerical value represents a single CVP, and the k-th CVP, where k = 1, 2, ..., 8 is indicated, will be specifically referred to as CVPk.
[0451] In this example, we focus on one object, OBJ81, and consider performing interpolation processing for the listening position LP81 based on the object position information and gain at CVP1 to CVP8 using equations (7) to (11) described above.
[0452] In such cases, for example, when the gain of object OBJ81 is the same non-zero value in CVP1 through CVP8, neither Case A nor Case B described above occurs.
[0453] In contrast, if, for example, the gain of object OBJ81 in CVP1, CVP2, CVP6, and CVP8 is greater than the gain of object OBJ81 in CVP3 and CVP4, then case A may occur.
[0454] This is because the distance from the listening position LP81 to CVP1, CVP2, CVP6, and CVP8 is long (far), so the apportionment ratio of those CVPs is low, but their original gain is high, which greatly affects the listener's reference gain.
[0455] Furthermore, for example, in CVP3, object OBJ81 is designated as a mute object, but if the horizontal angle Azimuth of the object's position information is -180 degrees, which is significantly different from the angle Azimuth in other CVPs, then Case B occurs. This is because, when calculating the listener-referenced object position information, the influence of the object position information in CVP3, which is closest to the listening position LP81, becomes significant.
[0456] Here, Figure 39 shows an example of the arrangement of CVPs, etc., in a common absolute coordinate system when the free viewpoint space is essentially a two-dimensional plane and the positional relationship of the listener and CVP is as shown in Figure 38 (CVP arrangement pattern PTT2). Note that in Figure 39, the parts corresponding to the case in Figure 38 are denoted by the same reference numerals, and their explanations are omitted as appropriate.
[0457] In Figure 39, the horizontal and vertical axes represent the X and Y axes in the common absolute coordinate system. If we represent the position (coordinates) in the common absolute coordinate system as (x,y), then, for example, the listening position LP81 is represented as (-0.1768, 0.176777).
[0458] Figure 40 shows examples of the object position and gain of object OBJ81 in each CVP when Case A and Case B occur, as well as the listener-referenced object position information and listener-referenced gain of object OBJ81, for such arrangements in a common absolute coordinate system. In the example in Figure 40, the contribution rate dp(i) is obtained by the same calculation as equations (4) to (6) described above.
[0459] In Figure 40, the column labeled "CaseA" shows an example of the object position and gain of object OBJ81 when Case A occurs.
[0460] Specifically, "azi(0)" represents the angle Azimuth as object position information for object OBJ81, and "Gain(0)" represents the gain of object OBJ81.
[0461] In this example, the gain (Gain(0)) at CVP1, CVP2, CVP6, and CVP8 is "1", while the gain (Gain(0)) at CVP3 and CVP4 is "0.2". Therefore, the listener-referenced gain (Gain(0)) at listening position LP81 is "0.501194".
[0462] In this example, the listening position LP81 is located between CVP3 and CVP4, which have a gain of 0.2. Therefore, the listener reference gain should ideally be close to the gain of "0.2" at CVP3 and CVP4. However, in reality, due to the influence of CVP1 and CVP2, which have higher gains, it becomes a value of "0.501194", which is greater than "0.2".
[0463] Additionally, the column labeled "CaseB" shows an example of the object position and gain of object OBJ81 when Case B occurs.
[0464] Specifically, "azi(1)" represents the angle Azimuth as object position information for object OBJ81, and "Gain(1)" represents the gain of object OBJ81.
[0465] In this example, the gain (Gain(1)) and angle Azimuth (azi(1)) in CVPs other than CVP3 are "1" and "0", respectively. In other words, in CVPs other than CVP3, object OBJ81 is not a muted object.
[0466] In contrast, the gain (Gain(1)) and angle Azimuth (azi(1)) in CVP3 are "0" and "120," respectively. In other words, in CVP3, object OBJ81 is a muted object.
[0467] Furthermore, the angle Azimuth(azi(1)) as the listener-referenced object position information at the listening position LP81 is "20.05743".
[0468] In this example, the gain (Gain(1)) at CVP3 is "0", so the angle Azimuth (azi(1)) of "120" at CVP3 should be ignored. However, in reality, the angle Azimuth at CVP3 is used to calculate the listener-referenced object position information. As a result, the angle Azimuth (azi(1)) at the listening position LP81 becomes "20.05743", which is significantly larger than "0".
[0469] In this embodiment, the occurrence of cases A and B described above is suppressed by methods SLA1, SLB1, and SLB2.
[0470] In method SLA1, the sensitivity coefficient is N, and the contribution rate dp(i) is calculated based on the reciprocal of the Nth power of the distance from the listening position to the CVP. In this case, for example, the content creator may specify an arbitrary positive real number as the sensitivity coefficient and store the sensitivity coefficient in the configuration information, or if the listener or others are permitted to change the sensitivity coefficient, the sensitivity coefficient may be set on the client 101 side.
[0471] Furthermore, in the SLB1 method, for each object in each frame, it is determined whether the object's gain in CVP is 0 or a value that can be considered 0. Then, CVP values for objects whose gain is 0 or a value that can be considered 0 are not used in the interpolation process. In other words, CVP is excluded from the vector addition operations in the interpolation process.
[0472] In the SLB2 method, a Mute flag is stored in the configuration information for each CVP, signaling whether it is a muted object. Objects with a Mute flag of 1, i.e., muted CVPs, are excluded from the vector sum calculation during the interpolation process.
[0473] In these methods, SLB1 and SLB2, CVP is selected for interpolation, as shown in Figure 41, for example.
[0474] In other words, for example, when methods SLB1 or SLB2 were not applied, the interpolation process to determine the listener-referenced object position information for the listening position LP91, as shown on the left side of the figure, used all of the CVP1 to CVP4 surrounding the listening position LP91. That is, the interpolation process was performed using the object position information of each CVP from CVP1 to CVP4.
[0475] In contrast, when applying methods SLB1 or SLB2, as shown on the right side of the figure, CVP4, which is designated as a muted object, is excluded from the interpolation process. That is, in the interpolation process to determine the listener-referenced object position information for listening position LP91, the object position information from the three CVP1 to CVP3, excluding CVP4, is used.
[0476] When method SLA1 and method SLB1 or method SLB2 are performed simultaneously, the interpolation process yields the results shown in Figures 42 and 43, respectively, in the examples shown in Figures 37 and 40. Note that explanations of the parts in Figures 42 and 43 that correspond to Figures 37 and 40 are omitted as appropriate.
[0477] Figure 42 shows an example of applying method SLA1 and method SLB1 or method SLB2 to the CVP arrangement pattern PTT1 shown in Figures 36 and 37.
[0478] In Figure 42, the section indicated by arrow Q71 shows the angle Azimuth and gain for each CVP for "Case A" and "Case B".
[0479] Furthermore, the section indicated by arrow Q72 shows the angle Azimuth and listener-referenced gain as listener-referenced object position information when the sensitivity coefficient value is changed for "Case A" and "Case B," and it can be seen that the occurrence of Case A and Case B is suppressed.
[0480] For example, in "Case A," if we set the sensitivity coefficient value to "3," that is, if we focus on the "1 / distance cube ratio" column, the listener-referenced gain (Gain(0)) at listening position LP71 is "0.205033."
[0481] In this example, the application of method SLA1 significantly reduces the influence of CVP1 to CVP3, which are located far from the listening position LP71, and the listener-referenced gain becomes an ideal value close to the gain of "0.2" at the nearby CVP5 and CVP6.
[0482] In other words, the listener's reference gain at the listening position LP71, which is located between CVP5 and CVP6, becomes close to the gain at CVP5 and CVP6, thereby suppressing the occurrence of unnatural sound image movement.
[0483] Furthermore, focusing on "Case B," the angle Azimuth(azi(1)) as listener-referenced object position information at listening position LP71 is "0" regardless of the sensitivity coefficient value.
[0484] In this example, the angle Azimuth(azi(1)) value of "120" at CVP6, where the gain is "0", is not used in the interpolation process due to the application of either method SLB1 or method SLB2. In other words, the angle Azimuth "120" at CVP6 is excluded from the interpolation process.
[0485] Therefore, the angle Azimuth (azi(1)) at the listening position LP71 is the same value "0" as the angle Azimuth at all CVPs that are not excluded from the target, indicating that appropriate listener-referenced object position information can be obtained.
[0486] Figure 43 shows an example of applying method SLA1 and method SLB1 or method SLB2 to the CVP placement pattern PTT2 shown in Figures 39 and 40.
[0487] In Figure 43, the section indicated by arrow Q81 shows the angle Azimuth and gain for each CVP for "Case A" and "Case B".
[0488] Furthermore, the section indicated by arrow Q82 shows the angle Azimuth and listener-referenced gain as listener-referenced object position information when the sensitivity coefficient value is changed for "Case A" and "Case B," and it can be seen that the occurrence of Case A and Case B is suppressed.
[0489] For example, in "Case A," if we consider the case where the sensitivity coefficient value is "3," the listener-referenced gain (Gain(0)) at the listening position LP81 is "0.25492."
[0490] In this example, the application of method SLA1 significantly reduces the influence of CVP1, CVP2, CVP6, and CVP8, which are located far from the listening position LP81, and the listener-referenced gain becomes an ideal value close to the gain of "0.2" at nearby CVP3 and CVP4.
[0491] In other words, by controlling the sensitivity coefficient, the listener-referenced gain at the listening position LP81, which is between CVP3 and CVP4, becomes close to the gain at CVP3 and CVP4, thereby suppressing the occurrence of unnatural sound image movement.
[0492] Furthermore, focusing on "Case B," the angle Azimuth(azi(1)) as listener-referenced object position information at listening position LP81 is "0" regardless of the sensitivity coefficient value.
[0493] In this example, the angle Azimuth(azi(1)) value of "120" at CVP3, where the gain is "0", is not used in the interpolation process when either method SLB1 or SLB2 is applied. In other words, the angle Azimuth "120" at CVP3 is excluded from the interpolation process.
[0494] Therefore, the angle Azimuth (azi(1)) at the listening position LP81 is the same value "0" as the angle Azimuth at all CVPs that are not excluded from the target, indicating that appropriate listener-referenced object position information can be obtained.
[0495] <Example of configuration information format> Furthermore, when applying the SLB2 method, the configuration information will store, for example, the configuration (information) shown in Figure 44.
[0496] Figure 44 shows an example of the format (syntax) of a portion of the configuration information when applying the SLB2 method.
[0497] More specifically, the configuration information includes the configuration shown in Figure 44, as well as the configuration shown in Figure 7. In other words, the configuration information includes the configuration shown in Figure 44 as part of the configuration shown in Figure 7. Alternatively, the configuration information shown in Figure 32 may include the configuration shown in Figure 44 as part of the configuration information shown in Figure 32.
[0498] In the example in Figure 44, "NumOfControlViewpoints" indicates the number of CVPs, i.e., the number of CVPs set by the creator, while "numOfObjs" indicates the number of objects.
[0499] The configuration information contains, for each CVP, a Mute flag "MuteObjIdx[i][j]" corresponding to the number of objects, for each CVP and object combination.
[0500] The Mute flag "MuteObjIdx[i][j]" is a flag that indicates whether the j-th object is a muted object (is muted) when viewed from the i-th CVP, that is, when the listening position (viewpoint) is at the i-th CVP. Specifically, a value of "0" for the Mute flag "MuteObjIdx[i][j]" indicates that the object is not a muted object, and a value of "1" for the Mute flag "MuteObjIdx[i][j]" indicates that the object is a muted object, that is, is muted.
[0501] In this example, we described how the Mute flag is stored in the configuration information as mute information to identify objects designated as muted in CVP. However, this is not the only way; for example, "MuteObjIdx[i][j]" could be used as index information to indicate an object designated as muted.
[0502] In such cases, it is not necessary to store "MuteObjIdx[i][j]" for all objects in the configuration information; it is sufficient to store "MuteObjIdx[i][j]" only for objects that have been muted. In this example as well, client 101 can determine whether each object is muted in CVP by referring to "MuteObjIdx[i][j]".
[0503] <Explanation of the calculation process for the contribution coefficient> Next, we will describe the operation of the information processing device 11 and the client 101 when applying method SLA1 and either method SLB1 or method SLB2.
[0504] For example, when the method SLB2 is applied, the information processing apparatus 11 performs the content production process described with reference to FIG. 14.
[0505] However, in this case, for example, the control unit 26 receives a designation operation as to whether to set an object as a mute object in the CVP at an arbitrary timing, and in step S23, generates configuration information including a Mute flag having a value corresponding to the designation operation.
[0506] Also, for example, when a sensitivity coefficient is stored in the configuration information, the control unit 26 receives a designation operation of the sensitivity coefficient at an arbitrary timing, and in step S23, generates configuration information including the sensitivity coefficient designated by the designation operation.
[0507] Also, when the method SLA1 and the method SLB1 or the method SLB2 are applied, the client 101 basically performs the reproduced audio data generation process described with reference to FIG. 18 or FIG. 34. However, in step S84 of FIG. 18 or step S126 of FIG. 34, interpolation processing based on the method SLA1 and the method SLB1 or the method SLB2 is performed.
[0508] Specifically, first, the client 101 calculates a contribution coefficient for obtaining a contribution rate by performing the contribution coefficient calculation process shown in FIG. 45.
[0509] Hereinafter, the contribution coefficient calculation process performed by the client 101 will be described with reference to the flowchart of FIG. 45.
[0510] In step S201, the position calculation unit 114 initializes an index cvpidx indicating the CVP to be processed. As a result, the value of the index cvpidx becomes 0.
[0511] In step S202, the position calculation unit 114 determines whether the value of the index cvpidx indicating the CVP to be processed is less than the number numOfCVP of all CVPs, that is, whether cvpidx < numOfCVP.
[0512] Note that the number numOfCVP of CVPs is the number of CVP candidates used in the interpolation process. Specifically, the number indicated by the CVP number information, the number of CVPs that satisfy specific conditions such as being around the listening position, the number of CVPs belonging to the CVP group corresponding to the target group area, etc. are regarded as numOfCVP.
[0513] If it is determined in step S202 that cvpidx < numOfCVP, since the contribution coefficients have not yet been calculated for all CVPs that are candidates for the interpolation process, the process proceeds to step S203.
[0514] In step S203, the position calculation unit 114 calculates the Euclidean distance from the listening position to the CVP to be processed based on the listener position information and the CVP position information of the CVP to be processed, and holds the calculation result as the distance information dist[cvpidx]. For example, the position calculation unit 114 calculates the distance information dist[cvpidx] by performing the same calculation as the above formula (5).
[0515] In step S204, the position calculation unit 114 calculates the contribution coefficient cvp_contri_coef[cvpidx] of the CVP to be processed based on the distance information dist[cvpidx] and the sensitivity coefficient WeightRatioFactor.
[0516] For example, the sensitivity coefficient WeightRatioFactor may be read from the configuration information, or may be specified by a designation operation of the listener or the like on an input unit (not shown). In addition, based on the positional relationship between the listening position and each CVP, and the gain of the object at each CVP, etc., the position calculation unit 114 may calculate the sensitivity coefficient WeightRatioFactor.
[0517] Here, the sensitivity coefficient WeightRatioFactor is, for example, a real number with a value of 2 or more. However, it is not limited to this, and the sensitivity coefficient WeightRatioFactor can be any value.
[0518] For example, the position calculation unit 114 calculates the power of the distance information dist[cvpidx] with the sensitivity coefficient WeightRatioFactor as the exponent, and divides 1 by the obtained power value, that is, obtains the reciprocal of the power value, thereby calculating the contribution coefficient cvp_contri_coef[cvpidx].
[0519] That is, by performing the operation of cvp_contri_coef[cvpidx]=1.0 / pow(dist[cvpidx],WeightRatioFactor), the contribution coefficient cvp_contri_coef[cvpidx] is obtained. Here, pow() represents a function for performing power calculation.
[0520] In step S205, the position calculation unit 114 increments the value of the CVP index cvpidx.
[0521] When the process of step S205 is performed, then the process returns to step S202, and the above-described process is repeated. That is, for the newly targeted CVP, the contribution coefficient cvp_contri_coef[cvpidx] is calculated.
[0522] Also, when it is determined in step S202 that cvpidx<numOfCVP is not true, all CVPs have been targeted and the contribution coefficient cvp_contri_coef[cvpidx] has been calculated, so the contribution coefficient calculation process ends.
[0523] As described above, the client 101 calculates the contribution coefficient according to the distance between the listening position and the CVP. By doing so, it becomes possible to perform the interpolation process based on the method SLA1, and the occurrence of unnatural audio-visual movement can be suppressed.
[0524] <Explanation of the calculation process for the normalized contribution coefficient> Furthermore, after the client 101 performs the contribution coefficient calculation process described with reference to Figure 45, it then performs a normalized contribution coefficient calculation process based on method SLB1 or method SLB2 to obtain the normalized contribution coefficient as the contribution rate.
[0525] Here, we will first refer to the flowchart in Figure 46 to explain the process of calculating the normalized contribution coefficient based on the SLB2 method, which is performed by client 101.
[0526] The normalization contribution coefficient calculation process based on the SLB2 method, as used here, is a process that calculates the normalization contribution coefficient based on the Mute flag included in the configuration information.
[0527] In step S231, the position calculation unit 114 initializes the index cvpidx, which indicates the CVP to be processed. As a result, the value of index cvpidx is set to 0.
[0528] In the normalized contribution coefficient calculation process, the same CVPs that were processed in the contribution coefficient calculation process shown in Figure 45 are processed as the target CVPs. Therefore, the number of CVPs to be processed, numOfCVP, is the same as in the contribution coefficient calculation process shown in Figure 45.
[0529] In step S232, the position calculation unit 114 initializes the index objidx, which indicates the object to be processed.
[0530] This sets the value of index objidx to 0. Here, the number of objects to be processed, numOfObjs, is the total number of objects that make up the content, that is, the number indicated by the object count information in the configuration information. Subsequently, the CVP indicated by index cvpidx and the object indicated by index objidx from the perspective of the CVP are processed in order.
[0531] In step S233, the position calculation unit 114 determines whether the value of the index objidx is less than the number numOfObjs of all objects, that is, whether objidx < numOfObjs.
[0532] If it is determined in step S233 that objidx < numOfObjs, then in step S234, the position calculation unit 114 initializes the value of the coefficient sum variable total_coef. As a result, the value of the coefficient sum variable total_coef for the object to be processed indicated by the index objidx is set to 0.
[0533] The coefficient sum variable total_coef is a coefficient used to normalize the contribution coefficient cvp_contri_coef[cvpidx] of each CVP for the object to be processed indicated by the index objidx. As will be described later, finally, the sum of the contribution coefficients cvp_contri_coef[cvpidx] of all CVPs used for interpolation processing for one object becomes the coefficient sum variable total_coef.
[0534] In step S235, the position calculation unit 114 determines whether the value of the index cvpidx indicating the CVP to be processed is less than the number numOfCVP of all CVPs, that is, whether cvpidx < numOfCVP.
[0535] If it is determined in step S235 that cvpidx < numOfCVP, then the process proceeds to step S236.
[0536] In step S236, the position calculation unit 114 determines whether the value of the Mute flag of the object to be processed indicated by the index objidx in the CVP indicated by the index cvpidx is 1, that is, whether it is a mute object.
[0537] If the value of the Mute flag is not determined to be 1 in step S236, i.e., if it is not a mute object, in step S237 the position calculation unit 114 updates the coefficient sum variable by adding the contribution coefficient of the CVP to be processed to the value of the coefficient sum variable it holds.
[0538] Specifically, total_coef+=cvp_contri_coef[cvpidx] is calculated. That is, the current value of total_coef, the sum of coefficients of the object to be processed indicated by the index objidx, which is held by the position calculation unit 114, is added to the contribution coefficient cvp_contri_coef[cvpidx] of the CVP to be processed indicated by the index cvpidx, and the result of this addition is the updated sum of coefficients variable total_coef.
[0539] Once the process in step S237 is completed, the process proceeds to step S238.
[0540] Furthermore, if the value of the Mute flag is determined to be 1 in step S236, i.e., if it is a muted object, the processing in step S237 is not performed, and the process then proceeds to step S238. This is because CVPs whose target object is a muted object are excluded from the interpolation process.
[0541] If the process in step S237 is performed, or if it is determined in step S236 that the value of the Mute flag is 1, in step S238 the position calculation unit 114 increments the index cvpidx which indicates the CVP to be processed.
[0542] Once the process in step S238 is completed, the process returns to step S235, and the process described above is repeated.
[0543] By repeating the processes of step S235 to step S238, for the object to be processed, the sum of the contribution coefficients of the CVPs that are not mute objects is obtained, and the obtained sum is set as the final coefficient sum variable of the object to be processed. This coefficient sum variable corresponds to the variable t in the above-described equation (6).
[0544] Also, when it is determined in step S235 that cvpidx < numOfCVP, in step S239, the position calculation unit 114 initializes the index cvpidx indicating the CVP to be processed. As a result, for the object to be processed, each CVP will be newly processed in order, and subsequent processing will be performed.
[0545] In step S240, the position calculation unit 114 determines whether cvpidx < numOfCVP.
[0546] When it is determined in step S240 that cvpidx < numOfCVP, the process proceeds to step S241.
[0547] ]>In step S2'41, the position calculation unit 114 determines whether the value of the Mute flag of the object to be processed indicated by the index objidx in the CVP indicated by the index cvpidx is 1.
[0548] When it is not determined in step S241 that the value of the Mute flag is 1, that is, when it is not a mute object, in step S242, the position calculation unit 114 calculates the normalized contribution coefficient contri_norm_ratio[objidx][cvpidx].
[0549] For example, the calculation of contri_norm_ratio[objidx][cvpidx] = cvp_contri_coef[cvpidx] / total_coef is performed to normalize the contribution coefficient, and the normalized contribution coefficient is set as the normalized contribution coefficient.
[0550] In other words, the position calculation unit 114 normalizes the contribution coefficient cvp_contri_coef[cvpidx] of the CVP to be processed indicated by index cvpidx by the sum of coefficients variable total_coef of the object to be processed indicated by index objidx. As a result, for the object to be processed indicated by index objidx, the normalized contribution coefficient contri_norm_ratio[objidx][cvpidx] of the CVP to be processed indicated by index cvpidx is obtained.
[0551] In this embodiment, the normalized contribution coefficient contri_norm_ratio[objidx][cvpidx] is used as the contribution rate dp(i) in equation (8), i.e., the contribution of CVP. In other words, the normalized contribution coefficient is used as the weight of CVP for each object in the interpolation process.
[0552] More specifically, in equation (8), the same contribution coefficient dp(i) common to all objects was used for the same CVP, but in this embodiment, since CVPs that are muted objects are excluded from the interpolation process, a normalized contribution coefficient (contribution coefficient dp(i)) is obtained for each object even for the same CVP.
[0553] In this case, since the contribution coefficient calculation process in Figure 45 is based on the contribution coefficient obtained by raising the distance information to a power of the sensitivity coefficient, it becomes possible to implement interpolation processing based on method SLA1.
[0554] Once the process in step S242 is completed, the process proceeds to step S244.
[0555] Furthermore, if the value of the Mute flag is determined to be 1 in step S241, that is, if it is a mute object, the processing in step S242 is not performed, and the process then proceeds to step S243.
[0556] In step S243, the position calculation unit 114 sets the value of the normalized contribution coefficient contri_norm_ratio[objidx][cvpidx] of the CVP to be processed indicated by the index cvpidx for the object to be processed indicated by the index objidx to 0.
[0557] As a result, the CVP for which the object is a mute object is excluded from the interpolation process target, and the interpolation process based on the method SLB2 can be realized.
[0558] When the process of step S242 or step S243 is performed, in step S244, the position calculation unit 114 increments the index cvpidx indicating the CVP to be processed.
[0559] When the process of step S244 is performed, then the process returns to step S240, and the above-described process is repeatedly performed.
[0560] By repeatedly performing the processes of steps S240 to S244, the normalized contribution coefficients of each CVP are obtained for the object to be processed.
[0561] Also, when it is determined in step S240 that cvpidx < numOfCVP is not satisfied, in step S245, the position calculation unit 114 increments the index objidx indicating the object to be processed. As a result, a new object that has not yet been the process target becomes the process target.
[0562] When the process of step S245 is performed, then the process returns to step S233, and the above-described process is repeatedly performed.
[0563] Also, when it is determined in step S233 that objidx < numOfObjs is not satisfied, for all objects, the normalization contribution coefficient of each CVP, that is, the contribution rate dp(i), is obtained, and thus the normalization contribution coefficient calculation process ends.
[0564] As described above, the client 101 calculates the normalization contribution coefficient of each CVP for each object according to the Mute flag of each object. By doing so, the interpolation process based on the method SLB2 can be realized, and appropriate listener reference object position information can be obtained.
[0565] In the above, the normalization contribution coefficient calculation process based on the method SLB2 has been described. However, the same process as in the case of the method SLB2 is performed as the normalization contribution coefficient calculation process based on the method SLB1.
[0566] Hereinafter, referring to the flowchart of FIG. 47, the normalization contribution coefficient calculation process based on the method SLB1 performed by the client 101 will be described.
[0567] In the normalization contribution coefficient calculation process based on the method SLB1 shown in FIG. 47, that is, steps S271 to S285, basically the same processes as steps S231 to S245 of the normalization contribution coefficient calculation process described with reference to FIG. 46 are performed.
[0568] However, in steps S276 and S281, instead of determining whether the value of the Mute flag is 1, it is determined whether the gain of the object to be processed indicated by the index objidx in the CVP indicated by the index cvpidx can be regarded as 0.
[0569] Specifically, when the value of the gain of the object is less than or equal to a predetermined threshold, it is determined that the gain of the object can be regarded as 0.
[0570] If it is determined in step S276 that the gain cannot be considered 0, the process proceeds to step S277 because it is not a muted CVP, and the coefficient sum variable is updated.
[0571] In contrast, if it is determined in step S276 that the gain can be considered to be 0, the CVP is a muted object, and therefore the CVP is excluded from the interpolation process, and the process then proceeds to step S278.
[0572] Furthermore, if it is determined in step S281 that the gain cannot be considered zero, the process proceeds to step S282 because it is not a muted CVP, and the normalized contribution coefficient is calculated.
[0573] In contrast, if it is determined in step S281 that the gain can be considered to be 0, the CVP is a muted object, and the process proceeds to step S283, where the normalization contribution coefficient is set to 0 and the CVP is excluded from the interpolation process.
[0574] By using the normalized contribution coefficient calculation process based on the SLB1 method described above, interpolation processing based on the SLB1 method can be implemented, and appropriate listener-referenced object position information can be obtained.
[0575] In step S84 in Figure 18, or step S126 in Figure 34, after the normalization contribution coefficient calculation process based on method SLB1 or method SLB2 is performed, the position calculation unit 114 then performs interpolation using the obtained normalization contribution coefficient.
[0576] In other words, the position calculation unit 114 obtains the object's 3D position vector by performing the calculation of equation (7), and instead of the contribution rate dp(i), it performs the calculation of equation (8) using the normalized contribution coefficients contri_norm_ratio[objidx][cvpidx] obtained by the above process. That is, the interpolation process of equation (8) is performed using the normalized contribution coefficients.
[0577] Furthermore, the position calculation unit 114 performs the calculation of equation (9) based on the calculation result of equation (8), and also performs corrections as appropriate based on the correction amounts obtained by the calculations of equation (10) and equation (11).
[0578] This results in the final listener-referenced object position information and listener-referenced gain, to which method SLA1 and either method SLB1 or method SLB2 have been applied.
[0579] Therefore, the occurrence of cases A and B described above is suppressed. In other words, it is possible to suppress the occurrence of unnatural sound image movement and obtain appropriate listener-referenced object position information.
[0580] In both Method SLB1 and Method SLB2, the position calculation unit 114 performs interpolation based on the CVP position information, object position information, and object gain of CVPs where the object is not substantially a muted object, and the listener position information, to calculate the listener-referenced object position information and the listener-referenced gain.
[0581] In this case, in method SLB2, the position calculation unit 114 identifies CVPs where the object is not a muted object based on the Mute flag as mute information. In contrast, in method SLB1, the position calculation unit 114 identifies CVPs where the object is not a muted object based on the object's gain as seen from the CVP, that is, based on the determination result of whether the gain is below a threshold.
[0582] <Fourth Embodiment> <Regarding object position information and gain interpolation processing> By the way, when performing interpolation processing to determine the listener-referenced object position information and listener-referenced gain, it may be possible to allow the playback side, i.e., the listener side, to intentionally select the CVP to be used for interpolation processing.
[0583] This would allow listeners to enjoy content by using only the CVPs that suit their preferences and the content they want to listen to. For example, it would be possible to play content using only CVPs where all artists, as objects, are positioned close to the listener.
[0584] Specifically, as shown in Figure 48, for example, the stage ST11, target position TP, and each CVP are positioned in the same location as in the example shown in Figure 11 in the free viewpoint space, and the listener (user) can select the CVP to be used for interpolation processing. Note that in Figure 48, the parts corresponding to those in Figure 11 are denoted by the same reference numerals, and their explanations are omitted as appropriate.
[0585] In the example shown in Figure 48, for example, let's assume that the original CVP configuration, i.e., the CVPs set up by the content creator, included CVP1 through CVP7, as shown on the left side of the figure.
[0586] In this case, for example, as shown on the right side of the diagram, suppose the listener selects CVP1, CVP3, CVP4, and CVP6 from among CVP1 to CVP7, which are located near stage ST11.
[0587] As a result, when the content is actually played, the artist as an object will feel closer to the listener compared to when all CVP is used for interpolation.
[0588] Furthermore, when a listener selects a CVP, a CVP selection screen, such as the one shown in Figure 49, may be displayed on the client 101.
[0589] In this example, the left side of the figure shows the CVP selection screen DSP11, which displays multiple viewpoint images showing what each CVP (Critical Point Venture) shown in Figure 48 looks like when viewed from that CVP to the target position TP, i.e., stage ST11.
[0590] For example, viewpoint images SPC11 through SPC14 are viewpoint images when CVP5, CVP7, CVP2, and CVP6 are set as the viewpoint position (listening position), respectively. The CVP selection screen DSP11 also displays the message "Please select the viewpoint you wish to play back" prompting the user to select a CVP.
[0591] When the CVP selection screen DSP11 is displayed, the listener (user) selects the CVP to be used for interpolation by selecting the viewpoint image corresponding to their preferred CVP. As a result, the display of the CVP selection screen DSP11 shown on the left in the figure is updated, and the CVP selection screen DSP12 shown on the right in the figure is displayed.
[0592] On the CVP selection screen DSP12, the viewpoint images of CVPs not selected by the listener are displayed in a different format than the viewpoint images of selected CVPs, such as being shown in a light gray.
[0593] Here, for example, CVP5, CVP7, and CVP2, which correspond to viewpoint images SPC11 through SPC13, are not selected, and their viewpoint images are displayed in gray. Also, the display of viewpoint images corresponding to CVP6, CVP1, CVP3, and CVP4, which were selected by the listener, remains the same as in the CVP selection screen DSP11.
[0594] By displaying such a CVP selection screen, listeners can visually confirm the view from the CVP and appropriately select the appropriate CVP. Furthermore, the CVP selection screen may also display an image of the entire venue, i.e., the entire free-viewpoint space, as shown in Figure 48.
[0595] Furthermore, if the listener is allowed to select the CVP used for interpolation processing, the configuration information may be used to store information regarding whether or not the CVP can be selected. This would allow the content creator's intentions to be transmitted to the listener (client 101).
[0596] If the configuration information contains information regarding the selectability of CVPs, the playback system will only consider CVPs that the listener is permitted to select, and the system will detect (identify) whether or not the listener has selected a CVP. If there are CVPs among those permitted to be selected that were not selected by the listener (hereinafter also referred to as unselected CVPs), the unselected CVPs will be excluded, and interpolation processing will be performed to calculate listener-referenced object position information and listener-referenced gain.
[0597] <Example of configuration information format> In this embodiment, the configuration information includes, for example, the information shown in Figure 50, which is stored as information regarding whether the CVP can be selected, i.e., as selectability information.
[0598] Figure 50 shows an example of the format (syntax) of a portion of the configuration information.
[0599] More specifically, the configuration information includes the configuration shown in Figure 50, as well as the configuration shown in Figure 7. In other words, the configuration information includes the configuration shown in Figure 50 as part of the configuration shown in Figure 7. Alternatively, the configuration information shown in Figure 32 may include the configuration shown in Figure 50, or the configuration information may also store the information shown in Figure 44.
[0600] In the example in Figure 50, "CVPSelectAllowPresentFlag" indicates the CVP selection information existence flag. The CVP selection information existence flag is a flag that indicates whether or not information about selectable CVPs exists in the configuration information, that is, whether or not the listener can select a CVP.
[0601] A value of "0" for the CVP selection information existence flag indicates that the configuration information does not contain (is not stored) information about selectable CVPs.
[0602] Furthermore, a CVP selection information existence flag value of "1" indicates that the configuration information includes information about selectable CVPs.
[0603] If the value of the CVP selection information existence flag is "1", the configuration information also stores "numOfAllowedCVP", which indicates the number of CVPs that can be selected by the listener, and "AllowedCVPIdx[i]", which is index information indicating the CVPs that can be selected by the listener.
[0604] For example, the index information "AllowedCVPIdx[i]" is the value of the CVP index "ControlViewpointIndex[i]" shown in Figure 9, which indicates the CVP that can be selected by the listener. In addition, the configuration information stores the index information "AllowedCVPIdx[i]" which indicates the number of CVPs that can be selected, as shown by "numOfAllowedCVP".
[0605] As described above, in the example in Figure 50, the configuration information includes the CVP selection information existence flag "CVPSelectAllowPresentFlag", the number of selectable CVPs "numOfAllowedCVP", and the index information "AllowedCVPIdx[i]", as selectability information regarding the selectability of CVPs used to calculate the listener-referenced object position information and listener-referenced gain.
[0606] Using this configuration information, client 101 can identify which of the CVPs that make up the content are the CVPs that are permitted to be selected.
[0607] Furthermore, this embodiment can be combined with any one or more of the first to third embodiments described above.
[0608] Even when the configuration information includes the configuration shown in Figure 50, the information processing device 11 performs the content creation process described with reference to Figure 14.
[0609] However, in this case, for example, the control unit 26 accepts a specification operation at any timing, such as in step S16, to determine whether or not to allow the selection of a CVP. Then, in step S23, the control unit 26 generates configuration information that includes the necessary information from among the CVP selection information existence flag, the number of selectable CVPs, and index information indicating the selectable CVPs, in response to the specification operation.
[0610] <Example of client structure> Furthermore, if the client 101 (listener) can select the CVP to be used for interpolation processing, the client 101 will have the configuration shown in Figure 51, for example. Note that in Figure 51, the same reference numerals are used for parts corresponding to those in Figure 17, and their explanations will be omitted as appropriate.
[0611] The configuration of client 101 shown in Figure 51 is the same as the configuration shown in Figure 17, but with the addition of an input unit 201 and a display unit 202.
[0612] The input unit 201 consists of input devices such as a touch panel, mouse, keyboard, and buttons, and supplies signals to the position calculation unit 114 in accordance with the input operations of the listener (user).
[0613] The display unit 202 consists of a display and displays various images, such as the CVP selection screen, in response to instructions from the position calculation unit 114, etc.
[0614] <Explanation of Selective Interpolation> Even if client 101 can appropriately select the CVP to be used for interpolation processing, client 101 basically performs the playback audio data generation process described with reference to Figure 18 or Figure 34.
[0615] However, in step S84 in Figure 18, or step S126 in Figure 34, the selective interpolation process shown in Figure 52 is performed to obtain the listener-referenced object position information and the listener-referenced gain.
[0616] The selective interpolation process performed by client 101 will be explained below with reference to the flowchart in Figure 52.
[0617] In step S311, the position calculation unit 114 obtains configuration information from the decoding unit 113.
[0618] In step S312, the position calculation unit 114 determines, based on the configuration information, whether the number of selectable CVPs is greater than 0, that is, whether numOfAllowedCVP>0.
[0619] If it is determined in step S312 that numOfAllowedCVP > 0, that is, if there are CVPs that can be selected by the listener, in step S313 the position calculation unit 114 presents the selectable CVPs and accepts the listener's selection of a CVP.
[0620] For example, the position calculation unit 114 generates a CVP selection screen based on the index information "AllowedCVPIdx[i]" which indicates selectable CVPs included in the configuration information, presenting the CVPs indicated by that index information as selectable CVPs, and displays it on the display unit 202. In this case, the display unit 202 displays, for example, the CVP selection screen shown in Figure 49.
[0621] The listener (user) selects the desired CVP to be used for interpolation processing by operating the input unit 201 while viewing the CVP selection screen displayed on the display unit 202.
[0622] Then, the input unit 201 supplies a signal to the position calculation unit 114 corresponding to the listener's selection operation, and the position calculation unit 114 updates the screen of the display unit 202 in accordance with the signal from the input unit 201. As a result, for example, the display on the display unit 202 is updated from the display shown on the left side of Figure 49 to the display shown on the right side of Figure 49.
[0623] Furthermore, the listener's selection of a CVP on the CVP selection screen may be performed before content playback, or it may be performed at any time during content playback, and any number of times.
[0624] In step S314, the position calculation unit 114 determines, based on the signal supplied from the input unit 201 in response to the listener's selection operation, whether or not there are any CVPs among the selectable CVPs that have been excluded from the interpolation process, that is, whether or not there are any CVPs that were not selected by the listener.
[0625] If it is determined in step S314 that there are CVPs that have been excluded, the process then proceeds to step S315.
[0626] In step S315, the position calculation unit 114 performs interpolation using the CVP that cannot be selected and the CVP selected by the listener to obtain listener-referenced object position information and listener-referenced gain.
[0627] More specifically, interpolation is performed based on the CVP position information, object position information, object gain, etc., of each of the multiple CVPs, which consist of CVPs that cannot be selected and CVPs selected by the listener, as well as the listener's position information.
[0628] Here, an unselectable CVP is a CVP whose configuration information does not include the index information "AllowedCVPIdx[i]". In other words, an unselectable CVP is a CVP that is not designated as selectable by the selectability information included in the configuration information.
[0629] Therefore, in step S315, interpolation is performed using all CVPs remaining after excluding the CVPs that were not selected by the listener, i.e., the unselected CVPs.
[0630] In other words, for example, all CVPs except for non-selected CVPs are used, and interpolation processing is performed in the same manner as in the first and third embodiments to obtain listener-referenced object position information and listener-referenced gain.
[0631] Furthermore, the interpolation process may be performed using CVPs obtained by excluding non-selected CVPs from all CVPs, for example, by excluding non-selected CVPs from CVPs that meet specific conditions, such as being around the listening position, or by excluding non-selected CVPs from CVPs belonging to a CVP group corresponding to the target group area.
[0632] Once step S315 is completed, the selective interpolation process is terminated.
[0633] Furthermore, if it is determined in step S312 that numOfAllowedCVP > 0, that is, if there are no selectable CVPs, or if it is determined in step S314 that there are no excluded CVPs, the process then proceeds to step S316.
[0634] In step S316, the position calculation unit 114 performs interpolation using all CVPs to obtain listener-referenced object position information and listener-referenced gain, and the selective interpolation process is completed.
[0635] In step S316, the interpolation process is performed in the same way as in step S315, except that the CVP used for the interpolation process is different. In addition, in step S316, the interpolation process may also be performed using CVPs that satisfy specific conditions or CVPs belonging to the CVP group corresponding to the target group region.
[0636] As described above, client 101 selectively uses CVP (Computer-Generated Video) based on the listener's selection and performs interpolation processing. In this way, it is possible to achieve content playback that reflects both the content creator's intentions and the listener's (user's) preferences.
[0637] <Example of computer configuration> Incidentally, the series of processes described above can be executed by hardware or by software. When the series of processes are executed by software, the programs that make up that software are installed on a computer. Here, a computer includes computers built into dedicated hardware, as well as general-purpose personal computers that can perform various functions by installing various programs.
[0638] Figure 53 is a block diagram showing an example of the hardware configuration of a computer that executes the series of processes described above by a program.
[0639] In a computer, the CPU (Central Processing Unit) 501, ROM (Read Only Memory) 502, and RAM (Random Access Memory) 503 are interconnected by a bus 504.
[0640] An input / output interface 505 is further connected to the bus 504. An input unit 506, an output unit 507, a recording unit 508, a communication unit 509, and a drive 510 are connected to the input / output interface 505.
[0641] The input unit 506 consists of a keyboard, mouse, microphone, image sensor, etc. The output unit 507 consists of a display, speaker, etc. The recording unit 508 consists of a hard disk, non-volatile memory, etc. The communication unit 509 consists of a network interface, etc. The drive 510 drives a removable recording medium 511 such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory.
[0642] In a computer configured as described above, the CPU 501 loads, for example, a program stored in the recording unit 508 into the RAM 503 via the input / output interface 505 and the bus 504, and executes it, thereby performing the series of processes described above.
[0643] The program executed by the computer (CPU 501) can be provided by recording it on a removable recording medium 511, such as a packaged media. The program can also be provided via wired or wireless transmission media, such as a local area network, the internet, or digital satellite broadcasting.
[0644] In a computer, a program can be installed in the recording unit 508 via the input / output interface 505 by inserting the removable recording medium 511 into the drive 510. Alternatively, the program can be received by the communication unit 509 via a wired or wireless transmission medium and installed in the recording unit 508. Furthermore, the program can be pre-installed in the ROM 502 or the recording unit 508.
[0645] The programs executed by the computer may be programs that are processed chronologically in the order described herein, or they may be programs that are processed in parallel or at necessary times, such as when a call is made.
[0646] Furthermore, the embodiments of this technology are not limited to those described above, and various modifications are possible without departing from the spirit of this technology.
[0647] For example, this technology can be configured as cloud computing, where a single function is shared and processed collaboratively by multiple devices via a network.
[0648] Furthermore, each step described in the flowchart above can be performed by a single device, or it can be divided and performed by multiple devices.
[0649] Furthermore, if a single step includes multiple processes, those processes can be executed by a single device or shared among multiple devices.
[0650] Furthermore, this technology can also be configured as follows:
[0651] (1) Multiple metadata sets are generated, each containing metadata for multiple objects, including object position information indicating the position of an object as seen from the control viewpoint, where the direction from the control viewpoint towards the target position in space is defined as the direction of the median plane. For each of the multiple control viewpoints, control viewpoint information is generated, which includes control viewpoint position information indicating the position of the control viewpoint in the space, and information indicating the metadata set associated with the control viewpoint from among the multiple metadata sets. Generate content data that includes multiple sets of metadata that are different from each other, and configuration information that includes the control viewpoint information of multiple control viewpoints. Equipped with a control unit Information processing device. (2) The metadata includes the gain of the object. (1) The information processing device described above. (3) The control viewpoint information includes control viewpoint orientation information indicating the direction from the control viewpoint to the target position in the space, or target position information indicating the target position in the space. The information processing device described in (1) or (2). (4) The configuration information includes at least one of the following: object count information indicating the number of objects constituting the content, control viewpoint count information indicating the number of control viewpoints, and metadata set count information indicating the number of metadata sets. An information processing device as described in any one of items (1) to (3). (5) The configuration information includes control viewpoint group information relating to a control viewpoint group consisting of the control viewpoints included in a predetermined group region within the space, The control viewpoint group information includes, for one or more control viewpoint groups, information indicating the control viewpoints belonging to the control viewpoint group, and information for identifying the group region corresponding to the control viewpoint group. An information processing device as described in any one of items (1) through (4). (6) The configuration information includes information indicating whether or not the control viewpoint group information is included. (5) The information processing device described above. (7) The control viewpoint group information includes at least one of the following: information indicating the number of control viewpoints belonging to the control viewpoint group, and information indicating the number of control viewpoint groups. The information processing device described in (5) or (6). (8) The configuration information includes mute information for identifying the object that is considered a mute object when viewed from the control viewpoint. An information processing device as described in any one of items (1) through (7). (9) The configuration information includes selectability information regarding the selectability of the control viewpoint used to calculate listener-referenced object position information, which indicates the position of the object as seen from the listening position, or to calculate the gain of the object as seen from the listening position. An information processing device as described in any one of items (1) through (8). (10) Information processing device, Multiple metadata sets are generated, each containing metadata for multiple objects, including object position information indicating the position of an object as seen from the control viewpoint, where the direction from the control viewpoint towards the target position in space is defined as the direction of the median plane. For each of the multiple control viewpoints, control viewpoint information is generated, which includes control viewpoint position information indicating the position of the control viewpoint in the space, and information indicating the metadata set associated with the control viewpoint from among the multiple metadata sets. Generate content data that includes multiple sets of metadata that are different from each other, and configuration information that includes the control viewpoint information of multiple control viewpoints. Information processing methods. (11) Multiple metadata sets are generated, each containing metadata for multiple objects, including object position information indicating the position of an object as seen from the control viewpoint, where the direction from the control viewpoint towards the target position in space is defined as the direction of the median plane. For each of the multiple control viewpoints, control viewpoint information is generated, which includes control viewpoint position information indicating the position of the control viewpoint in the space, and information indicating the metadata set associated with the control viewpoint from among the multiple metadata sets. Generate content data that includes multiple sets of metadata that are different from each other, and configuration information that includes the control viewpoint information of multiple control viewpoints. A program that instructs a computer to perform a process. (12) An acquisition unit that acquires object position information indicating the position of an object as seen from the control viewpoint, where the direction from the control viewpoint toward the target position in space is defined as the direction of the median plane, and control viewpoint position information indicating the position of the control viewpoint in space. A listener location information acquisition unit acquires listener location information indicating the listening position in the aforementioned space, A position calculation unit calculates listener-referenced object position information indicating the position of the object as seen from the listening position, based on the listener position information, the control viewpoint position information of the plurality of control viewpoints, and the object position information of the plurality of control viewpoints. An information processing device equipped with the following features. (13) The acquisition unit acquires a metadata set consisting of metadata for a plurality of objects including the object position information, the control viewpoint position information, and designation information indicating the metadata set associated with the control viewpoint. The position calculation unit calculates the listener-referenced object position information based on the object position information included in the metadata set indicated by the specified information, from among a plurality of mutually different metadata sets. (12) The information processing device described above. (14) The position calculation unit calculates the listener-referenced object position information by interpolation processing based on the listener position information, the control viewpoint position information of the plurality of control viewpoints, and the object position information of the plurality of control viewpoints. The information processing device described in (12) or (13). (15) The interpolation process is vector synthesis. (14) The information processing device described above. (16) The position calculation unit performs vector synthesis using weights obtained from the listener position information and the control viewpoint position information of the multiple control viewpoints. (15) The information processing device described above. (17) The position calculation unit performs the interpolation process based on the control viewpoint position information and object position information of the control viewpoint where the object is not a muted object. An information processing device as described in any one of paragraphs (14) to (16). (18) The acquisition unit further acquires mute information to identify the object that is designated as the mute object when viewed from the control viewpoint, Based on the mute information, the position calculation unit identifies the control viewpoint where the object is not the mute object. (17) The information processing device described above. (19) The acquisition unit further acquires the gain of the object as seen from each of the control viewpoints for each of the control viewpoints, The position calculation unit identifies the control viewpoint in which the object is not the muted object, based on the gain. (17) The information processing device described above. (20) The acquisition unit further acquires selection availability information regarding the selection availability of the control viewpoint used in calculating the listener reference object position information, The position calculation unit performs the interpolation process based on the control viewpoint position information and object position information of the control viewpoint selected by the listener from among the control viewpoints that can be selected according to the selectability information. An information processing device as described in any one of paragraphs (14) through (19). (twenty one) The position calculation unit performs the interpolation process based on the control viewpoint position information and object position information of the control viewpoint that is not selectable according to the selectability information, and the control viewpoint position information and object position information of the control viewpoint selected by the listener. (20) The information processing device described above. (twenty two) The listener location information acquisition unit acquires listener orientation information indicating the orientation of the listener in the space, The position calculation unit calculates the listener reference object position information based on the listener orientation information, the listener position information, the control viewpoint position information of the plurality of control viewpoints, and the object position information of the plurality of control viewpoints. An information processing device as described in any one of paragraphs (12) through (21). (twenty three) The acquisition unit further acquires control viewpoint orientation information for a plurality of control viewpoints, indicating the direction from the control viewpoint in the space toward the target position. The position calculation unit calculates the listener reference object position information based on the control viewpoint orientation information of the plurality of control viewpoints, the listener orientation information, the listener position information, the control viewpoint position information of the plurality of control viewpoints, and the object position information of the plurality of control viewpoints. (22) The information processing device described above. (twenty four) The acquisition unit further acquires the gain of the object as seen from each of the control viewpoints for each of the control viewpoints, The position calculation unit calculates the gain of the object as seen from the listening position by interpolation processing based on the listener position information, the control viewpoint position information of the plurality of control viewpoints, and the gains of the plurality of control viewpoints. An information processing device as described in any one of paragraphs (12) to (23). (twenty five) The position calculation unit performs the interpolation process based on a weight obtained from the reciprocal of the power of the distance from the listening position to the control viewpoint, with a predetermined sensitivity coefficient as the exponent. (24) The information processing device described above. (26) The sensitivity coefficient is set for each control viewpoint, or for each object as viewed from the control viewpoint. (25) The information processing device described above. (27) The acquisition unit further acquires selectability information regarding the selectability of the control viewpoint used to calculate the gain of the object as viewed from the listening position, The position calculation unit performs the interpolation process based on the control viewpoint position information and gain of the control viewpoint selected by the listener from among the control viewpoints that can be selected according to the selectability information. An information processing device as described in any one of paragraphs (24) to (26). (28) The position calculation unit performs the interpolation process based on the control viewpoint position information and gain of the control viewpoint that is not selectable according to the selectability information, and the control viewpoint position information and gain of the control viewpoint selected by the listener. (27) The information processing device described above. (29) The system further includes a rendering processing unit that performs rendering processing based on the audio data of the object and the listener-referenced object position information. An information processing device as described in any one of paragraphs (12) through (28). (30) The aforementioned listener-referenced object position information is information indicating the position of the object, expressed in coordinates of a polar coordinate system with the listening position as the origin. An information processing device as described in any one of paragraphs (12) through (29). (31) The acquisition unit further acquires control viewpoint group information relating to a control viewpoint group consisting of the control viewpoints included in a predetermined group region within the space, which includes, for one or more of the control viewpoint groups, information indicating the control viewpoints belonging to the control viewpoint group, and information for identifying the group region corresponding to the control viewpoint group. The position calculation unit calculates the listener reference object position information based on the control viewpoint position information and object position information of the control viewpoint belonging to the control viewpoint group corresponding to the group region including the listening position, and the listener position information. An information processing device as described in any one of paragraphs (12) through (30). (32) The position calculation unit acquires configuration information including control viewpoint information for each of the plurality of control viewpoints, including the control viewpoint position information, and information indicating whether or not the control viewpoint group information is included. The configuration information includes the control viewpoint group information, depending on whether or not the control viewpoint group information is included. (31) The information processing device described above. (33) The control viewpoint group information includes at least one of the following: information indicating the number of control viewpoints belonging to the control viewpoint group, and information indicating the number of control viewpoint groups. The information processing device described in (31) or (32). (34) The acquisition unit is, Control viewpoint information for each of the plurality of control viewpoints, including the control viewpoint position information, At least one of the following: object count information indicating the number of objects constituting the content, control viewpoint count information indicating the number of control viewpoints, and metadata set count information indicating the number of metadata sets consisting of metadata for multiple objects including object position information. Get configuration information that includes this information. An information processing device as described in any one of paragraphs (12) through (33). (35) Information processing device, Object position information indicating the position of the object as seen from the control viewpoint, where the direction from the control viewpoint toward the target position in space is defined as the direction of the median plane, and control viewpoint position information indicating the position of the control viewpoint in space are obtained. The listener's position information indicating the listening position in the aforementioned space is acquired, Based on the listener position information, the control viewpoint position information of the multiple control viewpoints, and the object position information of the multiple control viewpoints, listener-referenced object position information indicating the position of the object as seen from the listener position is calculated. Information processing methods. (36) Object position information indicating the position of the object as seen from the control viewpoint, where the direction from the control viewpoint toward the target position in space is defined as the direction of the median plane, and control viewpoint position information indicating the position of the control viewpoint in space are obtained. The listener's position information indicating the listening position in the aforementioned space is acquired, Based on the listener position information, the control viewpoint position information of the multiple control viewpoints, and the object position information of the multiple control viewpoints, listener-referenced object position information indicating the position of the object as seen from the listener position is calculated. A program that instructs a computer to perform a process. [Explanation of symbols]
[0652] 11 Information processing unit, 21 Input unit, 22 Display unit, 24 Communication unit, 26 Control unit, 51 Server, 61 Communication unit, 62 Control unit, 71 Encoding unit, 101 Client, 111 Listener location information acquisition unit, 112 Communication unit, 113 Decoding unit, 114 Location calculation unit, 115 Rendering processing unit
Claims
1. Multiple metadata sets are generated, each containing metadata for multiple objects, including object position information indicating the position of an object as seen from the control viewpoint, where the direction from the control viewpoint towards the target position in space is defined as the direction of the median plane. For each of the multiple control viewpoints, control viewpoint information is generated, which includes control viewpoint position information indicating the position of the control viewpoint in the space, and information indicating the metadata set associated with the control viewpoint from among the multiple metadata sets. Generate content data that includes multiple sets of metadata that are different from each other, and configuration information that includes the control viewpoint information of multiple control viewpoints. Equipped with a control unit Information processing device.
2. The metadata includes the gain of the object. The information processing apparatus according to claim 1.
3. The control viewpoint information includes control viewpoint orientation information indicating the direction from the control viewpoint to the target position in the space, or target position information indicating the target position in the space. The information processing apparatus according to claim 1.
4. The configuration information includes at least one of the following: object count information indicating the number of objects constituting the content, control viewpoint count information indicating the number of control viewpoints, and metadata set count information indicating the number of metadata sets. The information processing apparatus according to claim 1.
5. The configuration information includes control viewpoint group information relating to a control viewpoint group consisting of the control viewpoints included in a predetermined group region within the space, The control viewpoint group information includes, for one or more control viewpoint groups, information indicating the control viewpoints belonging to the control viewpoint group, and information for identifying the group region corresponding to the control viewpoint group. The information processing apparatus according to claim 1.
6. The configuration information includes information indicating whether or not the control viewpoint group information is included. The information processing apparatus according to claim 5.
7. The control viewpoint group information includes at least one of the following: information indicating the number of control viewpoints belonging to the control viewpoint group, and information indicating the number of control viewpoint groups. The information processing apparatus according to claim 5.
8. The configuration information includes mute information for identifying the object that is considered a mute object when viewed from the control viewpoint. The information processing apparatus according to claim 1.
9. The configuration information includes selectability information regarding the selectability of the control viewpoint used to calculate listener-referenced object position information, which indicates the position of the object as seen from the listening position, or to calculate the gain of the object as seen from the listening position. The information processing apparatus according to claim 1.
10. Information processing device, Multiple metadata sets are generated, each containing metadata for multiple objects, including object position information indicating the position of an object as seen from the control viewpoint, where the direction from the control viewpoint towards the target position in space is defined as the direction of the median plane. For each of the multiple control viewpoints, control viewpoint information is generated, which includes control viewpoint position information indicating the position of the control viewpoint in the space, and information indicating the metadata set associated with the control viewpoint from among the multiple metadata sets. Generate content data that includes multiple sets of metadata that are different from each other, and configuration information that includes the control viewpoint information of multiple control viewpoints. Information processing methods.
11. Multiple metadata sets are generated, each containing metadata for multiple objects, including object position information indicating the position of an object as seen from the control viewpoint, where the direction from the control viewpoint towards the target position in space is defined as the direction of the median plane. For each of the multiple control viewpoints, control viewpoint information is generated, which includes control viewpoint position information indicating the position of the control viewpoint in the space, and information indicating the metadata set associated with the control viewpoint from among the multiple metadata sets. Generate content data that includes multiple sets of metadata that are different from each other, and configuration information that includes the control viewpoint information of multiple control viewpoints. A program that instructs a computer to perform a process.
12. An acquisition unit that acquires object position information indicating the position of an object as seen from the control viewpoint, where the direction from the control viewpoint toward the target position in space is defined as the direction of the median plane, and control viewpoint position information indicating the position of the control viewpoint in space. A listener position information acquisition unit acquires listener position information indicating the listening position in the space; and a position calculation unit calculates listener reference object position information indicating the position of the object as seen from the listening position, based on the listener position information, the control viewpoint position information of the plurality of control viewpoints, and the object position information of the plurality of control viewpoints. An information processing device equipped with the following features.
13. The acquisition unit acquires a metadata set consisting of metadata for a plurality of objects including the object position information, the control viewpoint position information, and designation information indicating the metadata set associated with the control viewpoint. The position calculation unit calculates the listener-referenced object position information based on the object position information included in the metadata set indicated by the specified information, from among a plurality of mutually different metadata sets. The information processing apparatus according to claim 12.
14. The position calculation unit calculates the listener-referenced object position information by interpolation processing based on the listener position information, the control viewpoint position information of the plurality of control viewpoints, and the object position information of the plurality of control viewpoints. The information processing apparatus according to claim 12.
15. The interpolation process is vector synthesis. The information processing apparatus according to claim 14.
16. The position calculation unit performs vector synthesis using weights obtained from the listener position information and the control viewpoint position information of the multiple control viewpoints. The information processing apparatus according to claim 15.
17. The position calculation unit performs the interpolation process based on the control viewpoint position information and object position information of the control viewpoint where the object is not a muted object. The information processing apparatus according to claim 14.
18. The acquisition unit further acquires mute information to identify the object that is designated as the mute object when viewed from the control viewpoint, Based on the mute information, the position calculation unit identifies the control viewpoint where the object is not the mute object. The information processing apparatus according to claim 17.
19. The acquisition unit further acquires the gain of the object as seen from each of the control viewpoints for each of the control viewpoints, The position calculation unit identifies the control viewpoint in which the object is not the muted object, based on the gain. The information processing apparatus according to claim 17.
20. The acquisition unit further acquires selection availability information regarding the selection availability of the control viewpoint used in calculating the listener reference object position information, The position calculation unit performs the interpolation process based on the control viewpoint position information and object position information of the control viewpoint selected by the listener from among the control viewpoints that can be selected according to the selectability information. The information processing apparatus according to claim 14.
21. The position calculation unit performs the interpolation process based on the control viewpoint position information and object position information of the control viewpoint that is not selectable according to the selectability information, and the control viewpoint position information and object position information of the control viewpoint selected by the listener. The information processing apparatus according to claim 20.
22. The listener location information acquisition unit acquires listener orientation information indicating the orientation of the listener in the space, The position calculation unit calculates the listener reference object position information based on the listener orientation information, the listener position information, the control viewpoint position information of the plurality of control viewpoints, and the object position information of the plurality of control viewpoints. The information processing apparatus according to claim 12.
23. The acquisition unit further acquires control viewpoint orientation information for a plurality of control viewpoints, indicating the direction from the control viewpoint in the space toward the target position. The position calculation unit calculates the listener reference object position information based on the control viewpoint orientation information of the plurality of control viewpoints, the listener orientation information, the listener position information, the control viewpoint position information of the plurality of control viewpoints, and the object position information of the plurality of control viewpoints. The information processing apparatus according to claim 22.
24. The acquisition unit further acquires the gain of the object as seen from each of the control viewpoints for each of the control viewpoints, The position calculation unit calculates the gain of the object as seen from the listening position by interpolation processing based on the listener position information, the control viewpoint position information of the plurality of control viewpoints, and the gains of the plurality of control viewpoints. The information processing apparatus according to claim 12.
25. The position calculation unit performs the interpolation process based on a weight obtained from the reciprocal of the power of the distance from the listening position to the control viewpoint, with a predetermined sensitivity coefficient as the exponent. The information processing apparatus according to claim 24.
26. The sensitivity coefficient is set for each control viewpoint, or for each object as viewed from the control viewpoint. The information processing apparatus according to claim 25.
27. The acquisition unit further acquires selectability information regarding the selectability of the control viewpoint used to calculate the gain of the object as viewed from the listening position, The position calculation unit performs the interpolation process based on the control viewpoint position information and gain of the control viewpoint selected by the listener from among the control viewpoints that can be selected according to the selectability information. The information processing apparatus according to claim 24.
28. The position calculation unit performs the interpolation process based on the control viewpoint position information and gain of the control viewpoint that is not selectable according to the selectability information, and the control viewpoint position information and gain of the control viewpoint selected by the listener. The information processing apparatus according to claim 27.
29. The system further includes a rendering processing unit that performs rendering processing based on the audio data of the object and the listener-referenced object position information. The information processing apparatus according to claim 12.
30. The aforementioned listener-referenced object position information is information indicating the position of the object, expressed in coordinates of a polar coordinate system with the listening position as the origin. The information processing apparatus according to claim 12.
31. The acquisition unit further acquires control viewpoint group information relating to a control viewpoint group consisting of the control viewpoints included in a predetermined group region within the space, which includes, for one or more of the control viewpoint groups, information indicating the control viewpoints belonging to the control viewpoint group, and information for identifying the group region corresponding to the control viewpoint group. The position calculation unit calculates the listener reference object position information based on the control viewpoint position information and object position information of the control viewpoint belonging to the control viewpoint group corresponding to the group region including the listening position, and the listener position information. The information processing apparatus according to claim 12.
32. The position calculation unit acquires configuration information including control viewpoint information for each of the plurality of control viewpoints, including the control viewpoint position information, and information indicating whether or not the control viewpoint group information is included. The configuration information includes the control viewpoint group information, depending on whether or not the control viewpoint group information is included. The information processing apparatus according to claim 31.
33. The control viewpoint group information includes at least one of the following: information indicating the number of control viewpoints belonging to the control viewpoint group, and information indicating the number of control viewpoint groups. The information processing apparatus according to claim 31.
34. The acquisition unit is, Control viewpoint information for each of the plurality of control viewpoints, including the control viewpoint position information, At least one of the following: object count information indicating the number of objects constituting the content, control viewpoint count information indicating the number of control viewpoints, and metadata set count information indicating the number of metadata sets consisting of metadata for multiple objects including object position information. Get configuration information that includes this information. The information processing apparatus according to claim 12.
35. Information processing device, Object position information indicating the position of the object as seen from the control viewpoint, where the direction from the control viewpoint toward the target position in space is defined as the direction of the median plane, and control viewpoint position information indicating the position of the control viewpoint in space are obtained. The listener's position information indicating the listening position in the aforementioned space is acquired, Based on the listener position information, the control viewpoint position information of the multiple control viewpoints, and the object position information of the multiple control viewpoints, listener-referenced object position information indicating the position of the object as seen from the listener position is calculated. Information processing methods.
36. Object position information indicating the position of the object as seen from the control viewpoint, where the direction from the control viewpoint toward the target position in space is defined as the direction of the median plane, and control viewpoint position information indicating the position of the control viewpoint in space are obtained. The listener's position information indicating the listening position in the aforementioned space is acquired, Based on the listener position information, the control viewpoint position information of the multiple control viewpoints, and the object position information of the multiple control viewpoints, listener-referenced object position information indicating the position of the object as seen from the listener position is calculated. A program that instructs a computer to perform a process.