Information processing device and method, and program

The information processing apparatus addresses the challenge of calculating the listener's orientation in free viewpoint audio systems by defining a prohibited space around the target point and changing the listener's position if necessary, ensuring appropriate UI display and artistic quality.

WO2025121113A1PCT designated stage expired Publication Date: 2025-06-12SONY GROUP CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/040766
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-05
Filing Date
2024-11-18
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Existing free viewpoint audio systems face challenges in appropriately displaying the orientation of a listener's face when the listener is at the same position as the target point, leading to calculation impossibilities and potential UI display failures.

Method used

An information processing apparatus and method that identify whether the user's position is within a prohibited space, defined around a target point, and change the user's position to a different location outside the prohibited space to prevent calculation failures and ensure appropriate UI display.

Benefits of technology

Enables the appropriate display of free viewpoint content by preventing calculation failures and ensuring the listener's orientation can be correctly determined, even when the listener is at the target point, thereby maintaining artistic quality in the content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024040766_12062025_PF_FP_ABST
    Figure JP2024040766_12062025_PF_FP_ABST
Patent Text Reader

Abstract

The present technology relates to an information processing device and method, and a program that make it possible to appropriately display content. This information processing device is provided with a position determination unit that, on the basis of position information indicating the position of a user in a space and prohibited space information for identifying a prohibited space which includes a prescribed target point and into which intrusion by the user is prohibited, identifies whether the position of the user indicated by the position information is within the prohibited space, and when the position of the user is within the prohibited space, changes the position of the user to a position that is outside the prohibited space and is different from the position indicated by the position information. The present technology can be applied to an information processing device.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, method, and program

[0001] The present technology relates to an information processing device, method, and program, and in particular to an information processing device, method, and program that enable appropriate display of content.

[0002] Free viewpoint technology, also known as 6DoF (Degree of Freedom), has been used in various content such as games and music. In free viewpoint content, users (listeners) can rotate their head up, down, left, right, and diagonally in a three-dimensional space, and can also move to any position within the space to view the content.

[0003] For example, as a technology relating to free viewpoint audio, a technology has been proposed that enables content playback based on the intentions of the content creator (see, for example, Patent Document 1).

[0004] In Patent Document 1, 3DoF content is produced on a free viewpoint audio system at multiple control viewpoints called CVPs (Control Viewpoints) set by the producer. The CVPs are, for example, desired listening positions when the content is played back.

[0005] In particular, in Patent Document 1, a target point (hereinafter also referred to as TP) is set as a point (position) common to all CVPs in a three-dimensional space, and 3DoF content is produced with the TP as the median plane. In other words, the 3DoF content is produced assuming that the listener in the CVP is facing (looking at) the TP.

[0006] International Publication No. 2023 / 085140

[0007] In Patent Document 1, a case is considered in which a listener in a three-dimensional space, that is, a free viewpoint space, moves around while always looking at the TP on the content playback side.

[0008] In such a case, the direction of the listener's face as displayed on a playback system using a three-dimensional display user interface (UI) can be uniquely determined from a calculation formula based on the position of the listener and the position of the TP in three-dimensional space.

[0009] However, if the listener and TP are in the same location, that is, if the listener is in the TP, it may become impossible to perform calculations according to the formula, and the UI may not be displayed properly (UI expression).

[0010] The present technology has been made in view of such circumstances, and makes it possible to appropriately display content.

[0011] An information processing device according to one aspect of the present technology includes a position determination unit that, based on position information indicating a user's position within a space and prohibited space information for identifying a prohibited space in which the user is prohibited from entering, including a predetermined target point, determines whether the user's position indicated by the position information is within the prohibited space, and, if the user's position is within the prohibited space, changes the user's position to a position outside the prohibited space that is different from the position indicated by the position information.

[0012] An information processing method or program according to one aspect of the present technology includes a step of determining whether the user's position indicated by the position information is within a prohibited space based on position information indicating the user's position within a space and prohibited space information for identifying a prohibited space into which the user is prohibited, including a predetermined target point, and if the user's position is within the prohibited space, changing the user's position to a position outside the prohibited space that is different from the position indicated by the position information.

[0013] In one aspect of the present technology, based on position information indicating the user's position within a space and prohibited space information for identifying a prohibited space, including a predetermined target point, into which the user is prohibited from entering, it is determined whether the user's position indicated by the position information is within the prohibited space, and if the user's position is within the prohibited space, the user's position is changed to a position outside the prohibited space that is different from the position indicated by the position information.

[0014] FIG. 1 is a diagram illustrating the direction of a listener's face. FIG. 2 is a diagram illustrating a prohibited space. FIG. 3 is a diagram illustrating an example of configuration information. FIG. 4 is a diagram illustrating a listener's entry into a prohibited space. FIG. 5 is a diagram illustrating an example of the configuration of a server. FIG. 6 is a flowchart illustrating encoding processing. FIG. 7 is a diagram illustrating an example of the configuration of a client. FIG. 8 is a flowchart illustrating output audio data generation processing. FIG. 9 is a flowchart illustrating metadata generation processing. FIG. 10 is a diagram illustrating vector synthesis. FIG. 11 is a diagram illustrating the contribution rate of each vector during vector synthesis. FIG. 12 is a flowchart illustrating metadata generation processing. FIG. 13 is a diagram illustrating an example of the configuration of a computer.

[0015] Hereinafter, embodiments to which the present technology is applied will be described with reference to the drawings.

[0016] First Embodiment About the Present Technology The present technology realizes free viewpoint content with artistic quality.

[0017] For example, the present technology has the following features T1 to T4.

[0018] (Feature T1) A space (area) in which the listener is prohibited from being present, centered on a TP (target point) in three-dimensional space, is predetermined as a prohibited space. (Feature T2) If a listener attempts to enter the prohibited space, a signal indicating that entry is prohibited is generated, an alternative listener position is determined, and object metadata for the alternative listener position is generated. (Feature T3) When object metadata for the alternative listener position is generated, the alternative listener position is set to the position on the surface of the boundary of the prohibited space that is closest to the listener position, and the listener is forcibly moved to the alternative listener position. (Feature T4) If the listener position overlaps the TP, the object metadata generated for the listener position immediately before overlapping with the TP is used.

[0019] Now, the present technology will be described below.

[0020] In an audio playback system using this technology, users can watch content by rotating their head up, down, left, right, or diagonally within the space, or by moving to any position within the space. This type of content is called 6DoF content or free viewpoint content.

[0021] The content described below may be, for example, audio content consisting of audio only, or content consisting of video and accompanying audio, but hereinafter these types of content will not be distinguished and will be simply referred to as content. Also, hereinafter, audio objects will also be simply referred to as objects.

[0022] In the content production process, in order to enhance the artistic (musical) quality, objects are sometimes intentionally placed in locations other than their visible locations, without being bound by the physical placement of the objects.

[0023] Specifically, in the case of a band performance, for example, rather than physically placing objects (audio objects) corresponding to the vocals or guitar at the positions of those vocals or guitar, the objects are placed in positions that take into account the musicality of the music. Taking into account the musicality means making the music easy to listen to.

[0024] As an example, rather than placing a monaural object directly at the physical position of a guitarist in three-dimensional space, objects are placed at two positions on the left and right in front of the user, based on musical knowledge. By placing objects with a certain width at each of the positions on the left and right in front of the user in this way, an audio presentation that envelops the listener (user) can be realized.

[0025] As described above, a creator creates free viewpoint content by arranging each object in a space based on musicality.

[0026] In three-dimensional space (hereinafter also referred to as free viewpoint space), the content creator determines positions that will be used as multiple control viewpoints (hereinafter also referred to as CVPs (Control Viewpoints)) and a position that will be used as one target point (hereinafter also referred to as TP (Target Point)).

[0027] A CVP is the position of a viewpoint that you want to represent within a free viewpoint space. For example, a CVP may be the desired listening position when playing content. Hereinafter, the i-th CVP will also be referred to as CVPi.

[0028] The TP (target point) is the reference position for interpolation processing of object position information, and in particular, all virtual listeners in each CVP are assumed to be facing the TP. The actual listener (user) in the free viewpoint space can also move freely to any position, but is assumed to always face the TP. In other words, the listener can move within the free viewpoint space while facing the TP.

[0029] The positions of these CVPs and TPs are expressed by absolute coordinates that indicate absolute positions within the free viewpoint space. For example, if the coordinate system of absolute coordinates that indicate absolute positions within the free viewpoint space is called a common absolute coordinate system, the common absolute coordinate system is an orthogonal coordinate system with an origin O at a specific position within the free viewpoint space and mutually orthogonal X-, Y-, and Z-axes as its axes.

[0030] The placement position of objects is determined for each CVP. For example, in each CVP, a polar coordinate space (hereinafter also referred to as the CVP polar coordinate space) is created with the CVP position as its center, and each object is placed in that CVP polar coordinate space.

[0031] Specifically, a position in the CVP polar coordinate space is expressed by the coordinates (polar coordinates) of a polar coordinate system (hereinafter also referred to as the CVP polar coordinate system) consisting of mutually orthogonal x-, y-, and z-axes, with the position of the CVP, i.e., the absolute position of the CVP in the free viewpoint space (common absolute coordinate system), as the origin O'.

[0032] In particular, the direction from the CVP position toward the TP is the positive y-axis, i.e., the front direction of the virtual listener in the CVP, the left-right axis as seen from the virtual listener in the CVP is the x-axis, and the up-down axis as seen from the virtual listener in the CVP is the z-axis.

[0033] Incidentally, the listener can move to any position within the free viewpoint space.

[0034] When playing back content, the playback system may display a three-dimensional user interface (hereinafter also referred to as a free viewpoint space UI) based on the position and orientation of the listener in the free viewpoint space.

[0035] For example, the free viewpoint space UI may be an image related to the free viewpoint space, such as an image of a three-dimensional space that imitates the free viewpoint space. Note that the free viewpoint space UI may also be a video (image) of the content being played.

[0036] As a specific example, the free viewpoint space UI may be an image of the entire free viewpoint space that displays the position and orientation of the listener's face, or an image of the free viewpoint space as viewed from the listener's position in the direction of the listener's face. In either case, the free viewpoint space UI is generated based on the position and orientation of the listener in the free viewpoint space.

[0037] In addition, it is assumed that the listener in the free viewpoint space moves while always looking at the TP. In such a case, the position and facial direction of the listener are used to generate and display the free viewpoint space UI, and the facial direction of the listener can be uniquely determined.

[0038] For example, let us assume that a predetermined position on the stage in the free viewpoint space is TP, and the position of TP is expressed by the coordinates (TpX, TpY, TpZ) of the common absolute coordinate system, as shown in Figure 1. In Figure 1, O indicates the origin of the common absolute coordinate system.

[0039] It is also assumed that the listener is at position O'' in the free viewpoint space, and the position of the listener in the free viewpoint space (position O'') is represented by coordinates (LpX, LpY, LpZ) in the common absolute coordinate system.

[0040] Now, let us say that the polar coordinate space has the listener position (position O'') as its origin (hereinafter also referred to as origin O'') is the listener polar coordinate space. Furthermore, positions within this listener polar coordinate space are expressed by coordinates (polar coordinates) in a polar coordinate system (hereinafter also referred to as listener polar coordinate system) that has position O'' as its origin and is made up of mutually orthogonal x-, y-, and z-axes.

[0041] In such a case, the orientation of the listener (orientation of the face) in the free viewpoint space can be expressed by a horizontal angle Yaw indicating the orientation in the horizontal direction and a vertical angle Pitch indicating the orientation in the vertical direction.

[0042] If the X'Y' plane is defined as a plane that includes the origin O'' of the listener, that is, the listener polar coordinate system, and is parallel to the XY plane that includes the X and Y axes of the common absolute coordinate system, then line LN11 is a line obtained by projecting the y axis of the listener polar coordinate system onto the X'Y' plane. Also, line LN12 is a line on the X'Y' plane that is parallel to the Y axis of the common absolute coordinate system.

[0043] At this time, the angle formed by the lines LN11 and LN12 is the horizontal angle Yaw indicating the horizontal orientation of the listener, and the angle formed by the y-axis and the line LN11 is the vertical angle Pitch indicating the vertical orientation of the listener.

[0044] These horizontal angle Yaw and vertical angle Pitch can be found by calculating the following equations (1) and (2) from the coordinates of the TP (TpX, TpY, TpZ) and the coordinates of the listener's position (LpX, LpY, LpZ).

[0045]

[0046]

[0047] However, when calculating Tmp1 in equation (1) or Tmp2 in equation (2), if the TP and listener positions have the same value, the denominator values ​​(numbers) of Tmp1 and Tmp2 will be 0, and the values ​​of Tmp1 and Tmp2 will become infinite.

[0048] If this happens, subsequent calculations become impossible, the horizontal angle Yaw and the vertical angle Pitch cannot be obtained, and as a result, the free viewpoint space UI cannot be generated and displayed, which means that the free viewpoint space UI cannot be displayed appropriately.

[0049] Therefore, in this technology, a space within the free viewpoint space into which the listener (user) is prohibited from entering, that is, a prohibited space into which the user cannot enter, is defined.

[0050] By defining such a prohibited space and prohibiting entry into that prohibited space, it is possible to prevent the listener's orientation from being unable to be calculated, and to always display the free viewpoint space UI appropriately. Defining a prohibited space also makes it possible to reflect the intentions of content creators, such as not wanting people to approach specific positions such as TPs or specific objects.

[0051] The prohibited space is a space of any shape that includes the TP.

[0052] Specifically, the prohibited space is a spherical space as shown in Fig. 2. In this example, one position in the rectangular parallelepiped free viewpoint space FVS is set as TP.

[0053] A sphere P11 having a predetermined radius and centered at TP is set as a prohibited space. In other words, the center position of the sphere P11, which is set as a prohibited space, is set as TP.

[0054] The listener can move freely within the free viewpoint space FVS while viewing the content, but cannot move (enter) inside the sphere P11, which is a prohibited space.

[0055] The prohibited space is not limited to a spherical space (sphere), but may be a space of any shape, such as a cube, a rectangular parallelepiped, a polyhedron, or a plane, as long as it contains the TP. In particular, the TP itself may be set as the prohibited space. Furthermore, the prohibited space does not have to be a space centered on the TP. In other words, the position of the TP is not limited to the center position of the prohibited space, and may be any position within the prohibited space. In addition, a space that includes the TP and a specific object, such as an object placed in the free viewpoint space that the listener is not desired to approach, or an object placed at a fixed position in the free viewpoint space regardless of the viewpoint position, may be set as the prohibited space.

[0056] In this embodiment, the description will be given assuming that the prohibited space is a spherical space, and the spherical space that is particularly designated as the prohibited space will also be referred to as a prohibited space sphere.

[0057] Configuration information of the format shown in FIG. 3, for example, is supplied from the server side that distributes the content to the playback side that plays back the content.

[0058] That is, Fig. 3 shows an example of the bitstream format of the configuration information. In more detail, the configuration information includes other information in addition to the information shown in Fig. 3, such as information about CVP, which will be described later, but this information is not shown in Fig. 3.

[0059] In the example of FIG. 3, the configuration information includes a prohibited space information presence flag "TPProhibitPresentFlag" that indicates whether or not a definition of a prohibited space exists, that is, whether or not prohibited space information exists.

[0060] Here, when the value of the prohibited space information existence flag is 1, it indicates that prohibited space information for identifying a prohibited space exists, i.e., prohibited space information is included in the configuration information. In contrast, when the value of the prohibited space information existence flag is 0, it indicates that prohibited space information does not exist, and that prohibited space information is not included in the configuration information.

[0061] When the value of the prohibited space information existence flag is 1, the configuration information stores radius information "TPProhibitRadius" indicating the radius of the prohibited space sphere as prohibited space information.

[0062] For example, the value of the radius indicated by the radius information is the real value of the radius of the forbidden space sphere when normalized with the longest side of a rectangular parallelepiped defined as the three-dimensional free viewpoint space set to 1.0.

[0063] The radius information may be a value set by the content production tool, or may be set by the playback side, or may be rewritable. For example, the radius information may be changed on the playback side according to object metadata such as the priority of an object, the size of an object, or the size of the object's area indicated by spread information. For example, if object metadata indicating the priority of an object is specified as a value from 0 to 7, the radius of the prohibited space sphere may be dynamically changed according to the priority value. Furthermore, the radius of the prohibited space sphere may be dynamically changed according to the resources, remaining battery power, etc. on the playback side.

[0064] Alternatively, it may be determined in advance that a prohibited space will be set, and only the radius information may be stored in the configuration information without storing a prohibited space information presence flag. Furthermore, if the value of the radius information is determined in advance, it may be determined that only the prohibited space information presence flag is stored in the configuration information without storing the radius information. Alternatively, the prohibited space may be determined independently on the content playback side.

[0065] On the content playback side, the position of the listener in the free viewpoint space is adjusted (changed) based on the configuration information, etc., so that the listener does not enter the prohibited space. In other words, the final position of the listener is determined.

[0066] For example, a position in the listener's free viewpoint space that is input by the user (listener) or determined according to the output of a sensor attached to the user and is input on the content playback side is also referred to as an input listener position, and information indicating this input listener position is referred to as input listener position information. Furthermore, head tracking information and eye tracking information may be input using a sensor mounted on an HMD (Head Mounted Display) used in AR (Augmented Reality) or VR (Virtual Reality), for example.

[0067] Furthermore, information indicating the final position of the listener in the free viewpoint space, which is determined by the content playback system from input listener position information and configuration information, etc., will be referred to as listener position information.

[0068] In this technology, for example, if the input listener position is within a prohibited space, the position of the listener is changed. That is, a position outside the prohibited space that is different from the input listener position is set as the final listener position. In this case, the final listener position (the changed listener position) is set as, for example, a position on the surface of the boundary of the prohibited space that is closest to the input listener position. Alternatively, a position outside the prohibited space that is different from a position on the surface of the boundary of the prohibited space and that is closest to the input listener position may be set as the final listener position.

[0069] Here, a specific example of a method for determining the final listener position will be described.

[0070] First, on the playback side, the distance Dist from the listener position (input listener position) to the TP is calculated.

[0071] Specifically, the coordinates indicating the input listener position U in the free viewpoint space, i.e., the common absolute coordinate system, are assumed to be (UpX, UpY, UpZ), and the coordinates indicating the position of TP in the common absolute coordinate system are assumed to be (TpX, TpY, TpZ). In this case, the direction vector (LVx, LVy, LVz) starting from TP and ending at the input listener position is expressed by the following equation (3).

[0072]

[0073] In addition, the distance Dist can be calculated by calculating the magnitude (length) of the direction vector (LVx, LVy, LVz) using the following equation (4).

[0074]

[0075] For example, consider the case shown in Figure 4 where there are positions A, B, and C, and position A or C is the input listener position. Note that in Figure 4, parts corresponding to those in Figure 2 are given the same reference numerals, and their explanation will be omitted where appropriate. Also, in Figure 4, sphere P11, which is a prohibited space sphere, is drawn large to make the figure easier to see.

[0076] 4, position A is located inside sphere P11, position B is located on the surface of sphere P11, and position C is located outside sphere P11. Also, assume that the coordinates of positions A, B, and C in the free viewpoint space FVS, i.e., the common absolute coordinate system, are (Ax, Ay, Az), (Bx, By, Bz), and (Cx, Cy, Cz).

[0077] For example, the distance Dist when the input listener position is position C outside the sphere P11 will be specifically referred to as distance DistCT, and the distance Dist when the input listener position is position A inside the sphere P11 will be specifically referred to as distance DistAT.

[0078] In this case, the distance DistCT is greater than the radius of the sphere P11, that is, the radius indicated by the radius information TPProhibitRadius, and therefore it is clear that position C, which is the input listener position, is outside the sphere P11, which is the prohibited space.

[0079] Therefore, in such a case, the input listener position can be used as the final listener position to generate object metadata and a free viewpoint space UI.

[0080] On the other hand, because the distance DistAT is smaller than the radius of the sphere P11, i.e., the radius indicated by the radius information TPProhibitRadius, it is clear that position A, which is the input listener position, is inside (inside) the sphere P11, which is the prohibited space. In other words, it is clear that the listener has entered the prohibited space.

[0081] Therefore, in such a case, the listener is forcibly moved to a position on the surface of the prohibited space sphere (sphere P11) using, for example, the technique described below, and the position after the movement is set as the final listener position. Then, based on the final listener position, object metadata is generated and a free viewpoint space UI is generated.

[0082] An example of a specific method for determining the destination of the listener (the changed listener position) will be described.

[0083] First, for position A, which is the input listener position shown in Fig. 4, the direction vector (LVx, LVy, LVz) connecting TP and position A is found by equation (3). In this case, UpX = Ax, UpY = Ay, UpZ = Az.

[0084] Next, the intersection of the direction vector, more specifically, the line including the direction vector, and the surface of the prohibited space sphere (sphere P11) is found. In this case, there are two intersections.

[0085] These intersections are designated XP1 and XP2, and the coordinates indicating the absolute positions of intersections XP1 and XP2 in the free viewpoint space FVS, i.e., the common absolute coordinate system, are designated as (XP1x, XP1y, XP1z) and (XP2x, XP2y, XP2z).

[0086] In this case, the coordinates (XP1x, XP1y, XP1z) of the intersection point XP1 can be obtained by the following formula (5), and the coordinates (XP2x, XP2y, XP2z) of the intersection point XP2 can be obtained by the following formula (6). A direction vector (LVx, LVy, LVz) is used in the calculations of formulas (5) and (6).

[0087]

[0088]

[0089] In addition, in equations (5) and (6), LvecPow is the same as LvecPow in equation (4), and tppr is the radius of the prohibited space sphere (sphere P11) indicated by the radius information TPProhibitRadius.

[0090] After the two intersection points XP1 and XP2 are found, the one closest to the input listener position is selected as the destination position of the listener, i.e., the final listener position after the movement. In other words, the intersection point that is closer to the input listener position is selected as the final listener position.

[0091] Specifically, if the distance from position A, which is the input listener position, to intersection XP1 is DistX1 and the distance from position A to intersection XP2 is DistX2, these distances DistX1 and DistX2 can be obtained by the following equation (7).

[0092]

[0093] In equation (7), the coordinates (UpX, UpY, UpZ) of the input listener position are (Ax, Ay, Az), i.e., UpX=Ax, UpY=Ay, UpZ=Az.

[0094] Based on the distances DistX1 and DistX2 thus obtained, the following equation (8) is calculated, and either the intersection XP1 or the intersection XP2 is determined as the final listener position.

[0095]

[0096] In equation (8), the coordinates of the final listener position are (FinalUPx, FinalUPy, FinalUPz), and the coordinates of the listener position determined in the previous processing, i.e., the immediately preceding listener position, are (PrevUPx, PrevUPy, PrevUPz).

[0097] In the calculation of equation (8), the intersection point with the smaller distance value is selected as the final listener position. That is, if the distance DistX2 is smaller than the distance DistX1, the intersection point XP2 is selected as the final listener position, and if the distance DistX1 is smaller than the distance DistX2, the intersection point XP1 is selected as the final listener position.

[0098] Furthermore, if the distance DistX1 and the distance DistX2 are the same, the listener position of the previous process, that is, the intersection point with the shorter distance from the immediately preceding listener position, is determined to be the final listener position.

[0099] Specifically, the distance DistPX1 between the intersection point XP1 and the immediately preceding listener position, and the distance DistPX2 between the intersection point XP2 and the immediately preceding listener position are calculated in the same manner as in equation (7).

[0100] If the distance DistPX2 is smaller than the distance DistPX1, the intersection XP2 is set as the final listener position, and if the distance DistPX1 is smaller than the distance DistPX2, the intersection XP1 is set as the final listener position.

[0101] For example, in the example shown in Figure 4, when the input listener position is position A, the final listener position is position B, which is the intersection point closest to position A among the points where the line passing through TP and position A, i.e., the direction vector, intersects with the surface of the prohibited space sphere (sphere P11).

[0102] In the example of Figure 4, the prohibited space sphere (sphere P11) is a spherical space centered at TP, so the final listener position, i.e., position B, which is the listener position after the change, is the position on the surface of the prohibited space sphere that is closest to position A, which is the listener position before the change.

[0103] In addition, the position of the listener after movement when the listener enters the prohibited space, i.e., the final listener position, may be any position outside the prohibited space, not limited to a position on the boundary part of the prohibited space (a position on the boundary), such as a position on the surface of the prohibited space sphere.

[0104] As described above, when the input listener position is within the prohibited space sphere, the listener position is moved to a position on the surface of the prohibited space sphere instead of the input listener position, and the position after this movement, i.e., the position indicated by the coordinates (FinalUPx, FinalUPy, FinalUPz), is set as the final listener position (current listener position). Then, based on this final listener position, object metadata is generated and a free viewpoint space UI is generated.

[0105] <Server Configuration Example> FIG. 5 is a diagram showing a configuration example of an embodiment of a server to which the present technology is applied.

[0106] The server 11 is, for example, an information processing device, and functions as an encoder that encodes and distributes content.

[0107] The server 11 includes a meta-encoder 21 , an audio encoder 22 , and a communication unit 23 .

[0108] The meta-encoder 21 is supplied with metadata of one or more objects that make up the content, that is, metadata of the audio data of the object (hereinafter referred to as object metadata), and configuration information.

[0109] Object metadata is prepared for each object for each CVP (control viewpoint). The object metadata includes at least object position information that indicates the position of the object and object gain, which is the gain of the object's audio data.

[0110] The object position information indicates the position of an object in the CVP polar coordinate space, with the CVP position as the reference (origin), and the object position is described by the coordinates (polar coordinates) of the CVP polar coordinate system. Therefore, the position of an object indicated by the object position information is the position of the object as seen from the CVP.

[0111] In the server 11, object metadata is prepared for each CVP, so even for the same object, the object's placement position in the free viewpoint space (common absolute coordinate system) and object gain value may differ for each CVP.

[0112] The object metadata may include priority information indicating the priority of the object, sound source type information indicating the type of the object (sound source), and spread information indicating the extent of the object's spread.

[0113] The configuration information includes, for example, at least one of the prohibited space information presence flag and the radius information described above.

[0114] The configuration information also includes, as appropriate, object number information indicating the number of objects that make up the content, CVP number information indicating the number of CVPs prepared in advance, CVP information related to the CVPs, etc. In the following, the description will be given assuming that the configuration information includes a prohibited space information presence flag, radius information, object number information, CVP number information, and CVP information.

[0115] For example, the CVP information includes a CVP index, CVP position information, and CVP orientation information.

[0116] The CVP index is ID information that uniquely identifies the CVP. The CVP position information is position information that indicates the absolute position of the CVP in the free viewpoint space, and this absolute position of the CVP is described by the coordinates of the common absolute coordinate system.

[0117] The CVP orientation information indicates the orientation of the face of a virtual listener in the CVP in the free viewpoint space (common absolute coordinate system). In other words, the CVP orientation information indicates the orientation of the CVP, more specifically, the positive direction of the y-axis of the CVP polar coordinate system in the free viewpoint space.

[0118] For example, the CVP orientation information consists of CVP Yaw information, which is a horizontal angle indicating the horizontal position of the TP as seen from the position of the CVP in the free viewpoint space, and CVP Pitch information, which is a vertical angle indicating the vertical position of the TP as seen from the position of the CVP in the free viewpoint space.

[0119] For example, in the example shown in FIG. 1, if the CVP is at position O'', the angle (Yaw) between the lines LN11 and LN12 becomes the CVP Yaw information, and the angle (Pitch) between the y-axis and the line LN11 becomes the CVP Pitch information.

[0120] The CVP orientation information may include not only CVP Yaw information and CVP Pitch information, but also CVP Roll information that indicates the rotation angle (Roll) of the CVP polar coordinate system relative to the common absolute coordinate system, with the y-axis of the CVP polar coordinate system as the rotation axis.

[0121] On the content playback side, if there is CVP information for multiple CVPs, it is possible to identify the position of the TP in the free viewpoint space (common absolute coordinate system) from that CVP information. In this case, if a line that passes through the CVP and points in the direction indicated by the CVP orientation information is found for each CVP, and the position of the intersection of those lines is found, the position of that intersection becomes the position of the TP.

[0122] Note that instead of the CVP orientation information, TP position information, which is the coordinates (absolute coordinates) indicating the position of the TP in the common absolute coordinate system, may be transmitted to the content playback side. In such a case, the playback side can obtain the CVP orientation information from the TP position information and CVP position information included in the configuration information.

[0123] The meta-encoder 21 encodes (multiplexes) the object metadata for each CVP of each object supplied, and supplies the resulting multiplexed object metadata and the supplied configuration information to the audio encoder 22.

[0124] The audio data of each object for reproducing the content, i.e., the audio data for reproducing the sound of each object, is supplied to the audio encoder 22. The audio encoder 22 encodes the supplied audio data of each object to generate encoded audio data.

[0125] In addition, the audio encoder 22 multiplexes the encoded audio data obtained by encoding with the multiplexed object metadata and configuration information supplied from the meta encoder 21 to generate an encoded bitstream, and supplies it to the communication unit 23.

[0126] The communication unit 23 transmits (outputs) the coded bit stream supplied from the audio encoder 22 to a client, which is a device on the playback side, via a network or the like.

[0127] Here, an example will be described in which object metadata is stored for each CVP in the coded bitstream, but this is not limiting and it is sufficient if the object metadata for each CVP can be obtained on the client side.

[0128] For example, it is conceivable to prepare in advance a relative arrangement pattern of each object as seen from a virtual listener (CVP) and a metadata set that is a set of object metadata of each object for each arrangement pattern.

[0129] In this case, for example, on the server 11 side, an appropriate metadata set is selected from among multiple metadata sets (arrangement patterns) for each CVP, and a metadata set index indicating the selected metadata set is stored in the CVP information.

[0130] In this way, if the server 11 and the client share a group of metadata sets by some means, the client can obtain the appropriate object metadata for each object for each CVP from the metadata set index included in the CVP information.

[0131] <Description of Encoding Process> A description will be given of the operation of the server 11. That is, the encoding process performed by the server 11 will be described below with reference to the flowchart of FIG.

[0132] In step S11, the meta-encoder 21 multiplexes (encodes) the object metadata for each CVP of each object provided to generate multiplexed object metadata. The meta-encoder 21 also provides the multiplexed object metadata and the provided configuration information to the audio encoder 22. Note that the configuration information may also be encoded as appropriate.

[0133] In step S12, the audio encoder 22 encodes the audio data of each object supplied thereto to generate encoded audio data.

[0134] For example, the audio encoder 22 encodes the audio data in accordance with an encoding method used in international standards such as MPEG (Moving Picture Experts Group)-I and MPEG-H 3D Audio. Note that any encoding method may be used for encoding the audio data.

[0135] In step S13, the audio encoder 22 multiplexes the encoded audio data obtained by encoding with the multiplexed object metadata and configuration information supplied from the meta encoder 21 to generate an encoded bitstream and supplies it to the communication unit 23.

[0136] In step S14, the communication unit 23 transmits the encoded bitstream supplied from the audio encoder 22 to the client, and the encoding process ends.

[0137] In this manner, the server 11 generates an encoded bitstream including configuration information in which a prohibited space information presence flag and radius information are stored, and transmits the encoded bitstream to the client.

[0138] By doing this, the client side can use the configuration information to process the prohibited space when playing back the content, and can appropriately display the content. In other words, the free viewpoint space UI can be appropriately displayed.

[0139] <Configuration Example of Client> FIG. 7 is a diagram showing a configuration example of a client that receives (acquires) the coded bitstream transmitted (output) from the server 11 and performs processing related to the playback of the content.

[0140] The client 61 shown in FIG. 7 is an information processing device such as a personal computer, a smartphone, a tablet, a head-mounted display, or a game device, and functions as a decoder.

[0141] The client 61 performs decoding of the coded bit stream received from the server 11, generates output audio data for playing back the sound of the content, and supplies the generated output audio data to the audio output unit 62. The client 61 also generates image data (video data) of a free viewpoint space UI related to the content based on the configuration information and the like, and supplies the generated image data to the presentation unit 63.

[0142] The audio output unit 62 is, for example, a speaker system consisting of a plurality of speakers, and outputs (plays) the sound of the content based on the output audio data supplied from the client 61. The audio output unit 62 may also be headphones, earphones, a hearing aid, a sound collector, or the like.

[0143] The presentation unit 63 includes a display or the like, and displays a free viewpoint space UI based on image data supplied from the client 61. The presentation unit 63 may also include a vibration motor or the like that presents information to the user (listener) by vibration or the like. The audio output unit 62 and the presentation unit 63 may also be provided in the client 61.

[0144] The client 61 includes a communication unit 71 , an audio decoder 72 , a listener information acquisition unit 73 , a meta decoder 74 , a rendering processing unit 75 , and a presentation control unit 76 .

[0145] The communication unit 71 receives (acquires) the coded bitstream transmitted from the server 11 and supplies it to the audio decoder 72 .

[0146] The audio decoder 72 demultiplexes the coded bitstream supplied from the communication unit 71 to extract coded audio data, multiplexed object metadata, and configuration information, and decodes the coded audio data to obtain audio data.

[0147] The audio decoder 72 supplies the audio data of each object obtained by decoding to a rendering processing unit 75 , and also supplies the multiplexed object metadata and configuration information to a meta decoder 74 .

[0148] The listener information acquisition unit 73 acquires input listener position information indicating an input listener position, which is the input position of the listener in a free viewpoint space, in response to an operation by the user (listener) on an input unit (not shown), and supplies the information to the meta decoder 74. Here, in addition to the listener position, the listener information acquisition unit 73 may also acquire, for example, head tracking information (head rotation information) and eye tracking information (gaze detection information) from a sensor mounted on an HMD used in AR, VR, or the like.

[0149] The meta decoder 74 demultiplexes (decodes) the multiplexed object metadata supplied from the audio decoder 72, and obtains object metadata for each CVP (control viewpoint) of each object.

[0150] In addition, the meta decoder 74 determines the final listener position based on the configuration information and object metadata supplied from the audio decoder 72 and the input listener position information supplied from the listener information acquisition unit 73, and generates listener reference object metadata.

[0151] The metadecoder 74 functions as a position determination unit that determines whether the listener (user) position is within a prohibited space, changes the listener position as appropriate based on the determination result, and generates listener reference object metadata.

[0152] The listener reference object metadata is object metadata of the object at the final position of the listener in the free viewpoint space, and the meta decoder 74 supplies the generated listener reference object metadata to the rendering processing unit 75 .

[0153] The listener-reference object metadata includes listener-reference object position information that indicates the relative position of the object as seen from the final listener position, and listener-reference object gain, which is the gain of each object relative to the final listener position. In other words, the listener-reference object metadata is metadata of the object when the listener is at the position determined by the meta decoder 74, and is generated for each object.

[0154] The listener reference object metadata may also include information about the priority of the object, information about the type of sound source, and spread information.

[0155] Furthermore, when generating the listener reference object metadata, the meta decoder 74 also generates listener position information indicating the final position of the listener in the free viewpoint space, a TP warning flag indicating whether the specified (input) input listener position is within a prohibited space, and TP position information, and supplies (outputs) these to the presentation control unit 76. For example, the listener position indicated by the listener position information is described using coordinates (absolute coordinates) in a common absolute coordinate system.

[0156] Here, the TP warning flag is flag information (warning information) that indicates whether or not the input listener position is within a prohibited space, and an example will be described in which the TP warning flag is always output to the presentation control unit 76. However, the present invention is not limited to this, and warning information to that effect may be output to the presentation control unit 76 only when the input listener position is within a prohibited space. Note that, in addition to presenting characters (messages), colors, images, etc. based on the warning information, notification based on the warning information may also be made by, for example, sound or vibration.

[0157] The rendering processing unit 75 performs rendering processing based on the audio data of each object supplied from the audio decoder 72 and the listener reference object metadata of each object supplied from the meta decoder 74, and generates output audio data. The rendering processing unit 75 supplies (outputs) the generated output audio data to the audio output unit 62, which plays back the sound of the content.

[0158] The rendering processing unit 75 performs rendering processing in a polar coordinate system defined by MPEG-H 3D Audio, such as processing using VBAP (Vector Based Amplitude Panning), to generate output audio data. This output audio data is audio data for playing back the sounds of the free viewpoint content, including the sounds of all objects.

[0159] The rendering processing unit 75 may perform convolution processing such as HRTF (Head Related Transfer Function), BRIR (Binaural Room Impulse Response), RIR (Room Impulse Response), ITD (Interaural Time Difference), IID (Interaural Intensity Difference), or processing using HOA (Higher Order Ambisonics).

[0160] The presentation control unit 76 is realized by, for example, a higher-level system of the system that realizes the meta-decoder 74, etc., and controls the display of various images (videos) on the presentation unit 63. Furthermore, for example, the presentation control unit 76 may control the presentation of information by vibration on the presentation unit 63, or the presentation of information by sound on the audio output unit 62.

[0161] For example, the presentation control unit 76 generates image data (video data) of the free viewpoint space UI based on the listener position information, TP warning flag, TP position information, etc. supplied from the meta decoder 74 , and supplies it to the presentation unit 63 .

[0162] <Description of Output Audio Data Generation Process> A description will now be given of the operation of the client 61. That is, the output audio data generation process performed by the client 61 will be described below with reference to the flowchart of FIG.

[0163] In step S 41 , the communication unit 71 receives (acquires) the coded bit stream transmitted from the server 11 and supplies it to the audio decoder 72 .

[0164] In step S 42 , the audio decoder 72 decodes the coded bit stream supplied from the communication unit 71 .

[0165] Specifically, the audio decoder 72 demultiplexes the coded bitstream to extract coded audio data, multiplexed object metadata, and configuration information, and decodes the coded audio data to obtain the audio data.

[0166] The audio decoder 72 supplies the audio data of each object obtained by decoding to a rendering processing unit 75 , and also supplies the multiplexed object metadata and configuration information to a meta decoder 74 .

[0167] In step S43, the listener information acquisition unit 73 acquires input listener position information indicating the input listener position input by the listener in response to an input operation by the user (listener), and supplies the information to the meta decoder 74. For example, the input listener position information may be coordinates (absolute coordinates) indicating the absolute input listener position in a free viewpoint space, i.e., a common absolute coordinate system.

[0168] In step S 44 , the meta decoder 74 generates listener reference object metadata based on the configuration information and multiplexed object metadata supplied from the audio decoder 72 and the input listener position information supplied from the listener information acquisition unit 73 .

[0169] For example, the meta decoder 74 determines the final listener position based on the prohibited space information presence flag, radius information, and CVP information contained in the configuration information, and the input listener position information, and generates listener position information indicating the position of the listener.

[0170] As will be described in detail later, when the input listener position is outside the prohibited space (but includes the boundary surface of the prohibited space), the input listener position is used as the final listener position. In other words, the listener position is not changed.

[0171] On the other hand, when the input listener position is within the prohibited space (inside the prohibited space), a position outside the prohibited space that is different from the input listener position is set as the final listener position. In other words, the listener position is changed to a position outside the prohibited space that is different from the input listener position.

[0172] The meta decoder 74 also generates a TP warning flag and TP position information in the process of generating the listener position information, and supplies the listener position information, TP warning flag, and TP position information to the presentation control unit 76 .

[0173] For example, if the listener enters a prohibited space, that is, if the input listener position is within the prohibited space (if a position within the prohibited space is specified as the input listener position), the value of the TP warning flag is set to "true." The "true" value of the TP warning flag indicates that the listener has entered a prohibited space, or more specifically, is about to enter a prohibited space.

[0174] On the other hand, for example, if the listener does not enter the prohibited space, that is, if the input listener position is outside the prohibited space (if a position outside the prohibited space is specified as the input listener position), the value of the TP warning flag is set to "false." The value "false" of the TP warning flag indicates that the listener is outside the prohibited space.

[0175] Supplying such a TP warning flag to the presentation control unit 76 can also be said to be a notification that the listener position has been moved from the input listener position, that is, that the listener position has been changed.

[0176] The metadecoder 74 generates listener reference object metadata for each object based on the listener location information, the CVP information contained in the configuration information, and the object metadata at each CVP.

[0177] The listener reference object metadata includes listener reference object position information and listener reference object gain. For example, the position of an object indicated by the listener reference object position information is a position within the listener polar coordinate space and is described by coordinates (polar coordinates) in the listener polar coordinate system.

[0178] The details of generating the listener reference object metadata will be described later. For example, in step S44, the listener reference object metadata is generated by performing an interpolation process such as vector synthesis.

[0179] At this time, object metadata of all CVPs or several CVPs around the listener position is used, and interpolation processing is performed according to the positional relationship between the listener position and the CVPs identified from the CVP information and listener position information, and listener reference object metadata is generated.

[0180] The meta decoder 74 supplies the generated listener reference object metadata to the rendering processing unit 75 .

[0181] In step S 45 , the rendering processing unit 75 performs rendering processing based on the audio data of each object supplied from the audio decoder 72 and the listener reference object metadata supplied from the meta decoder 74 .

[0182] For example, the rendering processing unit 75 performs gain correction on the audio data of each object based on the listener reference object gain of each object.

[0183] The rendering processing unit 75 then performs rendering processing such as VBAP based on the audio data of each object after gain correction and the listener reference object position information, and generates output audio data.

[0184] The rendering processing unit 75 outputs the generated output audio data to the audio output unit 62, which plays back the sound of the content. In this way, it becomes possible to play back content in which the listener (listening position) is any position outside the prohibited space within the free viewpoint space, i.e., free viewpoint content with multiple viewpoints (6DoF content).

[0185] The presentation control unit 76 also generates image data for the free viewpoint space UI based on the listener position information, TP warning flag, TP position information, etc. supplied from the meta decoder 74, and supplies the generated image data to the presentation unit 63, thereby displaying the free viewpoint space UI. The free viewpoint space UI may be generated using CVP information, listener reference object metadata, radius information, etc., as necessary.

[0186] The presentation control unit 76 calculates the direction of the listener's face in the free viewpoint space based on the listener position information and the TP position information, and also uses the calculation result to generate the free viewpoint space UI. In other words, the presentation control unit 76 generates a free viewpoint space UI, which is an image related to the free viewpoint space, based on the listener position information, the direction of the listener's face in the free viewpoint space, and the TP warning flag, and displays it on the presentation unit 63. The direction of the listener's face can be obtained by calculations similar to those of the above-mentioned equations (1) and (2).

[0187] For example, an image of the entire or part of the free viewpoint space can be displayed as the free viewpoint space UI. In this case, TPs, prohibited spaces, objects, etc. can be displayed on the free viewpoint space UI.

[0188] In the free viewpoint space UI, the TP may be displayed in a conspicuous color such as red or may blink to discourage the listener from approaching the TP. Similarly, an area (space) that is a prohibited space may also be displayed in a specific color or may blink.

[0189] Furthermore, for example, if the value of the TP warning flag is "true," a message indicating that entry into the prohibited space is prohibited may be displayed on the free viewpoint space UI. The notification (notification control) by the presentation control unit 76 that entry into the prohibited space is prohibited is not limited to displaying a message, but may also be performed by flashing the TP or prohibited space, sounding an alarm, vibrating, or a combination thereof. The aforementioned change in color or flashing of the TP, or the warning notification when the value of the TP warning flag is "true," may be set automatically by the system, for example, or may be set arbitrarily by the user.

[0190] In step S46, the client 61 determines whether or not to end the process.

[0191] For example, in step S46, it is determined that the processing is to end when the encoded bitstream for all frames of the content has been received and output audio data has been generated, or when the listener (user) has instructed to end playback of the content.

[0192] If it is determined in step S46 that the process is not yet finished, the process returns to step S41, and the above-described process is repeated.

[0193] On the other hand, if it is determined in step S46 that the process is to be ended, the client 61 ends the operation of each unit, and the output audio data generation process ends.

[0194] In this way, the client 61 determines the final listener position based on the configuration information and input listener position information, and generates a free viewpoint space UI and output audio data according to the determination result.

[0195] This allows the free viewpoint space UI to be displayed appropriately. That is, the content can be displayed appropriately. Furthermore, it is possible to realize musical content playback based on the intentions of the content creator, according to the listener's position, rather than simply the physical relationship between the listener and the object, and to fully convey the enjoyment of the content to the listener.

[0196] <Explanation of Metadata Generation Process> While the output audio data generation process described with reference to FIG. 8 is being performed, the client 61 performs a metadata generation process shown in FIG. 9 as a process corresponding to the process of step S44.

[0197] The metadata generation process by the client 61 will be described below with reference to the flowchart of FIG.

[0198] In step S81, the meta-decoder 74 sets the TP warning flag to "false" and holds the TP warning flag. That is, the meta-decoder 74 temporarily sets the value of the TP warning flag to "false."

[0199] In step S82, the meta decoder 74 determines whether or not there is prohibited space information.

[0200] For example, the meta decoder 74 determines that prohibited space information is present if the configuration information supplied from the audio decoder 72 includes the prohibited space information presence flag "TPProhibitPresentFlag" described with reference to Fig. 3 and the value of the prohibited space information presence flag is "1." If the value of the prohibited space information presence flag is "1," the configuration information includes radius information "TPProhibitRadius" as prohibited space information.

[0201] If it is determined in step S82 that there is no prohibited space information, that is, if no prohibited space (prohibited space sphere) is set on the server 11 side, then the process proceeds to step S83.

[0202] In step S 83 , the meta decoder 74 generates listener reference object metadata at the current listener position, that is, the input listener position, and supplies it to the rendering processing unit 75 .

[0203] Specifically, the meta decoder 74 takes the input listener position indicated by the input listener position information supplied from the listener information acquisition unit 73, which is the current listener position, as the final listener position and generates listener position information indicating that listener position.

[0204] In addition, the meta decoder 74 generates listener reference object metadata for each object placed in the free viewpoint space by interpolation processing or the like, based on the listener position information, the CVP information included in the configuration information, and the object metadata for each CVP.

[0205] Furthermore, the metadecoder 74 identifies the position of the TP based on the CVP information of each CVP included in the configuration information, particularly the CVP position information and CVP orientation information included in the CVP information, and also generates TP position information indicating the position of the TP.

[0206] The meta decoder 74 supplies the listener reference object metadata to the rendering processing unit 75, and also supplies the listener position information, the TP warning flag whose value is "false", and the TP position information to the presentation control unit 76. Then, the process proceeds to step S91. Note that if it is determined in step S82 that there is no prohibited space information, prohibited space information (radius information) may be set on the client 61 side, and the processes of steps S84 to S90, which will be described later, may be performed.

[0207] If it is determined in step S82 that there is prohibited space information, the prohibited space has been set on the server 11 side, and the process proceeds to step S84.

[0208] In step S84, the metadecoder 74 calculates the distance UT between the input listener position, that is, the input listener position, and the TP, based on the input listener position information and configuration information.

[0209] Specifically, the meta decoder 74 calculates the TP position information based on the CVP information of each CVP included in the configuration information, and calculates the distance UT by performing calculations similar to those of the above-mentioned equations (3) and (4) based on the input listener position information and the TP position information. In this case, the distance Dist obtained by equations (3) and (4) corresponds to the distance UT.

[0210] In step S85, the meta decoder 74 compares the distance UT calculated in step S84 with the radius of the prohibited space sphere indicated by the radius information as prohibited space information included in the configuration information, and determines whether the distance UT is less than the radius of the prohibited space sphere.

[0211] This determination process can be said to be a process of identifying (determining) whether the position of the listener indicated by the input listener position information is within a prohibited space based on the input listener position information and the prohibited space information (radius information) and CVP information contained in the configuration information.

[0212] If it is determined in step S85 that the distance UT is not less than the radius of the prohibited space sphere, the input listener position is outside the prohibited space sphere and there is no need to change the listener position (forced movement of the listener), so processing then proceeds to step S83.

[0213] In this case, the meta decoder 74 does not change the listener position, but performs the above-described process in step S83, and then the process proceeds to step S91.

[0214] On the other hand, if it is determined in step S85 that the distance UT is less than the radius of the prohibited space sphere, the input listener position is inside the prohibited space sphere and a change in the listener position (forced movement of the listener) is necessary, and then processing proceeds to step S86.

[0215] In step S86, the meta decoder 74 sets the TP warning flag to "true" and holds the TP warning flag. That is, the meta decoder 74 changes the value of the TP warning flag it holds from "false" to "true."

[0216] In step S87, the meta decoder 74 calculates a direction vector connecting the current listener position, i.e., the input listener position, and the TP, based on the input listener position information and the TP position information. For example, the meta decoder 74 calculates the direction vector by performing a calculation similar to that of equation (3) above. Note that this direction vector is obtained during the calculation of the distance UT performed in step S84, so the direction vector obtained in step S84 may be retained and used as the processing result of step S87.

[0217] In step S88, the meta decoder 74 calculates the intersection between the direction vector and the surface of the prohibited space sphere based on the radius information, TP position information, and the direction vector obtained in step S87 as prohibited space information contained in the configuration information.

[0218] For example, the meta-decoder 74 calculates the two intersection points by performing calculations similar to those of the above-mentioned equations (5) and (6).

[0219] In step S89, the meta-decoder 74 selects the intersection closest to the input listener position from the two intersections calculated in step S88.

[0220] For example, the metadecoder 74 performs calculations similar to the above-described formulas (7) and (8) based on the input listener position information and the positions of the two intersections to identify the intersection closest to the input listener position. At this time, the metadecoder 74 also uses, as appropriate, the listener position information of the listener position immediately before the input listener position was input, i.e., the final listener position determined in the most recent step S83 or step S89.

[0221] Depending on the shape of the prohibited space, there may be three or more intersections between the direction vector and the surface of the prohibited space. In such cases, the intersection closest to the input listener position is selected and set as the changed listener position (final listener position). When the processing of step S89 is performed, it may also be the case that the distance UT was determined to be less than the radius of the prohibited space sphere in step S85 at the immediately preceding timing, such as the immediately preceding frame, and the processing of step S89 was performed. In such cases, the meta decoder 74 may set the listener position immediately before the input listener position was input, i.e., the final listener position determined in the most recent execution of step S89, as the final listener position after this change. In this case, if the user continues to input positions within the prohibited space sphere as the input listener position, the same position on the surface of the prohibited space sphere will continue to be set as the final listener position.

[0222] In step S90, the meta decoder 74 generates listener reference object metadata with the position of the intersection point selected in step S89 as the destination position of the listener, that is, the final listener position after the change, and supplies the generated listener reference object metadata to the rendering processing unit 75.

[0223] Specifically, the meta decoder 74 sets the position of the selected intersection as the final listener position and generates listener position information indicating that listener position. That is, the listener position is changed from the input listener position to the selected intersection position, and listener position information indicating the changed listener position is generated.

[0224] In addition, for each object placed in the free viewpoint space, the meta decoder 74 generates listener reference object metadata by interpolation processing or the like based on the listener position information, CVP information, and object metadata for each CVP, in the same manner as in step S83.

[0225] The meta decoder 74 supplies the listener reference object metadata to the rendering processing unit 75, and also supplies the listener position information, the TP warning flag whose value is "true," and the TP position information to the presentation control unit 76. Then, the process proceeds to step S91. Note that the TP position information calculated in step S84 may be used.

[0226] After the process of step S90 or step S83 is performed, the process of step S91 is performed.

[0227] In step S91, the meta decoder 74 determines whether there is data to process. For example, in step S91, if all object metadata of the content has been processed, it is determined that there is no data to process.

[0228] If it is determined in step S91 that there is data to be processed, the process returns to step S81, and the above-described process is repeated.

[0229] On the other hand, if it is determined in step S91 that there is no data to be processed, the metadata generation process ends.

[0230] In this way, the client 61 determines the final listener position depending on whether the input listener position is inside the prohibited space, and generates listener reference object metadata depending on the determination result. In this way, it is possible to prevent the listener orientation from being unable to be calculated, and the free viewpoint space UI can always be displayed appropriately.

[0231] <Interpolation Processing> Here, a specific example of the interpolation processing performed when generating listener reference object metadata in step S83 or step S90 in Fig. 9 will be described. In particular, the case where polar coordinate vector synthesis is performed will be described.

[0232] For example, as shown on the left side of Fig. 10, assume that position LP21 in the free viewpoint space is the listener position. Also assume that CVP1 to CVP5 are set in the free viewpoint space, and that listener reference object position information is obtained using these CVP1 to CVP5. Note that, for simplicity of explanation, Fig. 10 shows an example in which CVPs are arranged on a two-dimensional plane.

[0233] In this example, CVP1 to CVP5 are located around TP, with the TP at the center. The positive y-axis of the CVP polar coordinate system for each CVP points from the CVP to the TP.

[0234] Furthermore, the positions of the same object (hereinafter also referred to as the object of interest) when viewed from each of CVP1 to CVP5 are positions OBP1 to OBP5, respectively. In other words, the polar coordinates of the CVP polar coordinate system indicating positions OBP1 to OBP5 are the object position information of CVP1 to CVP5.

[0235] At this time, it is assumed that each CVP is rotated so that the y-axis of the CVP polar coordinate system is in the vertical direction, that is, in the upward direction in the drawing.

[0236] Furthermore, if the object of interest is repositioned so that the origin O' of the CVP polar coordinate system of each CVP after rotation becomes the origin of the same CVP polar coordinate system, the relationship between the positions OBP1 to OBP5 of the object of interest in each CVP as viewed from the origin of the CVP polar coordinate system will be as shown on the right side of the figure. That is, the right side of the figure shows the object positions when each CVP is the origin and the median plane is the positive direction of the Y axis.

[0237] In the CVP polar coordinate system of each CVP, the median plane (y-axis) is constrained to point toward the TP, so the positional relationship shown on the right side of the figure can be easily determined.

[0238] Furthermore, on the right side of the figure, vectors starting from the position of the CVP, ie, the origin, and ending at positions OBP1 to OBP5 of the target object are defined as vectors V41 to V45.

[0239] In this case, as shown in Fig. 11, by weighting each vector (CVP) and combining vectors V41 to V45, vector V51 indicating the position of the target object as viewed from listener position LP21 is obtained. Note that in Fig. 11, the weight of each CVP is set to 1 for ease of explanation.

[0240] Furthermore, if the weight of each CVP during vector synthesis is called a contribution rate, the contribution rate may be calculated, for example, from the ratio of the distance from the listener position to the CVP in the free viewpoint space (common absolute coordinate system).

[0241] 12, for example, assume that the listener position indicated by the listener position information is position F, and the positions of the three CVPs used in the interpolation process are positions A to C. Also assume that the absolute coordinates of position F in the common absolute coordinate system are (xf, yf, zf), and the absolute coordinates of positions A, B, and C in the common absolute coordinate system are (xa, ya, za), (xb, yb, zb), and (xc, yc, zc). Note that the absolute coordinates indicating the position of each CVP in the common absolute coordinate system can be obtained from the CVP position information included in the configuration information.

[0242] The metadecoder 74 calculates the ratio (distance ratio) between the distance AF from position F to position A, the distance BF from position F to position B, and the distance CF from position F to position C, and determines the inverse of this distance ratio as the ratio (dependence ratio) of the contribution rates of the CVPs at each position.

[0243] That is, the meta decoder 74 calculates the following equation (9) by setting AF:BF:CF = a:b:c and defining the dependency of each CVP at positions A to C on the listener position (listener reference object position information) as dp(AF), dp(BF), and dp(CF).

[0244]

[0245] However, a, b, and c in equation (9) are as shown in the following equation (10).

[0246]

[0247] Furthermore, the meta decoder 74 normalizes the dependencies dp(AF), dp(BF), and dp(CF) shown in equation (9) by calculating the following equation (11), and determines the normalized dependencies ndp(AF), ndp(BF), and ndp(CF) as the final contribution rates. Note that a, b, and c in equation (11) are also determined by equation (10).

[0248]

[0249] Regarding the contribution rates ndp(AF) to ndp(CF) obtained in this manner, the shorter the distance from the listener position to the CVP, the closer the contribution rate of that CVP becomes to 1. Note that the contribution rate of each CVP is not limited to the above example, and may be obtained by any other method.

[0250] The metadecoder 74 calculates the contribution rate of each CVP by finding the ratio of the distance from the listener position to the CVP based on the listener position information and the CVP position information.

[0251] Based on the above, a specific example of calculating listener reference object metadata, more specifically, listener reference object position information and listener reference object gain, by interpolation processing will be described.

[0252] First, the metadecoder 74 selects a CVP to be used for interpolation based on the listener position information and the CVP position information included in the configuration information. The CVPs to be used for interpolation may be a portion of all CVPs located around the listener position, or all CVPs may be used for interpolation.

[0253] For each selected CVP, the meta decoder 74 calculates, based on the object position information, a three-dimensional object position vector with the CVP as the start point and the position of the object as seen from the CVP as the end point.

[0254] For example, if the object three-dimensional position vector for the jth object as viewed from the i-th CVPi is (Obj_vector_x[i][j], Obj_vector_y[i][j], Obj_vector_z[i][j]), the object three-dimensional position vector can be obtained by calculating the following equation (12).

[0255] Here, the polar coordinates indicated by the object position information of the jth object as viewed from the i-th CVPi are assumed to be (Azi[i][j], Ele[i][j], rad[i][j]).

[0256]

[0257] Next, the metadecoder 74 performs calculations similar to those of the above equations (9) to (11) based on the listener position information and the CVP position information of each CVPi included in the configuration information to determine the contribution rate ndp(i) of each CVPi, which serves as a weighting factor for the interpolation process. The contribution rate ndp(i) is a weighting factor determined by the ratio of the distances from the listener position to the CVPi, more specifically, the reciprocal ratio of the distances.

[0258] Furthermore, the meta decoder 74 calculates the following equation (13) based on the object three-dimensional position vector obtained by calculating equation (12), the contribution rate ndp(i) of each CVPi, and the object gain Obj_gain[i][j] of the j-th object as seen from CVPi, thereby obtaining the listener-reference object position information (Intp_x(j), Intp_y(j), Intp_z(j)) and listener-reference object gain Intp_gain(j) for the j-th object.

[0259]

[0260] Equation (13) calculates a weighted vector sum. That is, the sum of the object three-dimensional position vectors of each CVPi multiplied by the contribution rate ndp(i) is calculated as the listener-reference object position information. Also, the sum of the object gains of each CVPi multiplied by the contribution rate ndp(i) is calculated as the listener-reference object gain.

[0261] The listener reference object position information obtained by equation (13) is an absolute coordinate in an absolute coordinate system with the listener position as the origin and the direction from the listener position to the TP as the positive direction of the y-axis, i.e., the direction of the median plane.

[0262] However, when the rendering processing unit 75 performs rendering processing in a polar coordinate system, listener reference object position information expressed in polar coordinates is required.

[0263] Therefore, the metadecoder 74 converts the listener reference object position information (Intp_x(j), Intp_y(j), Intp_z(j)) in absolute coordinates obtained by equation (13) into listener reference object position information (Intp_azi(j), Intp_ele(j), Intp_rad(j)) in polar coordinates by calculating the following equation (14):

[0264]

[0265] The meta decoder 74 uses the thus obtained listener reference object position information (Intp_azi(j), Intp_ele(j), Intp_rad(j)) as the final listener reference object position information.

[0266] The listener reference object position information expressed in polar coordinates obtained by equation (14) is a polar coordinate system in which the listener position is the origin and the direction from the listener position to the TP is the positive direction of the y-axis, i.e., the direction of the median plane.

[0267] Through the above processing, listener reference object metadata including listener reference object position information and listener reference object gain can be obtained.

[0268] In this technology, by setting one TP common to all CVPs, the desired listener-reference object position information and listener-reference object gain can be obtained by simple calculation.

[0269] In addition, when the listener reference object metadata includes priority information, spread information, etc., the priority information and spread information may also be generated by interpolation processing, etc. Furthermore, for example, the priority information and spread information included in each of the plurality of object metadata used to generate the listener reference object metadata, the highest value, the lowest value, or the CVP closest to the listener position, etc. may be used as the priority information and spread information of the listener reference object metadata. Furthermore, the median or average value of the priority information and spread information may be used as the priority information and spread information of the listener reference object metadata.

[0270] Second Embodiment Explanation of Metadata Generation Process As described above, the prohibited space may be a space that includes at least a TP, and the TP (target point) itself may be set as the prohibited space.

[0271] If the TP itself is considered to be a prohibited space, when the listener position overlaps with the TP, that is, when the input listener position becomes the TP, the listener position is changed (forced movement).

[0272] In this case, the previous listener position can be made the final listener position, and when the previous listener position is made the final listener position, the listener reference object metadata obtained for the previous listener position is used as is.

[0273] If it is known between the server 11 and the client 61 that the TP is a prohibited space, the configuration information does not necessarily need to include a prohibited space information presence flag or radius information.

[0274] If the TP itself is considered to be a prohibited space, while the output audio data generation process described with reference to FIG. 8 is being performed, the client 61 performs the metadata generation process shown in FIG. 13 as a process corresponding to the process of step S44.

[0275] The metadata generation process by the client 61 will be described below with reference to the flowchart of FIG.

[0276] In step S141, the meta decoder 74 sets an initial value to a variable for storing past listener positions, that is, a variable PrevPos indicating the immediately preceding listener position.

[0277] For example, the metadecoder 74 sets the coordinates of the common absolute coordinate system indicating any position other than the position of the TP in the free viewpoint space as the initial value of the variable PrevPos. The initial value of the variable PrevPos may be a predetermined value or a randomly determined value, as long as it is a value other than the value of the TP position information.

[0278] In step S142, the meta decoder 74 sets the TP warning flag to "false" and holds the TP warning flag. That is, the meta decoder 74 temporarily sets the value of the TP warning flag to "false."

[0279] In step S143, the meta decoder 74 determines whether the input listener position is the same as the TP.

[0280] That is, the meta decoder 74 identifies the position of the TP based on the CVP information of each CVP contained in the configuration information supplied from the audio decoder 72, particularly the CVP position information and CVP direction information contained in the CVP information, and generates TP position information indicating the position of the TP.

[0281] The meta decoder 74 compares the position of the TP indicated by the TP position information with the input listener position indicated by the input listener position information supplied from the listener information acquisition unit 73 to determine whether the input listener position is the same as the TP.

[0282] If it is determined in step S143 that the input listener position is the same as the TP position, a change in listener position (forced movement of the listener) is necessary, and the process then proceeds to step S144.

[0283] In step S144, the meta decoder 74 sets the listener position to PrevPos. That is, the meta decoder 74 sets the position immediately before the listener indicated by the variable PrevPos as the final listener position (the changed listener position), rather than the input listener position, and generates listener position information indicating that listener position.

[0284] As a result, the position of the listener that was about to move to the TP is forcibly moved to the position where it was just before. In other words, the movement to the input listener position instructed by the input operation of the listener etc. is not accepted, and the listener remains in the position where it was just before without moving.

[0285] In step S145, the meta decoder 74 sets the TP warning flag to "true." That is, the meta decoder 74 changes the value of the TP warning flag from "false" to "true."

[0286] After the process of step S145 is performed, the process proceeds to step S146.

[0287] Also, if it is determined in step S143 that the input listener position is not the same position as the TP, there is no need to change the listener position, so steps S144 and S145 are not performed and the process proceeds to step S146.

[0288] In this case, the meta decoder 74 sets the input listener position as the final listener position and generates listener position information indicating the input listener position.

[0289] If the process of step S145 is performed, or if it is determined in step S143 that the input listener position is not the same position as the TP, the process of step S146 is performed.

[0290] In step S 146 , the meta decoder 74 generates listener reference object metadata based on the listener position information and the configuration information and multiplexed object metadata (object metadata) supplied from the audio decoder 72 .

[0291] 9 is performed to generate listener reference object metadata, which is then supplied to the rendering processing unit 75. The meta decoder 74 also supplies the TP warning flag, listener position information, and TP position information that it holds to the presentation control unit 76.

[0292] In this case, for example, if it is determined in step S143 that the input listener position is not the same as the TP, the input listener position is regarded as the final listener position, and listener reference object metadata is generated. Also, a TP warning flag with a value of "false" is supplied to the presentation control unit 76.

[0293] In contrast, when the processing of step S145 is performed, a position different from the input listener position is set as the final listener position, listener reference object metadata is generated, and a TP warning flag with a value of "true" is supplied to the presentation control unit 76.

[0294] When the process of step S145 is performed, the immediately preceding listener position is set as the final listener position, i.e., the current listener position, in step S144, so the listener has not moved from its immediately preceding position. Therefore, more specifically, the listener reference object metadata that has already been generated for the immediately preceding listener position is supplied again to the rendering processing unit 75.

[0295] In step S147, the meta decoder 74 updates the variable PrevPos that it holds.

[0296] Specifically, the meta decoder 74 updates the variable PrevPos so that the value of the variable PrevPos becomes the value indicating the final listener position, that is, the value of the listener position information.

[0297] Therefore, for example, if it is determined in step S143 that the input listener position is not the same as the TP position, the value of the variable PrevPos will be set to a value indicating the input listener position. In this case, when the next process is performed, the input listener position input in this process will be set to the immediately previous listener position.

[0298] On the other hand, if it is determined in step S143 that the input listener position is the same as the TP position, the value of the variable PrevPos is not actually updated, meaning that the listener has not moved from the previous listener position.

[0299] In step S148, the meta-decoder 74 determines whether there is data to process. In step S148, the same determination as in step S91 of FIG.

[0300] If it is determined in step S148 that there is data to be processed, the process returns to step S142, and the above-described process is repeated.

[0301] On the other hand, if it is determined in step S148 that there is no data to process, the metadata generation process ends.

[0302] In this way, the client 61 determines the final listener position depending on whether the input listener position overlaps with the TP, which is a prohibited space, and generates listener reference object metadata depending on the determination result. In this way, it is possible to prevent the listener orientation from being unable to be calculated, and the free viewpoint space UI can always be displayed appropriately.

[0303] The metadata generation process described in the first and second embodiments above may be performed by the meta-encoder 21 or the like on the server 11 side, rather than on the client 61 side. For example, when the server 11 acquires input listener position information from the client 61 and generates listener reference object metadata, the server 11 performs the same metadata generation process as shown in Fig. 9 or 13. Then, the listener reference object metadata, TP warning flag, final listener position information, and the like are stored in the coded bitstream.

[0304] As described above, according to the present technology, when a listener moves in a three-dimensional free viewpoint space with the direction of the TP as the line of sight, i.e., the direction of the listener's face, even if the same position as the TP is input as the input listener position, it is possible to prevent the direction of the listener's face from becoming incalculable. In other words, it is possible to prevent the Yaw and Pitch, which indicate the direction of the listener's face, from becoming infinite. This makes it possible to always display the free viewpoint space UI appropriately.

[0305] <Example of Computer Configuration> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs constituting the software are installed on a computer. Here, the computer includes a computer built into dedicated hardware, and a general-purpose personal computer, for example, that can execute various functions by installing various programs.

[0306] FIG. 14 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program.

[0307] In the computer, a CPU (Central Processing Unit) 501, a ROM (Read Only Memory) 502, and a RAM (Random Access Memory) 503 are interconnected by a bus 504.

[0308] An input / output interface 505 is further connected to the bus 504. An input unit 506, an output unit 507, a recording unit 508, a communication unit 509, and a drive 510 are connected to the input / output interface 505.

[0309] The input unit 506 includes a keyboard, a mouse, a microphone, an image sensor, etc. The output unit 507 includes a display, a speaker, etc. The recording unit 508 includes a hard disk, a non-volatile memory, etc. The communication unit 509 includes a network interface, etc. The drive 510 drives a removable recording medium 511 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.

[0310] In a computer configured as described above, the CPU 501 loads, for example, a program recorded in the recording unit 508 into the RAM 503 via the input / output interface 505 and the bus 504, and executes the program, thereby performing the above-described series of processes.

[0311] The program executed by the computer (CPU 501) can be provided by being recorded on a removable recording medium 511 such as a package medium, for example. The program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.

[0312] In a computer, a program can be installed in the recording unit 508 via the input / output interface 505 by inserting a removable recording medium 511 into the drive 510. The program can also be received by the communication unit 509 via a wired or wireless transmission medium and installed in the recording unit 508. Alternatively, the program can be installed in the ROM 502 or the recording unit 508 in advance.

[0313] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.

[0314] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present technology.

[0315] For example, the present technology can be configured as a cloud computing system in which a single function is shared and processed collaboratively by a plurality of devices via a network.

[0316] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by a plurality of devices.

[0317] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.

[0318] Furthermore, the present technology can also be configured as follows.

[0319] (1) An information processing device comprising a position determination unit that determines whether the user's position indicated by the position information is within a prohibited space based on position information indicating the user's position within a space and prohibited space information for identifying a prohibited space, the prohibited space including a predetermined target point, into which the user is prohibited from entering, and if the user's position is within the prohibited space, changes the user's position to a position outside the prohibited space that is different from the position indicated by the position information. (2) The information processing device described in (1), in which the position determination unit determines the user's position after the change to be a position on the surface of or outside the prohibited space that is closest to the user's position indicated by the position information. (3) The information processing device described in (1), in which the position determination unit determines the user's position after the change to be the position of an intersection between the user's position indicated by the position information and the target point and the surface of the prohibited space. (4) The information processing device described in (3), in which the position determination unit determines the user's position after the change to be the position of an intersection closest to the user's position indicated by the position information, among the multiple intersections. (5) The information processing device according to any one of (1) to (4), wherein the prohibited space is a spherical space. (6) The information processing device according to any one of (1) to (5), wherein the prohibited space is a space including a specific object arranged within the space. (7) The information processing device according to (1), wherein the prohibited space is the target point. (8) The information processing device according to (1), wherein, when the user's position is within the prohibited space, the position determination unit sets the user's previous position as the user's position after change. (9) The information processing device according to any one of (1) to (8), wherein the position determination unit outputs warning information indicating whether the user's position is within the prohibited space. (10) The information processing device according to (9), further comprising a presentation control unit that displays an image related to the space based on position information indicating the user's final position, the orientation of the user within the space, and the warning information.(11) The information processing device according to (10), wherein at least one of the prohibited space, the target point, and an object is displayed in the image. (12) The information processing device according to (10) or (11), wherein a notification that entry into the prohibited space is prohibited is displayed in the image. (13) The information processing device according to any one of (1) to (12), wherein, when the user's position is within the prohibited space, the position determination unit generates metadata for an object arranged in the space based on position information indicating the user's position after the change. (14) The information processing device according to (13), wherein a plurality of positions in the space are defined as control viewpoints, and the position determination unit generates the metadata for the user's position after the change based on the metadata for each of the plurality of control viewpoints and position information indicating the user's position after the change. (15) The information processing device according to (14), wherein the position determination unit generates the metadata for the user's position after the change by interpolation. (16) The information processing device according to any one of (13) to (15), wherein the metadata includes at least one of position information of the object in the space, a gain of the object, priority information of the object, sound source type information indicating a type of the object, and spread information indicating a degree of spread of the object. (17) The information processing device according to any one of (13) to (16), further comprising a rendering processing unit that performs rendering processing based on the metadata of the object at the user's position after the change and audio data of the object, to generate output audio data. (18) The information processing device according to (17), wherein the rendering processing is processing using at least one of HRTF, BRIR, RIR, ITD, IID, HOA, or VBAP. (19) The information processing device according to any one of (1) to (18), wherein the position determination unit does not change the user's position when the user's position is outside the prohibited space.(20) The information processing device according to (9), further comprising a presentation control unit that controls a notification that entry into the prohibited space is prohibited based on the warning information. (21) An information processing method in which an information processing device determines whether the user's position indicated by the position information is within the prohibited space based on position information indicating the user's position within a space and prohibited space information for specifying a prohibited space into which the user is prohibited from entering, the prohibited space including a predetermined target point, and if the user's position is within the prohibited space, changes the user's position to a position outside the prohibited space that is different from the position indicated by the position information. (22) A program that causes a computer to execute processing including the steps of: determining whether the user's position indicated by the position information is within the prohibited space based on position information indicating the user's position within a space and prohibited space information for specifying a prohibited space into which the user is prohibited from entering, the prohibited space including a predetermined target point, and if the user's position is within the prohibited space, changes the user's position to a position outside the prohibited space that is different from the position indicated by the position information.

[0320] REFERENCE SIGNS LIST 11 Server, 21 Meta-encoder, 22 Audio encoder, 23 Communication unit, 61 Client, 71 Communication unit, 72 Audio decoder, 73 Listener information acquisition unit, 74 Meta-decoder, 75 Rendering processing unit, 76 Presentation control unit

Claims

1. An information processing device comprising a position determination unit that, based on position information indicating a user's position within a space and prohibited space information for identifying a prohibited space in which the user is prohibited from entering, including a specified target point, determines whether the user's position indicated by the position information is within the prohibited space, and if the user's position is within the prohibited space, changes the user's position to a position outside the prohibited space that is different from the position indicated by the position information.

2. The information processing device according to claim 1, wherein the position determination unit determines the position closest to the user's position indicated by the position information on the surface of or outside the prohibited space as the changed position of the user.

3. The information processing device according to claim 1, wherein the position determination unit determines the position of the intersection of the vector connecting the user's position indicated by the position information and the target point with the surface of the prohibited space as the changed user's position.

4. The information processing device according to claim 3, wherein the position determination unit determines the position of the intersection among the multiple intersections that is closest to the user's position indicated by the position information as the changed position of the user.

5. The information processing device according to claim 1, wherein the prohibited space is a spherical space.

6. The information processing device according to claim 1, wherein the prohibited space is a space that includes a specific object placed within the space.

7. The information processing device according to claim 1, wherein the prohibited space is the target point.

8. The information processing device according to claim 1, wherein, when the user's position is within the prohibited space, the position determination unit determines the previous position of the user as the changed position of the user.

9. The information processing device according to claim 1, wherein the position determination unit outputs warning information indicating whether or not the user's position is within the prohibited space.

10. The information processing device according to claim 9, further comprising a presentation control unit that displays an image relating to the space based on location information indicating the user's final position, the user's orientation within the space, and the warning information.

11. The information processing device according to claim 10, wherein at least one of the prohibited space, the target point, and an object is displayed in the image.

12. The information processing device according to claim 10, wherein the image displays a notification that entry into the prohibited space is prohibited.

13. The information processing device according to claim 1, wherein, when the user's position is within the prohibited space, the position determination unit generates metadata for an object placed within the space based on position information indicating the user's position after the change.

14. The information processing device of claim 13, wherein a plurality of positions within the space are defined as control viewpoints, and the position determination unit generates the metadata at the user's position after the change based on the metadata for each of the plurality of control viewpoints and position information indicating the user's position after the change.

15. The information processing device according to claim 14, wherein the position determination unit generates the metadata for the changed user position by an interpolation process.

16. The information processing device of claim 13, wherein the metadata includes at least one of position information of the object in the space, gain of the object, priority information of the object, sound source type information indicating the type of the object, and spread information indicating the degree of spread of the object.

17. The information processing device according to claim 13, further comprising a rendering processing unit that performs rendering processing based on the metadata of the object at the changed user position and the audio data of the object, and generates output audio data.

18. The information processing device according to claim 17, wherein the rendering process is a process using at least one of HRTF, BRIR, RIR, ITD, IID, HOA, and VBAP.

19. An information processing method in which an information processing device determines whether the user's position indicated by the position information is within a space based on position information indicating the user's position within a space and prohibited space information for identifying a prohibited space into which the user is prohibited from entering, including a specified target point, and if the user's position is within the prohibited space, changes the user's position to a position outside the prohibited space that is different from the position indicated by the position information.

20. A program that causes a computer to execute a process including the steps of: determining whether the user's position indicated by the position information is within a prohibited space based on position information indicating the user's position within a space and prohibited space information for identifying a prohibited space into which the user is prohibited from entering, including a specified target point; and, if the user's position is within the prohibited space, changing the user's position to a position outside the prohibited space that is different from the position indicated by the position information.

Citation Information

Patent Citations

  • Method and system for handling local transitions between listening positions in a virtual reality environment

    JP2021507558A

  • Method, apparatus and computer program for spatial audio

    JP2021508197A

  • Rendering metadata to control user movement based audio rendering

    US20200304935A1

  • An apparatus and associated methods for audio presented as spatial audio

    WO2019057530A1

  • Information processing device and method, and program

    WO2021140951A1