Information Processing Apparatus, Information Processing Method, and Program
The system dynamically adjusts object positions and gains in free viewpoint spaces to align with content creator intent, enhancing immersion by focusing on listener-object relationships beyond physical proximity.
Patent Information
- Application Number
- CN202080091452.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-03
- Filing Date
- 2020-12-25
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2040-12-25
AI Technical Summary
In the prior art, the physical relationship between the listener and the object cannot fully convey the entertainment possibility of the content, resulting in insufficient reproduction effect of the desired content.
By acquiring the listener position information and object position information of multiple reference viewpoints, the position information of the object at the listener's viewpoint is calculated using the information processing device, and content reproduction is performed based on the intention of the content creator, including the object position calculation unit and the rendering processing unit, to realize the precise calculation and reproduction of the object position and gain information.
The content reproduction is realized according to the intention of the content creator, which enhances the listener's entertainment experience, ensures the precise position and gain information calculation of the object at different viewpoints, and improves the interesting communication of the content.
Smart Images

Figure CN114930877B_ABST
Abstract
Description
Technical Field
[0001] The present technology relates to an information processing device, an information processing method, and a program, and more particularly, to an information processing device, an information processing method, and a program capable of realizing content reproduction based on the intention of a content creator. Background Art
[0002] For example, in a free viewpoint space, each object arranged in the space using an absolute coordinate system is fixedly arranged (see, for example, Patent Document 1).
[0003] In this case, the direction of each object viewed from an arbitrary listening position is uniquely obtained based on the coordinate position, face direction, and relationship with the object of the listener in the absolute space, and the gain of each object is uniquely obtained based on the distance from the listening position, and the sound of each object is reproduced.
[0004] Citation List
[0005] Patent Literature
[0006] Patent Document 1: WO 2019 / 198540 A. Summary of the Invention
[0007] Problems to be Solved by the Invention
[0008] On the other hand, there are points to be emphasized as the content of art and the listener.
[0009] For example, there are cases where it is desired that an object is in the front, such as with respect to music content, an instrument or a performer at a certain listening point where the desired content is emphasized in its essential aspect, or with respect to sports content, a player to be emphasized.
[0010] In view of the above, there is a possibility that the mere physical relationship between the listener and the object as described above cannot fully convey the entertainment of the content.
[0011] In view of this situation, the present technology has been created, and the present technology realizes content reproduction based on the intention of the content creator while following the free position of the listener.
[0012] Solution to the Problem
[0013] An information processing apparatus according to an aspect of the present technology includes: a listener position information acquisition unit that acquires listener position information of a listener's viewpoint; a reference viewpoint information acquisition unit that acquires position information of a first reference viewpoint and object position information of an object at the first reference viewpoint, and position information of a second reference viewpoint and object position information of the object at the second reference viewpoint; and an object position calculation unit that calculates object position information at the listener's viewpoint based on the listener position information, the position information of the first reference viewpoint and the object position information at the first reference viewpoint, and the position information of the second reference viewpoint and the object position information at the second reference viewpoint.
[0014] An information processing method or program according to an aspect of the present technology includes the steps of: acquiring listener position information of a listener's viewpoint; acquiring position information of a first reference viewpoint and object position information of an object at the first reference viewpoint, and position information of a second reference viewpoint and object position information of the object at the second reference viewpoint; and calculating object position information at the listener's viewpoint based on the listener position information, the position information of the first reference viewpoint and the object position information at the first reference viewpoint, and the position information of the second reference viewpoint and the object position information at the second reference viewpoint.
[0015] According to an aspect of the present technology, acquire listener position information of a listener's viewpoint; acquire position information of a first reference viewpoint and object position information of an object at the first reference viewpoint, and position information of a second reference viewpoint and object position information of the object at the second reference viewpoint; and calculate object position information at the listener's viewpoint based on the listener position information, the position information of the first reference viewpoint and the object position information at the first reference viewpoint, and the position information of the second reference viewpoint and the object position information at the second reference viewpoint. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a diagram showing the configuration of a content reproduction system.
[0017] Figure 2 is a diagram showing the configuration of a content reproduction system.
[0018] Figure 3 is a diagram describing a reference viewpoint.
[0019] Figure 4 is a diagram showing an example of system configuration information.
[0020] Figure 5 is a diagram showing an example of system configuration information.
[0021] Figure 6 is a diagram describing coordinate transformation.
[0022] Figure 7 It is a diagram for describing coordinate axis transformation processing.
[0023] Figure 8 It is a diagram showing an example of the transformation result by coordinate axis transformation processing.
[0024] Figure 9 It is a diagram for describing interpolation processing.
[0025] Figure 10 It is a diagram showing a sequence example of a content reproduction system.
[0026] Figure 11 It is a diagram for describing an example of arranging an object closer at a reference viewpoint.
[0027] Figure 12 It is a diagram for describing the interpolation of the absolute coordinate position information of an object.
[0028] Figure 13 It is a diagram for describing the internal division ratio in a viewpoint-side triangular mesh.
[0029] Figure 14 It is a diagram for describing the calculation of an object position based on the internal division ratio.
[0030] Figure 15 It is a diagram for describing the calculation of gain information based on the internal division ratio.
[0031] Figure 16 It is a diagram for describing the selection of a triangular mesh.
[0032] Figure 17 It is a diagram showing the configuration of a content reproduction system.
[0033] Figure 18 It is a flowchart for describing the provision processing and the reproduction audio data generation processing.
[0034] Figure 19 It is a flowchart for describing the viewpoint selection processing.
[0035] Figure 20 It is a diagram showing an example of the configuration of a computer. Detailed implementation manners
[0036] The following describes the implementation manners applying the present technology with reference to the accompanying drawings.
[0037] <First implementation manner>
[0038] <Example of the configuration of a content reproduction system>
[0039] The present technology has the following features F1 to F6.
[0040] (Feature F1)
[0041] Characteristics of preparing in advance the object arrangement and gain information at a plurality of reference viewpoints in a free viewpoint space.
[0042] (Characteristic F2)
[0043] Characteristics of obtaining the object position and gain information at an arbitrary listening point based on the object arrangement and gain information at a plurality of reference viewpoints sandwiching or surrounding the arbitrary listening point (listening position).
[0044] (Characteristic F3)
[0045] In the case of obtaining the object position and gain amount at an arbitrary listening point, obtaining a ratio based on a plurality of reference viewpoints sandwiching or surrounding the arbitrary listening point and the arbitrary listening point, and using the ratio to obtain the object position relative to the arbitrary listening point.
[0046] (Characteristic F4)
[0047] Characteristics that the object arrangement information of a plurality of pre - prepared reference viewpoints uses a polar coordinate system and is transmitted.
[0048] (Characteristic F5)
[0049] Characteristics that the object arrangement information of a plurality of pre - prepared reference viewpoints uses an absolute coordinate system and is transmitted.
[0050] (Characteristic F6)
[0051] In the case of calculating the object position at an arbitrary listening point, the listener can listen to the object arrangement closer to any reference viewpoint by using a specific deviation coefficient.
[0052] First, a content reproduction system to which this technology is applied will be described.
[0053] The content reproduction system includes a server and a client that encode, transmit, and decode each piece of data.
[0054] For example, the listener position information is transmitted from the client side to the server as needed, and some object position information is transmitted from the server side to the client based on the result. Then, based on some object position information received on the client side, a rendering process is performed on each object, and the content including the sound of each object is reproduced.
[0055] For example, as Figure 1 configured such a content reproduction system.
[0056] That is, Figure 1 the content reproduction system shown includes a server 11 and a client 12.
[0057] The server 11 includes a configuration information sending unit 21 and an encoded data sending unit 22.
[0058] The configuration information sending unit 21 sends (transmits) the pre-prepared system configuration information to the client 12, receives the viewpoint selection information etc. transmitted from the client 12, and provides this information to the encoded data sending unit 22.
[0059] In the content reproduction system, a plurality of listening positions on a predetermined common absolute coordinate space are pre-specified (set) by the content creator as the positions of the reference viewpoints (hereinafter, also referred to as reference viewpoint positions).
[0060] Here, the content creator pre-specifies (sets) the position on the common absolute coordinate space where the content creator hopes the listener will be as the listening position during content reproduction, and the direction of the face that the content creator hopes the listener will face at this position (i.e., the viewpoint from which the content creator hopes the listener will listen to the sound of the content) as the reference viewpoint.
[0061] In the server 11, system configuration information as information about each reference viewpoint and object polar coordinate encoded data for each reference viewpoint are pre-prepared.
[0062] Here, the object polar coordinate encoded data for each reference viewpoint is obtained by encoding the object polar coordinate position information indicating the relative position of the object observed from the reference viewpoint. In the object polar coordinate position information, the position of the object observed from the reference viewpoint is represented by polar coordinates. Note that even for the same object, the absolute arrangement position of the object within the common absolute coordinate space changes with each reference viewpoint.
[0063] After the start of operation of the content reproduction system, i.e., immediately after the connection with the client 12 is established for example, the configuration information sending unit 21 sends the system configuration information to the client 12 via a network or the like.
[0064] Based on the viewpoint selection information provided by the configuration information sending unit 21, the encoded data sending unit 22 selects two reference viewpoints from among the plurality of reference viewpoints, and sends the object polar coordinate encoded data for each of the two selected reference viewpoints to the client 12 via a network or the like.
[0065] Here, the viewpoint selection information is, for example, information indicating two reference viewpoints selected on the client 12 side.
[0066] Therefore, in the encoded data sending unit 22, the object polar coordinate encoded data of the reference viewpoint requested by the client 12 is acquired and sent to the client 12. Note that the number of reference viewpoints selected by the viewpoint selection information is not limited to two, but can be three or more.
[0067] In addition, the client 12 includes a listener position information acquisition unit 41, a viewpoint selection unit 42, a configuration information acquisition unit 43, an encoded data acquisition unit 44, a decoding unit 45, a coordinate transformation unit 46, an axis transformation processing unit 47, an object position calculation unit 48, and a polar coordinate transformation unit 49.
[0068] The listener position information acquisition unit 41 acquires listener position information indicating the absolute position (listening position) of the listener in the common absolute coordinate space according to a specified operation of the user (listener) or the like, and provides the listener position information to the viewpoint selection unit 42, the object position calculation unit 48, and the polar coordinate transformation unit 49.
[0069] For example, in the listener position information, the position of the listener in the common absolute coordinate space is represented by absolute coordinates. Note that hereinafter, the coordinate system of the absolute coordinates indicated by the listener position information is also referred to as the common absolute coordinate system.
[0070] The viewpoint selection unit 42 selects two reference viewpoints based on the system configuration information provided by the configuration information acquisition unit 43 and the listener position information provided by the listener position information acquisition unit 41, and provides viewpoint selection information indicating the selection result to the configuration information acquisition unit 43.
[0071] For example, the viewpoint selection unit 42 specifies intervals from the position of the listener (listening position) and the assumed absolute coordinate positions of each reference viewpoint, and selects two reference viewpoints based on the results of the specified intervals.
[0072] The configuration information acquisition unit 43 receives the system configuration information transmitted from the server 11 and provides the system configuration information to the viewpoint selection unit 42 and the axis transformation processing unit 47, and transmits the viewpoint selection information provided by the viewpoint selection unit 42 to the server 11 via a network or the like.
[0073] Note that here, an example of setting the viewpoint selection unit 42 that selects reference viewpoints based on listener position information and system configuration information in the client 12 will be described, but the viewpoint selection unit 42 may be set on the server 11 side.
[0074] The encoded data acquisition unit 44 receives the object polar coordinate encoded data transmitted from the server 11, and provides the object polar coordinate encoded data to the decoding unit 45. That is, the encoded data acquisition unit 44 acquires the object polar coordinate encoded data from the server 11.
[0075] The decoding unit 45 decodes the object polar coordinate encoded data provided by the encoded data acquisition unit 44, and provides the obtained object polar coordinate position information to the coordinate transformation unit 46.
[0076] The coordinate transformation unit 46 performs coordinate transformation on the object polar coordinate position information provided by the decoding unit 45, and provides the obtained object absolute coordinate position information to the coordinate axis transformation processing unit 47.
[0077] The coordinate transformation unit 46 performs a coordinate transformation that transforms polar coordinates into absolute coordinates. Accordingly, the object polar coordinate position information, which is polar coordinates indicating the position of an object observed from a reference viewpoint, is transformed into object absolute coordinate position information, which is absolute coordinates indicating the position of the object in an absolute coordinate system with the position of the reference viewpoint as the origin.
[0078] The coordinate axis transformation processing unit 47 performs coordinate axis transformation processing on the object absolute coordinate position information provided by the coordinate transformation unit 46 based on the system configuration information provided by the configuration information acquisition unit 43.
[0079] Here, the coordinate axis transformation processing is a process performed by combining a coordinate transformation (coordinate axis transformation) and an offset shift, and the object absolute coordinate position information indicating the absolute coordinates of the object projected onto a common absolute coordinate space is obtained through the coordinate axis transformation processing. That is, the object absolute coordinate position information obtained through the coordinate axis transformation processing is the absolute coordinates of the common absolute coordinate system indicating the absolute position of the object on the common absolute coordinate space.
[0080] The object position calculation unit 48 performs interpolation processing based on the listener position information provided by the listener position information acquisition unit 41 and the object absolute coordinate position information provided by the coordinate axis transformation processing unit 47, and provides the obtained final object absolute coordinate position information to the polar coordinate transformation unit 49. The final object absolute coordinate position information mentioned here is information indicating the position of the object in the common absolute coordinate system when the listener's viewpoint is at the listening position indicated by the listener position information.
[0081] The object position calculation unit 48 calculates the absolute position (i.e., the absolute coordinates of the common absolute coordinate system) of the object in the common absolute coordinate space corresponding to the listening position from the listening position indicated by the listener position information and the positions of two reference viewpoints indicated by the viewpoint selection information, and determines the absolute position as the final object absolute coordinate position information. At this time, the object position calculation unit 48 obtains the system configuration information from the configuration information acquisition unit 43, and obtains the viewpoint selection information from the viewpoint selection unit 42 as needed.
[0082] The polar coordinate transformation unit 49 performs polar coordinate transformation on the object absolute coordinate position information provided by the object position calculation unit 48 based on the listener position information provided by the listener position information acquisition unit 41, and outputs the obtained polar coordinate position information to a subsequent rendering processing unit (not shown).
[0083] The polar coordinate transformation unit 49 performs a polar coordinate transformation that transforms the object absolute coordinate position information, which is the absolute coordinates of the common absolute coordinate system, into polar coordinate position information, which is the polar coordinates indicating the relative position of the object observed from the listening position.
[0084] Note that although the example of preparing the object polar coordinate encoded data in advance for each reference viewpoint in the server 11 has been described above, the object absolute coordinate position information to be the output of the coordinate axis transformation processing unit 47 may also be prepared in advance in the server 11.
[0085] In this case, for example, as Figure 2 shown, the content reproduction system is configured. Note that in Figure 2 the part corresponding to the part in Figure 1 is denoted by the same reference numeral, and the description is appropriately omitted.
[0086] Figure 2 The content reproduction system shown includes a server 11 and a client 12.
[0087] In addition, the server 11 includes a configuration information sending unit 21 and an encoded data sending unit 22. However, in this example, the encoded data sending unit 22 acquires the object absolute coordinate encoded data of the two reference viewpoints indicated by the viewpoint selection information, and sends the object absolute coordinate encoded data to the client 12.
[0088] That is, in the server 11, for each of the multiple reference viewpoints, object absolute coordinate encoded data obtained by encoding the object absolute coordinate position information, which is the output of the coordinate axis transformation processing unit 47 as Figure 1 shown, is prepared in advance.
[0089] Therefore, in this example, the client 12 is not provided with Figure 1 the coordinate transformation unit 46 or the coordinate axis transformation processing unit 47 shown.
[0090] That is, Figure 2 the client 12 shown includes a listener position information acquisition unit 41, a viewpoint selection unit 42, a configuration information acquisition unit 43, an encoded data acquisition unit 44, a decoding unit 45, an object position calculation unit 48, and a polar coordinate transformation unit 49.
[0091] Figure 2 The configuration of the client 12 shown is different from the configuration of the client 12 shown in Figure 1 that the coordinate transformation unit 46 and the coordinate axis transformation processing unit 47 are not provided, and is the same as the configuration of the client 12 shown in Figure 1 in other points.
[0092] The encoded data acquisition unit 44 receives the object absolute coordinate encoded data transmitted from the server 11 and provides the object absolute coordinate encoded data to the decoding unit 45.
[0093] The decoding unit 45 decodes the object absolute coordinate encoded data provided from the encoded data acquisition unit 44 and provides the obtained object absolute coordinate position information to the object position calculation unit 48.
[0094] <Regarding the present technology>
[0095] Next, the present technology will be further described.
[0096] First, the process of creating the content provided from the server 11 to the client 12 will be described.
[0097] First, an example of a transmission method using a polar coordinate system will be described, that is, an example of transmitting object polar coordinate encoded data as Figure 1 shown.
[0098] Content creation for 3D audio is performed using a polar coordinate system based on a fixed viewpoint, and has the advantage that such a creation method can be used as it is.
[0099] According to the intention of the creator, a plurality of reference viewpoints that the content creator (hereinafter, also simply referred to as the creator) wants the listener to listen to are set in the three-dimensional space.
[0100] Specifically, for example, as Figure 3 shown, four reference viewpoints are set in the common absolute coordinate space as the three-dimensional space. Here, the four positions P11 to P14 specified by the creator are reference viewpoints, and more specifically, the positions of the reference viewpoints.
[0101] The reference viewpoint information is information about each reference viewpoint, and the reference viewpoint information includes reference viewpoint position information and listener direction information. The reference viewpoint position information is the absolute coordinates of the common absolute coordinate system indicating the standing position (that is, the position of the reference viewpoint) in the common absolute coordinate space, and the listener direction information indicates the direction of the listener's face.
[0102] Here, the listener direction information includes, for example, the rotation angle (horizontal angle) of the listener's face in the horizontal direction at the reference viewpoint and the vertical angle indicating the direction of the listener's face in the vertical direction.
[0103] In Figure 3 the arrows drawn adjacent to the corresponding positions P11 to P14 indicate the listener direction information at the reference viewpoints represented by the corresponding positions P11 to P14, that is, the direction of the listener's face.
[0104] In addition, in Figure 3 , the region R11 indicates an example of a region where an object exists, and it can be seen that in this example, at each reference viewing point, the direction of the face of the listener indicated by the listener direction information is the direction of the region R11. For example, at the position P14, the direction of the face of the listener indicated by the listener direction information is backward.
[0105] Next, the creator sets object polar coordinate position information representing the position of each object of each of the set plurality of reference viewing points in polar coordinate format and the gain amount of each object of each reference viewing point. For example, the object polar coordinate position information includes the horizontal angle and the vertical angle of the object viewed from the reference viewing point, and the radius indicating the distance from the reference viewing point to the object.
[0106] When the position of the object and the like are set in this way for each of the plurality of reference viewing points, the following information IFP1 to information IFP5 are obtained as information regarding the reference viewing points.
[0107] (Information IFP1)
[0108] Number of objects
[0109] (Information IFP2)
[0110] Number of reference viewing points
[0111] (Information IFP3)
[0112] Direction (horizontal angle and vertical angle) of the face of the listener at the reference viewing point
[0113] (Information IFP4)
[0114] Absolute coordinate position of the reference viewing point in the absolute space (common absolute coordinate space)
[0115] (Information IFP5)
[0116] Polar coordinate position (horizontal angle, vertical angle, and radius) and gain amount of each object as observed from information IFP3 and information IFP4.
[0117] Here, information IFP3 is the above-mentioned listener direction information, and information IFP4 is the above-mentioned reference viewing point position information.
[0118] In addition, the polar coordinate position, that is, information IFP5, includes the horizontal angle, the vertical angle, and the radius, and is object polar coordinate position information indicating the relative position of the object based on the reference viewing point. Since the object polar coordinate position information corresponds to the polar coordinate encoding information of the Moving Picture Experts Group (MPEG)-H, the encoding system of MPEG-H can be utilized.
[0119] The information including each of the information from Information IFP1 to Information IFP4 among Information IFP1 to Information IFP5 is the above system configuration information.
[0120] Before transmitting data related to an object (i.e., object polar coordinate encoded data or encoded audio data obtained by encoding audio data of the object), the system configuration information is transmitted to the client 12 side.
[0121] Specific examples of the system configuration information are shown, for example, in Figure 4 as follows.
[0122] In Figure 4 the example shown, "NumOfObjs" indicates the number of objects that is the number of objects constituting the content, i.e., the above Information IFP1, and "NumfOfRefViewPoint" indicates the number of reference viewpoints, i.e., the above Information IFP2.
[0123] In addition, Figure 4 the system configuration information shown includes reference viewpoint information corresponding to the number "NumfOfRefViewPoint" of reference viewpoints.
[0124] That is, "RefViewX[i]", "RefViewY[i]", and "RefViewZ[i]" respectively indicate the X coordinate, Y coordinate, and Z coordinate of a common absolute coordinate system, and the common absolute coordinate system indicates the position of the reference viewpoint that is the reference viewpoint position information constituting the i-th reference viewpoint as Information IFP4.
[0125] In addition, "ListenerYaw[i]" and "ListenerPitch[i]" are the horizontal angle (yaw angle) and vertical angle (pitch angle) of the listener direction information constituting the i-th reference viewpoint as Information IFP3.
[0126] In addition, in this example, the system configuration information includes information "ObjectOverLapMode[i]" indicating a reproduction mode in the case where the positions of the listener and the object overlap with each other for each object (i.e., the listener (listening position) and the object are at the same position).
[0127] Next, an example of a transmission method using an absolute coordinate system will be described, i.e., an example of transmitting object absolute coordinate encoded data as shown in Figure 2 the following.
[0128] In addition, in the case of transmitting object absolute coordinate encoded data, similar to the case of transmitting object polar coordinate encoded data, the object position with respect to each reference viewing point is recorded as absolute coordinate position information. That is, the object absolute coordinate position information for each object is prepared by the creator for each reference viewing point.
[0129] However, in this example, different from the example of the transmission method using the polar coordinate system, there is no need to transmit listener direction information indicating the direction of the listener's face.
[0130] In the example of using the transmission method using the absolute coordinate system, the following information IFA1 to information IFA4 is obtained as information about the reference viewing point.
[0131] (Information IFA1)
[0132] The number of objects
[0133] (Information IFA2)
[0134] The number of reference viewing points
[0135] (Information IFA3)
[0136] The absolute coordinate position of the reference viewing point in the absolute space
[0137] (Information IFA4)
[0138] When the listener exists at the absolute coordinate position indicated in Information IFA3, the absolute coordinate position and gain amount of each object
[0139] Here, Information IFA1 and Information IFA2 are the same information as the above Information IFP1 and Information IFP2, and Information IFA3 is the above reference viewing point position information.
[0140] In addition, the absolute coordinate position of the object represented by Information IFA4 is object absolute coordinate position information indicating the absolute position of the object on the common absolute coordinate space represented by the absolute coordinates of the common absolute coordinate system
[0141] Note that when transmitting object absolute coordinate encoded data from the server 11 to the client 12, object absolute coordinate position information indicating the position of the object that precisely corresponds to the positional relationship between the listener and the object, such as the distance from the listener to the object, can be generated and transmitted. In this case, the amount of information (bit depth) of the object absolute coordinate position information can be reduced without causing a feeling of deviation in the sound image position.
[0142] For example, as the distance from the listener to the object is shorter, object absolute coordinate position information (object absolute coordinate encoded data) with higher precision is generated, that is, object absolute coordinate position information indicating a more precise position.
[0143] This is because, although the position of the object deviates according to the quantization precision (quantization step) at the time of encoding, the longer the distance from the listener to the object, the larger the size (tolerance) of the position deviation that does not cause a feeling of deviation in the localization position of the sound image.
[0144] Specifically, for example, object absolute coordinate encoded data obtained by encoding object absolute coordinate position information with the highest precision is prepared in advance and stored in the server 11.
[0145] Then, by extracting a part of the object absolute coordinate encoded data with the highest precision, object absolute coordinate encoded data obtained by quantizing object absolute coordinate position information with an arbitrary quantization precision can be obtained.
[0146] Therefore, the encoded data sending unit 22 extracts a part or all of the object absolute coordinate encoded data with the highest precision according to the distance from the listening position to the object, and sends the obtained object absolute coordinate encoded data with a predetermined precision to the client 12. Note that in this case, the encoded data sending unit 22 can obtain the listener position information from the listener position information acquisition unit 41 through the configuration information sending unit 21, the configuration information acquisition unit 43, and the viewpoint selection unit 42.
[0147] In addition, in Figure 2 the content reproduction system shown, system configuration information is prepared in advance, and this system configuration information includes each piece of information from information IFA1 to information IFA3 among information IFA1 to information IFA4.
[0148] Before transmitting data related to the object (that is, object absolute coordinate encoded data or encoded audio data), this system configuration information is transmitted to the client 12 side.
[0149] For example, Figure 5 a specific example of such system configuration information is shown in
[0150] In Figure 5 the example shown, similar to the example shown in Figure 4 the system configuration information includes the number of objects "NumOfObjs" and the number of reference viewpoints "NumfOfRefViewPoint".
[0151] In addition, the system configuration information includes reference viewpoint information corresponding to the number of reference viewpoints "NumfOfRefViewPoint".
[0152] That is, the system configuration information includes the X coordinate "RefViewX[i]", Y coordinate "RefViewY[i]", and Z coordinate "RefViewZ[i]" of the common absolute coordinate system indicating the position of the reference view point that constitutes the i-th reference view point. As described above, in this example, the reference view point information does not include the listener direction information, but only includes the reference view point position information.
[0153] In addition, the system configuration information includes the reproduction mode "ObjectOverLapMode[i]" when the positions of the listener and the object overlap with each other for each object.
[0154] The system configuration information obtained as described above, the object polar coordinate encoded data or object absolute coordinate encoded data for each object for each reference view point, and the encoded gain information obtained by encoding the gain information indicating the gain amount are held in the server 11.
[0155] Note that hereinafter, when it is not particularly necessary to distinguish between the object polar coordinate position information and the object absolute coordinate position information, the object polar coordinate position information and the object absolute coordinate position information are also simply referred to as object position information. Similarly, hereinafter, when it is not particularly necessary to distinguish between the object polar coordinate encoded data and the object absolute coordinate encoded data, the object polar coordinate encoded data and the object absolute coordinate encoded data are also simply referred to as object coordinate encoded data.
[0156] When the operation of the content reproduction system starts, the configuration information transmission unit 21 of the server 11 transmits the system configuration information to the client 12 side before transmitting the object coordinate encoded data. Therefore, the client 12 side can understand the number of objects constituting the content, the number of reference view points, the position of the reference view points in the common absolute coordinate space, and the like.
[0157] Next, the view point selection unit 42 of the client 12 selects a reference view point based on the listener position information, and the configuration information acquisition unit 43 transmits the view point selection information indicating the selection result to the server 11.
[0158] Note that as described above, the view point selection unit 42 can also be provided in the server 11, and the reference view point can be selected on the server 11 side.
[0159] In this case, the view point selection unit 42 selects a reference view point based on the listener position information received from the client 12 through the configuration information transmission unit 21 and the system configuration information, and provides the view point selection information indicating the selection result to the encoded data transmission unit 22.
[0160] At this time, the viewpoint selection unit 42 designates and selects, for example, two (or two or more) reference viewpoints sandwiching the listening position indicated by the listener position information. In other words, two reference viewpoints are selected such that the listening position is located between the two reference viewpoints.
[0161] Therefore, the object coordinate encoded data of each of the selected multiple reference viewpoints is transmitted to the client 12 side. Further, in more detail, the encoded data transmission unit 22 transmits not only the object coordinate encoded data but also the encoded gain information regarding the two reference viewpoints indicated by the viewpoint selection information to the client 12.
[0162] On the client 12 side, based on the object coordinate encoded data, the encoded gain information at each of the multiple reference viewpoints received from the server 11, and the listener position information, the object absolute coordinate position information and the gain information at an arbitrary viewpoint of the current listener are calculated by interpolation processing or the like.
[0163] Here, a specific example of calculating the final object absolute coordinate position information and the gain information at an arbitrary viewpoint of the current listener will be described.
[0164] Specifically, an example of interpolation processing using a data set of reference viewpoints in a polar coordinate system as two reference viewpoints sandwiching the listener will be described below.
[0165] In this case, the client 12 executes the following-described processes PC1 to PC4 to obtain the final object absolute coordinate position information and the gain information at the viewpoint of the listener.
[0166] (Process PC1)
[0167] In Process PC1, each reference viewpoint is set as the origin from a data set of two reference viewpoints in the polar coordinate system, and the objects included in each data set are transformed to the absolute coordinate system positions. That is, the coordinate transformation unit 46 performs coordinate transformation as Process PC1 on the object polar coordinate position information of each object of each reference viewpoint, and generates object absolute coordinate position information.
[0168] For example, as Figure 6 shown, it is assumed that there is an object OBJ11 in a polar coordinate system space based on the origin O. Further, a three-dimensional orthogonal coordinate system (absolute coordinate system) with the origin O as a reference (origin) and the x-axis, y-axis, and z-axis as respective axes is referred to as the xyz coordinate system.
[0169] In this case, the position of the object OBJ11 in the polar coordinate system can be represented by polar coordinates, which include a horizontal angle θ as the angle in the horizontal direction, a vertical angle γ as the angle in the vertical direction, and a radius r indicating the distance from the origin O to the object OBJ11. In this example, the polar coordinates (θ, γ, r) are the object polar coordinate position information of the object OBJ11.
[0170] Note that the horizontal angle θ is the angle in the horizontal direction starting from the origin O (i.e., in front of the listener). In this example, when the straight line (line segment) connecting the origin O and the object OBJ11 is LN and the straight line obtained by projecting the straight line LN onto the xy plane is LN′, the angle formed by the y-axis and the straight line LN′ is the horizontal angle θ.
[0171] Furthermore, the vertical angle γ is the angle in the vertical direction starting from the origin O (i.e., in front of the listener), and in this example, the angle formed by the straight line LN and the xy plane is the vertical angle γ. Moreover, the radius r is the distance from the listener (origin O) to the object OBJ11, that is, the length of the straight line LN.
[0172] When the position of such an object OBJ11 is represented by the coordinates (x, y, z) of the xyz coordinate system (i.e., absolute coordinates), this position is represented by the following formula (1).
[0173] [Mathematics 1]
[0174] x = -r * sinθ * cosγ
[0175] y = r * cosθ * cosγ
[0176] z = r * sinγ...(1)
[0177] In the process PC1, based on the object polar coordinate position information as polar coordinates, the object absolute coordinate position information as absolute coordinates indicating the position of the object in the xyz coordinate system (absolute coordinate system) with the position of the reference viewpoint as the origin O is calculated by formula (1).
[0178] Specifically, in the process PC1, for each of the two reference viewpoints, coordinate transformation is performed on the object polar coordinate position information of each object at the reference viewpoint.
[0179] (Process PC2)
[0180] In the process PC2, for each of the two reference viewpoints, axis transformation processing is performed on the object absolute coordinate position information obtained through the process PC1 for each object. That is, the axis transformation processing unit 47 performs the axis transformation processing as the process PC2.
[0181] The object absolute coordinate position information at each of the two reference viewpoints obtained by the above processing PC1 (i.e., obtained by the coordinate transformation unit 46) indicates the position in the xyz coordinate system with the reference viewpoint as the origin O. Therefore, the coordinates (coordinate systems) of the object absolute coordinate position information are different for each reference viewpoint.
[0182] Accordingly, as processing PC2, an axis transformation process is performed to synthesize the object absolute coordinate position information of each reference viewpoint into absolute coordinates in a common absolute coordinate system, that is, the absolute coordinates in the common absolute coordinate system (common absolute coordinate space).
[0183] To perform this axis transformation process, in addition to the data set for each reference viewpoint (i.e., the object absolute coordinate position information of each object for each reference viewpoint), the absolute position information (reference viewpoint position information) of the listener and the listener direction information indicating the direction of the listener's face are also required.
[0184] That is, the axis transformation process requires the object absolute coordinate position information obtained by processing PC1 and the system configuration information including the reference viewpoint position information indicating the position of the reference viewpoint in the common absolute coordinate system and the listener direction information at the reference viewpoint.
[0185] Note that here, for simplicity of description, only the rotation angle in the horizontal direction is used as the direction of the face indicated by the listener direction information, but information on the up-and-down movement (pitch) of the face may also be added.
[0186] Now, assume that the common absolute coordinate system is an XYZ coordinate system having an X-axis, a Y-axis, and a Z-axis as corresponding axes, and the rotation angle according to the direction of the face indicated by the listener direction information is For example, as Figure 7 shown, the axis transformation process is performed.
[0187] That is, in Figure 7 the example shown, as the axis transformation process, an axis rotation that rotates the coordinate axes by the rotation angle and a process of shifting the origin of the coordinate axes from the position of the reference viewpoint to the origin position of the common absolute coordinate system are performed. More specifically, a process of shifting the position of the object according to the positional relationship between the reference viewpoint and the origin of the common absolute coordinate system is performed.
[0188] In Figure 7 , the position P21 indicates the position of the reference viewpoint, and the arrow Q11 indicates the direction of the listener's face indicated by the listener direction information at the reference viewpoint. Specifically, here, the X coordinate and Y coordinate of the position P21 in the common absolute coordinate system (XYZ coordinate system) are (Xref, Yref).
[0189] In addition, the position P22 indicates the position of the object when the reference viewing point is at the position P21. Here, the X coordinate and Y coordinate of the common absolute coordinate system indicating the position P22 of the object are (Xobj, Yobj), and the x coordinate and y coordinate of the xyz coordinate system indicating the position P22 of the object and having the reference viewing point as the origin are (xobj, yobj).
[0190] In addition, in this example, the angle formed by the X axis of the common absolute coordinate system (XYZ coordinate system) and the x axis of the xyz coordinate system is the rotation angle of the coordinate axis transformation obtained from the listener direction information
[0191] Therefore, for example, the transformed coordinate axes X (X coordinate) and Y (Y coordinate) are as shown in the following formula (2).
[0192] [Mathematics 2]
[0193] X = reference viewing point X coordinate value + x * cos(φ) + y * sin(φ)
[0194] Y = reference viewing point Y coordinate value - x * sin(φ) + y * cos(φ)...(2)
[0195] Note that in formula (2), x and y represent the x axis (x coordinate) and y axis (y coordinate) before transformation, that is, in the xyz coordinate system. In addition, the "reference viewing point X coordinate value" and "reference viewing point Y coordinate value" in formula (2) represent the X coordinate and Y coordinate indicating the position of the reference viewing point in the XYZ coordinate system (common absolute coordinate system), that is, the X coordinate and Y coordinate constituting the reference viewing point position information.
[0196] In view of the above, in the Figure 7 example, the X coordinate value Xobj and Y coordinate value Yobj representing the position of the object after the coordinate axis transformation processing can be obtained from formula (2).
[0197] That is, in formula (2), is set to the rotation angle obtained from the listener direction information at the position P21 and "Xref", "xobj", and "yobj" are respectively substituted into the "reference viewing point X coordinate value", "x", and "y" in formula (2), and the X coordinate value Xobj can be obtained.
[0198] In addition, in formula (2), is set to the rotation angle obtained from the listener direction information at the position P21 Also, "Yref", "xobj", and "yobj" are respectively substituted into "the Y - coordinate value of the reference view - point", "x", and "y" in formula (2), and the Y - coordinate value Yobj can be obtained.
[0199] Similarly, for example, when two reference view - points A and B are selected according to the view - point selection information, the X - coordinate value and the Y - coordinate value representing the position of the object after the coordinate - axis transformation process of those reference view - points are as shown in formula (3) described below.
[0200] [Mathematics 3]
[0201] xa = the X - coordinate value of reference view - point A+x * cos(φa)+y * sin(φa)
[0202] ya = the Y - coordinate value of reference view - point A−x * sin(φa)+y * cos(φa)
[0203] xb = the X - coordinate value of reference view - point B+x * cos(φb)+y * sin(φb)
[0204] yb = the Y - coordinate value of reference view - point B−x * sin(φb)+y * cos(φb)...(3)
[0205] Note that in formula (3), xa and ya represent the X - coordinate value and the Y - coordinate value of the XYZ coordinate system after the axis transformation (after the coordinate - axis transformation process) of reference view - point A, and φa represents the rotation angle of the axis transformation of reference view - point A, that is, the above - mentioned rotation angle φ.
[0206] Thus, when substituting the x - coordinate and the y - coordinate constituting the object absolute - coordinate position information at reference view - point A obtained in process PC1 into formula (3), the coordinates xa and ya are obtained as the X - coordinate and the Y - coordinate representing the position of the object in the XYZ coordinate system (common absolute coordinate system) at reference view - point A. The absolute coordinates including the obtained coordinates xa and ya and the Z - coordinate are the object absolute - coordinate position information output from the coordinate - axis transformation processing unit 47.
[0207] Note that in this example, since only the rotation angle φ in the horizontal direction is processed, no coordinate - axis transformation is performed on the Z - axis (Z - coordinate). Therefore, for example, if the z - coordinate constituting the object absolute - coordinate position information obtained in process PC1 is used as it is as the Z - coordinate indicating the position of the object in the common absolute coordinate system, that is sufficient.
[0208] Similar to reference view - point A, in formula (3), xb and yb represent the X - coordinate value and the Y - coordinate value of the XYZ coordinate system after the axis transformation (after the coordinate - axis transformation process) of reference view - point B, and φb represents the rotation angle of the axis transformation of reference view - point B (rotation angle φ).
[0209] In the coordinate axis transformation processing unit 47, the above-described coordinate axis transformation processing is executed as processing PC2.
[0210] Therefore, for example, when the coordinate axis transformation processing is executed for each of the four reference viewpoints shown in Figure 3 , the transformation results shown in Figure 8 are obtained. Note that the part corresponding to the part in Figure 8 and Figure 3 is denoted by the same reference numeral, and description thereof is omitted as needed.
[0211] In Figure 8 , each circle (ring) represents an object. Further, in Figure 8 , the upper side of the drawing shows the position of each object in the polar coordinate system indicated by the object polar coordinate position information, and the lower side of the drawing shows the position of each object in the common absolute coordinate system.
[0212] Specifically, in Figure 8 , the left end shows the result of the coordinate axis transformation of the reference viewpoint "origin" at the position P11 shown in Figure 3 , and the second from the left in Figure 8 shows the result of the coordinate axis transformation of the reference viewpoint "Near" at the position P12 shown in Figure 3 .
[0213] In addition, in Figure 8 , the third from the left shows the result of the coordinate axis transformation of the reference viewpoint "Far" at the position P13 shown in Figure 3 , and the right end in Figure 8 shows the result of the coordinate axis transformation of the reference viewpoint "Back" at the position P14 shown in Figure 3 .
[0214] For example, with respect to the reference viewpoint "origin", since its position is the origin viewpoint where the position of the origin of the polar coordinate system is the position of the origin of the common absolute coordinate system, the position of the object observed from the origin does not change before and after the transformation. On the other hand, at the remaining three reference viewpoints "Near", "Far", and "Back", it can be seen that the position of the object is shifted to the absolute coordinate position observed from each viewpoint position. Specifically, at the reference viewpoint "Back", since the direction of the listener's face indicated by the listener direction information is backward, after the coordinate axis transformation processing, the object is located behind the reference viewpoint.
[0215] (Processing PC3)
[0216] In the processing PC3, the ratio of the interpolation process is obtained from the positional relationship between the absolute coordinate positions of each of the two reference viewpoints (i.e., the positions indicated by the reference viewpoint position information included in the system configuration information) and any listening position sandwiched between the positions of the two reference viewpoints.
[0217] That is, the object position calculation unit 48 performs the process of obtaining the ratio (m:n) as the processing PC3 based on the listener position information provided by the listener position information acquisition unit 41 and the reference viewpoint position information included in the system configuration information.
[0218] Here, it is assumed that the reference viewpoint position information indicating the position of the reference viewpoint A as the first reference viewpoint is (x1, y1, z1), the reference viewpoint position information indicating the position of the reference viewpoint B as the second reference viewpoint is (x2, y2, z2), and the listener position information indicating the listening position is (x3, y3, z3).
[0219] In this case, the object position calculation unit 48 calculates the ratio (m:n), that is, m and n of the ratio, by calculating the following formula (4).
[0220] [Mathematics 4]
[0221] m = SQRT((x3 - x1)*(x3 - x1)+(y3 - y1)*(y3 - y1)+(z3 - z1)*(z3 - z1))
[0222] n = SQRT((x3 - x2)*(x3 - x2)+(y3 - y2)*(y3 - y2)+(z3 - z2)*(z3 - z2))...(4)
[0223] (Processing PC4)
[0224] Subsequently, the object position calculation unit 48 performs the interpolation process as the processing PC4 based on the ratio (m:n) obtained through the processing PC3 and the object absolute coordinate position information of each object of the two reference viewpoints provided by the coordinate axis transformation processing unit 47.
[0225] That is, in the processing PC4, the object position and the gain amount corresponding to any listening position are obtained by applying the ratio (m:n) obtained in the processing PC3 to the same object corresponding to the two reference viewpoints obtained in the processing PC2.
[0226] Here, the absolute coordinate position of a predetermined object observed from the reference viewpoint A (i.e., the object absolute coordinate position information of the reference viewpoint A obtained through the processing PC2) is (xa, ya, za), and the gain amount indicated by the gain information of the predetermined object of the reference viewpoint A is g1.
[0227] Similarly, the absolute coordinate position of the above-mentioned predetermined object observed from the reference viewing point B (i.e., the object absolute coordinate position information of the reference viewing point B obtained by processing PC2) is (xb, yb, zb), and the gain amount indicated by the gain information of the object at the reference viewing point B is g2.
[0228] In addition, the absolute coordinates indicating the position of the above-mentioned predetermined object in the XYZ coordinate system (common absolute coordinate system) and the gain amount corresponding to an arbitrary viewing point position between the reference viewing point A and the reference viewing point B (i.e., the listening position indicated by the listener position information) are set to (xc, yc, zc) and gain_c. The absolute coordinates (xc, yc, zc) are the final object absolute coordinate position information output from the object position calculation unit 48 to the polar coordinate transformation unit 49.
[0229] At this time, the final object absolute coordinate position information (xc, yc, zc) and the gain amount gain_c of the predetermined object can be obtained by calculating the following formula (5) using the ratio (m:n).
[0230] [Mathematical formula 5]
[0231] xc = (m * xb + n * xa) / (m + n)
[0232] yc = (m * yb + n * ya) / (m + n)
[0233] zc = (m * zb + n * za) / (m + n)
[0234] gain_c = (m * g2 + n * g1) / (m + n)...(5)
[0235] The positional relationship among the above-mentioned reference viewing point A, reference viewing point B, and the listening position, and the positional relationship of the same object at the corresponding positions of the reference viewing point A, reference viewing point B, and the listening position are as Figure 9 shown.
[0236] In Figure 9 , the horizontal axis and the vertical axis respectively represent the X-axis and the Y-axis of the XYZ coordinate system (common absolute coordinate system). Note that for the sake of brief description, only the X-axis direction and the Y-axis direction are shown here.
[0237] In this example, the position P51 is the position represented by the reference viewing point position information (x1, y1, z1) of the reference viewing point A, and the position P52 is the position represented by the reference viewing point position information (x2, y2, z2) of the reference viewing point B.
[0238] In addition, the position P53 between the reference viewing point A and the reference viewing point B is the listening position indicated by the listener position information (x3, v3, z3).
[0239] In the above formula (4), the ratio (m:n) is obtained based on the positional relationship among the reference viewpoint A, the reference viewpoint B, and the listening position.
[0240] In addition, the position P61 is the position represented by the object absolute coordinate position information (xa, ya, za) of the reference viewpoint A, and the position P62 is the position represented by the object absolute coordinate position information (xb, yb, zb) of the reference viewpoint B.
[0241] In addition, the position P63 between the position P61 and the position P62 is the position indicated by the object absolute coordinate position information (xc, yc, zc) at the listening position.
[0242] By performing the calculation of formula (5) in this way, that is, the interpolation process, the object absolute coordinate position information indicating the appropriate object position can be obtained for any listening position.
[0243] Note that an example of using the ratio (m:n) to obtain the object position (i.e., the final object absolute coordinate position information) has been described above, but it is not limited to this, and the final object absolute coordinate position information can be estimated using machine learning or the like.
[0244] In addition, in the case of using the absolute coordinate system editor, that is, in Figure 2 the case of the content reproduction system shown, the position of each object of each reference viewpoint (i.e., the position represented by the object absolute coordinate position information) is a position on a common absolute coordinate system. In other words, the positions of the objects of each reference viewpoint are represented by the absolute coordinates of the common absolute coordinate system.
[0245] Therefore, in Figure 2 the content reproduction system shown, it is sufficient if the object absolute coordinate position information obtained by decoding by the decoding unit 45 is used as the input in the above processing PC3. That is, it is sufficient if the calculation of formula (4) is performed based on the object absolute coordinate position information obtained by decoding.
[0246] <Regarding the operation of the content reproduction system>
[0247] Next, with reference to Figure 10 , the processing flow (sequence) executed in the above content reproduction system will be described.
[0248] In addition, here, an example of selecting the reference viewpoint on the server 11 side and preparing the object polar coordinate encoded data in advance on the server 11 side will be described. That is, an example in which the viewpoint selection unit 42 in the content reproduction system shown in Figure 1 is provided on the server 11 side will be described.
[0249] First, on the server 11 side, for all reference viewpoints, the polar coordinate system editor generates and saves the polar coordinate system object position information, that is, the object polar coordinate encoding data, and also generates and saves the system configuration information.
[0250] Then, the configuration information sending unit 21 transmits the system configuration information to the client 12 via a network or the like.
[0251] Then, the configuration information acquisition unit 43 of the client 12 receives the system configuration information transmitted from the server 11 and provides the system configuration information to the coordinate axis transformation processing unit 47. At this time, the client 12 decodes (decodes) the received system configuration information and initializes the client system.
[0252] Subsequently, when the listener position information acquisition unit 41 acquires the listener position information and provides the listener position information to the configuration information acquisition unit 43, the configuration information acquisition unit 43 transmits the listener position information provided from the listener position information acquisition unit 41 to the server 11.
[0253] In addition, the configuration information sending unit 21 receives the listener position information transmitted from the client 12 and provides the listener position information to the viewpoint selection unit 42. Then, the viewpoint selection unit 42 selects the reference viewpoints required for interpolation processing based on the listener position information and the system configuration information provided from the configuration information sending unit 21, that is, for example, two reference viewpoints sandwiching the above-mentioned listening position, and provides the viewpoint selection information indicating the selection result to the encoding data sending unit 22.
[0254] The encoding data sending unit 22 prepares the polar coordinate system object position information of the reference viewpoints required for interpolation processing according to the viewpoint selection information provided from the viewpoint selection unit 42.
[0255] That is, the encoding data sending unit 22 generates a bit stream by reading and reusing the object polar coordinate encoding data and the encoding gain information of the reference viewpoints indicated by the viewpoint selection information. Then, the encoding data sending unit 22 sends the generated bit stream to the client 12.
[0256] The encoding data acquisition unit 44 receives and demultiplexes the bit stream transmitted from the server 11, and provides the obtained object polar coordinate encoding data and encoding gain information to the decoding unit 45.
[0257] The decoding unit 45 decodes the object polar coordinate encoded data provided by the encoded data acquisition unit 44 and provides the obtained object polar coordinate position information to the coordinate transformation unit 46. In addition, the decoding unit 45 decodes the encoded gain information provided by the encoded data acquisition unit 44 and provides the obtained gain information to the object position calculation unit 48 via the coordinate transformation unit 46 and the coordinate axis transformation processing unit 47.
[0258] The coordinate transformation unit 46 converts the object polar coordinate position information provided by the decoding unit 45 from polar coordinate information into absolute coordinate position information centered on the listener.
[0259] That is, for example, the coordinate transformation unit 46 calculates the above formula (1) based on the object polar coordinate position information and provides the obtained object absolute coordinate position information to the coordinate axis transformation processing unit 47.
[0260] Subsequently, the coordinate axis transformation processing unit 47 performs the expansion from the absolute coordinate position information centered on the listener to the common absolute coordinate space through coordinate axis transformation.
[0261] For example, the coordinate axis transformation processing unit 47 performs coordinate axis transformation processing by calculating the above formula (3) based on the system configuration information provided by the configuration information acquisition unit 43 and the object absolute coordinate position information provided by the coordinate transformation unit 46, and provides the obtained object absolute coordinate position information to the object position calculation unit 48.
[0262] The object position calculation unit 48 calculates the ratio for interpolation processing based on the current listener position and the reference viewpoints.
[0263] For example, the object position calculation unit 48 calculates the above formula (4) based on the listener position information provided by the listener position information acquisition unit 41 and the reference viewpoint position information of multiple reference viewpoints selected by the viewpoint selection unit 42, and calculates the ratio (m∶n).
[0264] In addition, the object position calculation unit 48 calculates the object position and gain amount corresponding to the current listener position using the ratio of the object position and gain amount from the reference viewpoints sandwiching the listener position.
[0265] For example, the object position calculation unit 48 performs interpolation processing by calculating the above formula (5) based on the object absolute coordinate position information and gain information provided by the coordinate axis transformation processing unit 47 and the ratio (m∶n), and provides the obtained final object absolute coordinate position information and gain information to the polar coordinate transformation unit 49.
[0266] Then, thereafter, the client 12 performs rendering processing applying the calculated object position and gain amount.
[0267] For example, the polar coordinate transformation unit 49 transforms the absolute coordinate position information into polar coordinates.
[0268] That is, for example, the polar coordinate transformation unit 49 performs a polar coordinate transformation on the object absolute coordinate position information provided from the object position calculation unit 48 based on the listener position information provided from the listener position information acquisition unit 41.
[0269] The polar coordinate transformation unit 49 provides the polar coordinate position information obtained by the polar coordinate transformation and the gain information provided from the object position calculation unit 48 to the subsequent rendering processing unit.
[0270] Then, the rendering processing unit performs polar coordinate rendering processing on all objects.
[0271] That is, the rendering processing unit performs rendering processing in the polar coordinate system defined by, for example, MPEG-H based on the polar coordinate position information and gain information of all objects provided from the polar coordinate transformation unit 49, and generates reproduction audio data for reproducing the sound of the content.
[0272] Here, for example, vector-based amplitude panning (VBAP) or the like is performed as the rendering processing in the polar coordinate system defined by MPEG-H. Note that, more specifically, before the rendering processing, gain adjustment based on the gain information is performed on the audio sound data, but the gain adjustment may also be performed not by the rendering processing unit but by the previous polar coordinate transformation unit 49.
[0273] When the above processing is performed on a predetermined frame and reproduction audio data is generated, content reproduction based on the reproduction audio data is appropriately performed. Then, hereinafter, the listener position information is appropriately transmitted from the client 12 to the server 11, and the above processing is repeatedly performed.
[0274] As described above, the content reproduction system calculates the object absolute coordinate position information and gain information at an arbitrary listening position from the object position information at a plurality of reference viewpoints through interpolation processing. In this way, an object layout based on the intention of the content creator can be realized according to the listening position, rather than a simple physical relationship between the listener and the object. Therefore, content reproduction based on the intention of the content creator can be realized, and the interestingness of the content can be fully conveyed to the listener.
[0275] <Regarding the listener and the object>
[0276] Incidentally, as reference viewpoints, for example, two examples of assuming the viewpoint as the listener and assuming the viewpoint of the performer as the object can be conceived.
[0277] In the latter case, since the listener and the object overlap at the reference viewpoint, that is, the listener and the object are located at the same position, the following cases CA1 to CA3 can be conceived.
[0278] (Case CA1)
[0279] The listener is prohibited from overlapping with the object, or the listener is prohibited from entering a specific range.
[0280] (Case CA2)
[0281] The listener and the object are merged, and the sound generated from the object is output from all channels.
[0282] (Case CA3)
[0283] The sound generated from the overlapping object is muted or attenuated.
[0284] For example, in the case of Case CA2, the feeling of being located in the listener's head can be recreated.
[0285] In addition, in Case CA3, by muting or attenuating the sound of the object, the listener becomes the performer, and for example, it can be conceived to be used in the karaoke mode. In this case, in addition to the singing voice of the performer, the surrounding accompaniment, etc. surround the listener himself / herself, and the feeling of singing can be obtained.
[0286] In the case where the content creator has such an intention, the identifiers indicating Cases CA1 to CA3 can be stored in the encoded bitstream transmitted from the server 11 and can be transmitted to the client 12 side. For example, such an identifier is information indicating the above reproduction mode.
[0287] In addition, in the above content reproduction system, the listener can move between two reference viewpoints.
[0288] In this case, there may be a situation where some listeners expect to intentionally make the object (viewpoint) closer to the object arrangement on one (side) of the two reference viewpoints. Specifically, for example, there may be a request to maintain an angle that allows the listener's favorite artist to be always easily visible.
[0289] Therefore, for example, the degree of approach can be controlled by biasing through ratio processing of the internal distribution ratio. For example, as Figure 11 shown, this can be achieved by newly introducing a bias coefficient α into the formula (5) for obtaining the above interpolation.
[0290] Figure 11 The characteristics in the case of multiplying by the bias coefficient α are shown. Specifically, the upper side in the drawing shows an example of the configuration that makes the object closer to the viewpoint X1 side (i.e., the above reference viewpoint A side).
[0291] On the other hand, the lower side in the drawings shows an example of an arrangement that brings an object closer to the viewpoint X2 side (i.e., the side of the above-mentioned reference viewpoint B).
[0292] Note that in Figure 11 , the horizontal axis represents the position of a predetermined viewpoint X3 without introducing the bias coefficient α, and the vertical axis represents the position of the predetermined viewpoint X3 when the bias coefficient α is introduced. Further, here, the position of the reference viewpoint A (viewpoint X1) is "0", and the position of the reference viewpoint B (viewpoint X2) is "1".
[0293] In the example on the upper side in the drawings, for example, when the listener moves from the reference viewpoint A (viewpoint X1) side to the position of the reference viewpoint B (viewpoint X2), the smaller the bias coefficient α, the more difficult it is for the listener to feel that they have reached the position of the reference viewpoint B (viewpoint X2).
[0294] Conversely, in the example on the lower side in the drawings, for example, when the listener moves from the reference viewpoint A side to the position of the reference viewpoint B, the smaller the bias coefficient α, the more immediately the listener feels that they have reached the position of the reference viewpoint B.
[0295] For example, in the case of an arrangement that brings an object closer to the reference viewpoint A side, the final object absolute coordinate position information (xc, yc, zc) and the gain amount gain_c can be obtained by calculating the following formula (6).
[0296] On the other hand, in the case of an arrangement that brings an object closer to the reference viewpoint B side, the final object absolute coordinate position information (xc, yc, zc) and the gain amount gain_c can be obtained by calculating the following formula (7).
[0297] However, in formulas (6) and (7), m and n of the ratio (m:n) and the bias coefficient α are as shown in formula (8) described below.
[0298] [Mathematics 6]
[0299] xc = (m * xb + α * n * xa) / (m + α * n)
[0300] yc = (m * yb + α * n * ya) / (m + α * n)
[0301] zc = (m * zb + α * n * za) / (m + α * n)
[0302] gain_c = (m * g2 + α * n * g1) / (m + α * n)...(6)
[0303] [Mathematics 7]
[0304] xc = (α * m * xb + n * xa) / (α * m + n)
[0305] yc = (α * m * yb + n * ya) / (α * m + n)
[0306] zc = (α * m * zb + n * za) / (α * m + n)
[0307] gain_c = (α * m * g2 + n * g1) / (α * m + n)...(7)
[0308] [Mathematics 8]
[0309] m = SQRT((x3 - x1) * (x3 - x1) + (y3 - y1) * (y3 - y1) + (z3 - z1) * (z3 - z1))
[0310] n = SQRT((x3 - x2) * (x3 - x2) + (y3 - y2) * (y3 - y2) + (z3 - z2) * (z3 - z2))
[0311] 0 < α ≤ 1...(8)
[0312] Note that in formula (8), the reference viewpoint position information (x1, y1, z1), the reference viewpoint position information (x2, y2, z2), and the listener position information (x3, y3, z3) are similar to those in the above formula (4).
[0313] As in formulas (6) and (7), obtaining the final object absolute coordinate position information and the gain amount by using the bias coefficient α is performed by giving weights to the object absolute coordinate position information and the gain information with respect to a predetermined reference viewpoint and performing an interpolation process to obtain the final object absolute coordinate position information and the gain amount.
[0314] When the object position information of the absolute coordinates after the interpolation process obtained in this way (i.e., the object absolute coordinate position information) is combined with the listener position information and converted into polar coordinate information (polar coordinate position information), the polar coordinate rendering process used in the existing MPEG-H can be performed in a subsequent stage.
[0315] <Interpolation Process for Object Absolute Coordinate Position Information and Gain Information>
[0316] Meanwhile, as an example of the object position calculation unit 48 obtaining the object absolute coordinate position information and the gain information at an arbitrary viewpoint position (i.e., the listening position) through an interpolation process, the two-point interpolation using the information of two reference viewpoints has been described above.
[0317] However, it is not limited thereto, and the absolute coordinate position information and gain information of the object at an arbitrary listening position can be obtained by performing three-point interpolation using the information of three reference viewpoints. In addition, the absolute coordinate position information and gain information of the object at an arbitrary listening position can be obtained by using the information of four or more reference viewpoints. Hereinafter, a specific example in the case of performing three-point interpolation will be described.
[0318] For example, as shown on the left side of Figure 12 , it is considered that the absolute coordinate position information of the object at an arbitrary listening position F is obtained by interpolation processing.
[0319] In this example, there are three reference viewpoints: reference viewpoint A, reference viewpoint B, and reference viewpoint C, so as to surround the listening position F, and here, it is assumed that interpolation processing is performed using the information of reference viewpoints A to C.
[0320] Hereinafter, it is assumed that the X coordinate and Y coordinate of the listening position F in the common absolute coordinate system (i.e., the XYZ coordinate system) are (x f , y f ).
[0321] Similarly, it is assumed that the X coordinate and Y coordinate of the corresponding positions of reference viewpoint A, reference viewpoint B, and reference viewpoint C are (x a , y a ), (x b , y b ), and (x c , y c ).
[0322] In this case, as shown on the right side of Figure 12 , the object position F′ at the listening position F is obtained based on the coordinates of the object positions A′, B′, and C′ corresponding to reference viewpoint A, reference viewpoint B, and reference viewpoint C, respectively.
[0323] Here, for example, the object position A′ represents the position of the object when the viewpoint is at reference viewpoint A, that is, the position of the object in the common absolute coordinate system represented by the absolute coordinate position information of the object at reference viewpoint A.
[0324] In addition, the object position F′ represents the position of the object in the common absolute coordinate system when the listener is at the listening position F, that is, the position represented by the absolute coordinate position information that is the output of the object position calculation unit 48.
[0325] Hereinafter, it is assumed that the X coordinate and Y coordinate of the object positions A′, B′, and C′ are (x a′ , y a′ ), (x b′ , y b′) and (x c′ , y c′ ), and the X coordinate and Y coordinate of the object position F′ are (x f′ , y f′ ).
[0326] In addition, hereinafter, a triangular region surrounded by any three reference viewpoints (such as reference viewpoints A to C) (i.e., a region having a triangular shape formed by three reference viewpoints) is also referred to as a triangular mesh.
[0327] Since there are multiple reference viewpoints in the common absolute coordinate space, multiple triangular meshes with the reference viewpoints as vertices can be formed in the common absolute coordinate space.
[0328] Similarly, hereinafter, a triangular region surrounded (formed) by object absolute coordinate position information of any three reference viewpoints (such as object positions A′ to C′) is also referred to as a triangular mesh.
[0329] For example, in the example of two-point interpolation, the listener can move to any position on the line segment connecting two reference viewpoints and listen to the sound of the content.
[0330] On the other hand, in the case of performing three-point interpolation, the listener can move to any position in the region of the triangular mesh surrounded by three reference viewpoints and listen to the sound of the content. That is, a region other than the line segment connecting two reference viewpoints in the case of two-point interpolation can be covered as the listening position.
[0331] In addition, in the case of performing three-point interpolation, similar to the case of two-point interpolation, the coordinates indicating an arbitrary position in the common absolute coordinate system (XYZ coordinate system) can be obtained from the coordinates of an arbitrary position in the xyz coordinate system, the listener direction information, and the reference viewpoint position information by the above formula (2).
[0332] Note that here, it is assumed that the Z coordinate value of the XYZ coordinate system is the same as the z coordinate value of the xyz coordinate system. However, in the case where the Z coordinate value and the z coordinate value are different, it is sufficient to obtain the Z coordinate value representing an arbitrary position by adding the Z coordinate value representing the position of the reference viewpoint in the XYZ coordinate system to the z coordinate value of the arbitrary position.
[0333] Ceva's theorem proves that when the internal division ratio of each side of the triangular mesh is appropriately determined, an arbitrary listening position in the triangular mesh formed by three reference viewpoints is uniquely determined by the intersection points of the line segments from each of the three vertices of the triangular mesh to each of the internal division points of the three sides not adjacent to the vertices.
[0334] When determining the configuration of the internal division ratios of the three sides of a triangular mesh according to the proof formula, this holds for all triangular meshes regardless of their shape.
[0335] Therefore, when obtaining the internal division ratio of a triangular mesh including the listening position with respect to the viewpoint side (i.e., the reference viewpoint) and applying the internal division ratio to the triangular mesh on the object side (i.e., the object position), an appropriate object position for any listening position can be obtained.
[0336] Hereinafter, an example of obtaining object absolute coordinate position information indicating the position of an object at any listening position using such an internal division ratio property will be described.
[0337] In this case, first, the internal division ratio of the sides of the triangular mesh of the reference viewpoint on the XY plane of the XYZ coordinate system as a two-dimensional space is obtained.
[0338] Next, on the XY plane, the above internal division ratio is applied to the triangular mesh of the object positions corresponding to the three reference viewpoints, and the X coordinate and Y coordinate of the position of the object corresponding to the listening position on the XY plane are obtained.
[0339] In addition, the Z coordinate of the object corresponding to the listening position is obtained based on the three-dimensional plane including the positions of the three objects corresponding to the three reference viewpoints in the three-dimensional space (XYZ coordinate system) and the X coordinate and Y coordinate of the object at the listening position on the XY plane.
[0340] Here, reference will be made to Figures 13 to 15 Describe an example of obtaining object absolute coordinate position information and gain information representing the object position F′ through interpolation processing of the listening position F shown in Figure 12 For example, as shown in
[0341] First, the X coordinate and Y coordinate of the internal division points in the triangular mesh including the reference viewpoints A to C (including the listening position F) are obtained. Figure 13 Now, the intersection point of the straight line passing through the listening position F and the reference viewpoint C and the line segment AB from the reference viewpoint A to the reference viewpoint B is defined as point D, and the coordinates representing the position of point D on the XY plane are defined as (x
[0342] d ), y d d ). That is, point D is the internal division point on the line segment AB (side AB).
[0343] At this time, for the X coordinate and Y coordinate representing the position of any point on the line segment CF from the reference viewpoint C to the listening position F, and the X coordinate and Y coordinate representing the position of any point on the line segment AB, the relationship shown in the following formula (9) holds.
[0344] [Mathematics 9]
[0345] Line segment CF: γ = α1X - α1x c +y c , where α1 = (y c -y f ) / (x c -x f )
[0346] Line segment AB: γ = α2X - α2x a +y a , where α2 = (y a -y a ) / (x b -x a )...(9)
[0347] In addition, since point D is the intersection of the line passing through the reference viewing point C and the listening position F and line segment AB, the coordinates (x d , y d ) of point D on the XY plane can be obtained from formula (9), and the coordinates (x d , y d ) are as shown in formula (10) described below.
[0348] [Mathematics 10]
[0349] x d = (α1x c -y c -α2x a +y a ) / (α1 - α2)
[0350] y d = α1x d -α′1x c +y c ...(10)
[0351] Therefore, as shown in formula (11) below, based on the coordinates (x d , y d ) of point D, the coordinates (x a , y a ) of the reference viewing point A, and the coordinates (x b , y b ) of the reference viewing point B, the internal division ratio (m, n) of line segment AB passing through point D, that is, the division ratio, can be obtained.
[0352] [Mathematics 11]
[0353] m = sqrt((x a -x d )2 +(y a -y d ) 2 )
[0354] n = sqrt((x b -x d ) 2 +(y b -y d ) 2 )...(11)
[0355] Similarly, the intersection point of the straight line passing through the listening position F and the reference viewpoint B and the line segment AC from the reference viewpoint A to the reference viewpoint C is defined as point E, and the coordinates representing the position of point E on the XY plane are defined as (x e , y e ). That is, point E is an internal dividing point on the line segment AC (side AC).
[0356] At this time, for the X coordinate and Y coordinate representing the position of any point on the line segment BF from the reference viewpoint B to the listening position F, and the X coordinate and Y coordinate representing the position of any point on the line segment AC, the relationship shown in the following formula (12) holds.
[0357] [Mathematics 12]
[0358] Line segment BF: γ = α3X - α3x b +y b , where α3 = (y b -y f ) / (x b -x f )
[0359] Line segment AC: γ = α4X - α4x a +y a , where α4 = (y c -y a ) / (x c -x a )...(12)
[0360] In addition, since point E is the intersection point of the straight line passing through the reference viewpoint B and the listening position F and the line segment AC, the coordinates (x e , y e ) of point E on the XY plane can be obtained from formula (12), and the coordinates (x e , y e ) are as shown in formula (13) described below.
[0361] [Mathematics 13]
[0362] x e=(α3x b -y b -α4x a +y a ) / (α3 - α4)
[0363] y e =α3x e -α3x b +y b ...(13)
[0364] Therefore, as shown in formula (14) described below, based on the coordinates (x e , y e ) of point E, the coordinates (x a , y a ) of reference view point A, and the coordinates (x c , y c ) of reference view point C, the internal division ratios (k, l) of line segment AC passing through point E, that is, the division ratios, can be obtained.
[0365] [Mathematics 14]
[0366] k = sqrt((x a - x e ) 2 + (y a - y e ) 2 )
[0367] l = sqrt((x c - x e ) 2 + (y c - y e ) 2 )...(14)
[0368] Next, by applying the ratios of the two sides obtained in this way (that is, the internal division ratios (m, n) and the internal division ratios (k, l)) to the object - side triangular mesh as Figure 14 shown, the coordinates (x f′ , y f′ ) of object position F' on the XY plane are obtained.
[0369] Specifically, in this example, the point corresponding to point D on line segment A'B' connecting object position A' and object position B' is point D'.
[0370] Similarly, the point corresponding to point E on line segment A'C' connecting object position A' and object position C' is point E'.
[0371] In addition, the intersection point between the straight line passing through the object position C′ and the point D′ and the straight line passing through the object position B′ and the point E′ is the object position F′ corresponding to the listening position F.
[0372] Here, it is assumed that the internal division ratio of the line segment A′B′ with respect to the point D′ is the same internal division ratio (m, n) as in the case of the point D. At this time, the coordinates (x d ′, y d ′) of the point D′ on the XY plane can be obtained based on the internal division ratio (m, n), the coordinates (x a ′, y a ′) of the object position A′, and the coordinates (x b ′, y b ′) of the object position B′, as shown in the following formula (15).
[0373] [Mathematics 15]
[0374] x d ’ = (nx a ’ + mx b ’) / (m + n)
[0375] y d ’ = (ny a ’ + my b ’) / (m + n)...(15)
[0376] In addition, it is assumed that the internal division ratio of the line segment A′C′ passing through the point E′ is the same internal division ratio (k, l) as in the case of the point E. At this time, the coordinates (x e ′, y e ′) of the point E′ on the XY plane can be obtained based on the internal division ratio (k, l), the coordinates (x a ′, y a ′) of the object position A′, and the coordinates (x c ′, y c ′) of the object position C”, as shown in the following formula (16).
[0377] [Mathematics 16]
[0378] x e ’ = (lx a ’ + kx c ’) / (k + l)
[0379] y e ’ = (ly a ’ + ky c ′) / (k + l)...(16)
[0380] Therefore, for the X coordinate and Y coordinate representing the position of any point on the line segment B'E' from the object position B' to the point E', and the X coordinate and Y coordinate representing the position of any point on the line segment C'D' from the object position C' to the point D', the relationship shown in the following formula (17) holds.
[0381] [Mathematics 17]
[0382] Line segment B'E': γ = α5X + y b ′ - α5x b ′, where α1 = (y e ′ - y b ′) / (x e ′ - x b ′)
[0383] Line segment C'D': γ = α6X + y c ′ - α6x c ′, where α6 = (y d ′ - y c ′) / (x d ′ - x c ′)...(17)
[0384] Since the target object position F' is the intersection point of the line segment B'E' and the line segment C'D', the coordinates (x f ′, y f ′) of the object position F' can be obtained by the following formula (18) according to the relationship of formula (17).
[0385] [Mathematics 18]
[0386] x f ’ = (-y b ’ + α5x b ’ + y c ’ - α6x c ’) / (α5 - α6)
[0387] y f ’ = α6x f ’ + y c ’ - α6x c ’...(18)
[0388] Through the above processing, the coordinates (x f ′, y f ′) of the object position F' on the XY plane are obtained.
[0389] Subsequently, based on the coordinates (x f ′, y f ′) of the object position F' on the XY plane, the coordinates (x a ′, ya ′, z a ′), the coordinates (x b ′, y b ′, z b ′) of the object position B′, and the coordinates (x c ′, y c ′, z c ′) of the object position C′ in the XYZ coordinate system, to obtain the coordinates (x f ′, y f ′, z f ′) of the object position F′ in the XYZ coordinate system. That is, obtain the Z coordinate z f ′ of the object position F′ in the XYZ coordinate system.
[0390] For example, obtain a triangle on a three-dimensional space having the object positions A′, B′, and C′ as vertices in the XYZ coordinate system (common absolute coordinate space), that is, a three-dimensional plane A′B′C′ including the object positions A′, B′, and C′. Then, obtain a point having X and Y coordinates (x f ′, y f ′) on the three-dimensional plane A′B′C′, and the Z coordinate of this point is z f ′.
[0391] Specifically, set the vector starting from the object position A′ and ending at the object position B′ in the XYZ coordinate system as the vector A′B′ = (x ab ′, y ab ′, z ab ′).
[0392] Similarly, set the vector starting from the object position A′ and ending at the object position C′ in the XYZ coordinate system as the vector A′C′ = (x ac ′, y ac ′, z ac ′).
[0393] These vectors A′B′ and A′C′ can be obtained based on the coordinates (x a ′, y a ′, z a ′) of the object position A′, the coordinates (x b ′, y b ′, z b ′) of the object position B′, and the coordinates (x c ′, y c ′, z c ′) of the object position C′. That is, the vectors A′B′ and A′C′ can be obtained by the formula (19) described below.
[0394] [Mathematics 19]
[0395] Vector A'B': (x ab ', y ab ', z ab ') = (x b ' - x a ', y b ' - y a ', z b ' - z a )
[0396] Vector A'C': (x ac ', y ac ', z ac ') = (x c ' - x a ', y c ' - y a ', z c ' - z a )...(19)
[0397] In addition, the normal vector (s, t, u) of the three-dimensional plane A'B'C' is the cross product of vectors A'B' and A'C', and can be obtained by the formula (20) described below.
[0398] [Mathematics 20]
[0399] (s, t, u) = (y ab 'z ac ' - z ab 'y ac ', z ab 'x ac ' - x ab 'z ac ', x ab 'y ac ' - y ab 'x ac )...(20)
[0400] Therefore, based on the normal vector (s, t, u) and the coordinates (x a ', y a ', z a ) of the object position A', the plane equation of the three-dimensional plane A'B'C' is shown in the formula (21) described below.
[0401] [Mathematics 21]
[0402] s(X - x a ') + t(Y - y a ') + u(Z - z a ) = 0...(21)
[0403] Here, since the X coordinate x f ′ and the Y coordinate y f ′ of the object position F′ on the three-dimensional plane A′B′C′ have been obtained, by substituting the X coordinate x f ′ and the Y coordinate y f ′ into X and Y of the plane equation of formula (21), the Z coordinate z f ′ can be obtained as shown in formula (22) described below.
[0404] [Mathematics 22]
[0405] Z f ’ = (-s(x f ’ - x a ’)-t(y f ’ - y a ’)) / u + z a ’...(22)
[0406] Through the above calculation, the coordinates (xf′, yf′, zf′) of the target object position F′ are obtained. The object position calculation unit 48 outputs object absolute coordinate position information indicating the coordinates (xf′, yf′, zf′) of the object position F′ obtained in the above manner.
[0407] In addition, similar to the case of object absolute coordinate position information, the gain information can also be obtained by three-point interpolation.
[0408] That is, when the viewing point is at each of the reference viewing points A to C, the gain information of the object at the object position F′ can be obtained by performing interpolation processing based on the gain information of the object.
[0409] For example, as Figure 15 shown, consider obtaining the gain information G f ′ of the object at the object position F′ in the triangular mesh formed by the object positions A′, B′, and C′.
[0410] Now, assume that the gain information of the object at the object position A′ when the viewing point is at the reference viewing point A is G a ′, the gain information of the object at the object position B′ is G b ′, and the gain information of the object at the object position C′ is G c ′.
[0411] In this case, first, obtain the gain information G d ′ of the object at the point D′, which is the internal division point of the line segment A′B′ when the viewing point is virtually at the point D.
[0412] Specifically, the gain information G dIt can be obtained by calculating the following formula (23) based on the internal division ratio (m, n) of the above line segment A'B' and the gain information G of the object position A' a ' and the gain information G of the object position B' b '.
[0413] [Mathematics 23]
[0414] G d ’ = (m * G b ’ + n * G a ’) / (m + n)...(23)
[0415] That is, in formula (23), the gain information G of point D' is obtained through interpolation processing based on the gain information G a ’ and the gain information G b ’. d ’.
[0416] Next, based on the internal division ratio (o, p) of the object position F' of the line segment C'D' from the object position C' to point D' and the gain information G of the object position C' c ’ and the gain information G of point D' d ’, interpolation processing is performed, and the gain information G of the object position F' is obtained f ’. That is, the gain information G f ’ is obtained by performing the calculation of formula (24) described below.
[0417] [Mathematics 24]
[0418] G f ’ = (o * G c ’ + P * G d ’) / (o + p)
[0419] Where
[0420] o = SQRT((x d ’ - x f ’) 2 + (y d ’ - y f ’) 2 + (z d ’ - z f ') 2 )
[0421] p = SQRT((x c ’ - x f ') 2 + (y c ’ - y f ’) 2 + (z c ’ - zf ) 2 )...(24)
[0422] The obtained gain information G f ′ is output from the object position calculation unit 48 as the gain information for the object corresponding to the listening position F.
[0423] By performing the three-point interpolation as described above, the object absolute coordinate position information and the gain information for any listening position can be obtained.
[0424] Meanwhile, in the case of performing the three-point interpolation, when there are four or more reference viewpoints in the common absolute coordinate space, a plurality of triangular meshes can be configured by combinations of the three selected reference viewpoints.
[0425] For example, as shown on the left side of Figure 16 , it is assumed that there are reference viewpoints at five positions P91 to P95.
[0426] In this case, a plurality of triangular meshes, such as triangular meshes MS11 to MS13, are formed (configured).
[0427] Here, the triangular mesh MS11 is formed by the positions P91 to P93 as reference viewpoints, the triangular mesh MS12 is formed by the positions P92, P93, and P95, and the triangular mesh MS13 is formed by the positions P93, P94, and P95.
[0428] The listener can move freely within the area surrounded by the triangular meshes MS11 to MS13 (i.e., within the area surrounded by all the reference viewpoints).
[0429] Therefore, as the listener moves, that is, as the listening position moves (changes), the triangular mesh used to obtain the object absolute coordinate position information and the gain information at the listening position is switched.
[0430] Note that hereinafter, the viewpoint-side triangular mesh used to obtain the object absolute coordinate position information and the gain information at the listening position is also referred to as the selected triangular mesh. In addition, the object-side triangular mesh corresponding to the selected triangular mesh on the viewpoint side is also appropriately referred to as the selected triangular mesh.
[0431] Figure 16 The left side of
[0432] In the case of selecting a triangular mesh for performing three-point interpolation, basically, the sum (total) of the distances from the listening position to the corresponding vertices of the triangular mesh is obtained as the total distance, and the triangular mesh including the listening position and having the minimum total distance is selected as the selected triangular mesh.
[0433] That is, basically, the selected triangular mesh is determined by a conditional process of selecting the triangular mesh having the minimum total distance from the triangular meshes including the listening position. Hereinafter, the condition of having the minimum total distance among the triangular meshes including the listening position is also particularly referred to as the viewpoint-side selection condition.
[0434] When performing three-point interpolation, basically, the triangular mesh satisfying such a viewpoint-side selection condition is selected as the selected triangular mesh.
[0435] Therefore, in Figure 16 the example shown on the left side, when the listening position is at position P96, the triangular mesh MS11 is selected as the selected triangular mesh, and when the listening position moves to position P96', the triangular mesh MS13 is selected as the selected triangular mesh.
[0436] However, when the triangular mesh having the minimum total distance is simply selected as the selected triangular mesh, a discontinuous transition of the object position, that is, a jump in the position of the object, may occur.
[0437] For example, as Figure 16 shown in the center of, it is assumed that there are triangular meshes MS21 to MS23 as object-side triangular meshes, that is, triangular meshes including the object positions corresponding to each reference viewpoint.
[0438] In this example, the triangular mesh MS21 and the triangular mesh MS22 are adjacent to each other, and the triangular mesh MS22 and the triangular mesh MS23 are also adjacent to each other.
[0439] That is, the triangular mesh MS21 and the triangular mesh MS22 have a common edge with each other, and the triangular mesh MS22 and the triangular mesh MS23 also have a common edge with each other. Hereinafter, the common edge of two adjacent triangular meshes is also particularly referred to as the common edge.
[0440] On the other hand, since the triangular mesh MS21 and the triangular mesh MS23 are not adjacent to each other, these two triangular meshes do not have a common edge.
[0441] Here, it is assumed that the triangular mesh MS21 is the object-side triangular mesh corresponding to the viewpoint-side triangular mesh MS11. That is, it is assumed that the triangular mesh having each object position of the same object as vertices when the viewpoint (listening position) is at each of the positions P91 to P93 as the reference viewpoints is the triangular mesh MS21.
[0442] Similarly, the triangular mesh MS22 is the object-side triangular mesh corresponding to the viewpoint-side triangular mesh MS12, and the triangular mesh MS23 is the object-side triangular mesh corresponding to the viewpoint-side triangular mesh MS13.
[0443] For example, assume that the listening position moves from the position P96 to the position P96′, so that the triangular mesh selected on the viewpoint side switches from the triangular mesh MS11 to the triangular mesh MS13. In this case, on the object side, the selected triangular mesh switches from the triangular mesh MS21 to the triangular mesh MS23.
[0444] In the central example in the drawings, the position P101 represents the object position when the listening position is at the position P96, and this object position is obtained by performing three-point interpolation using the triangular mesh MS21 as the selected triangular mesh. Similarly, the position P101′ represents the object position when the listening position is at the position P96′, and this object position is obtained by performing three-point interpolation using the triangular mesh MS23 as the selected triangular mesh.
[0445] Therefore, in this example, when the listening position moves from the position P96 to the position P96′, the object position moves from the position P101 to the position P101′.
[0446] However, in this case, the triangular mesh MS21 including the position P101 and the triangular mesh MS23 including the position P101′ are not adjacent to each other and do not have a common edge in common. In other words, the object position moves (transfers) across the triangular mesh MS22 existing between the triangular meshes.
[0447] Therefore, in this case, a discontinuous movement (transition) of the object position occurs. This is because the triangular mesh MS21 and the triangular mesh MS23 do not have a common edge, and thus, the scale (measurement) of the relationship of the object positions corresponding to the respective reference viewpoints is different between the triangular meshes.
[0448] On the other hand, when the selected triangular meshes on the object side have a common edge, the continuity of the scale is maintained between the selected triangular meshes before and after the movement of the listening position, and the occurrence of discontinuous transfer of the object position can be suppressed.
[0449] Therefore, in the case of performing three-point interpolation, it is sufficient to add not only the above-described basic condition processing but also condition processing for selecting the triangular mesh selected on the viewpoint side after the movement so that the selected triangular meshes on the object side have a common edge before and after the movement of the listening position.
[0450] In other words, it suffices to select the selected triangular mesh for three-point interpolation at the post-move viewpoint based on the relationship between the triangular mesh selected on the object side through three-point interpolation at the viewpoint (listening position) before the movement and the object-side triangular mesh corresponding to the viewpoint-side triangular mesh including the post-move viewpoint position (listening position).
[0451] Hereinafter, the condition that the object-side triangular mesh before the movement of the listening position and the object-side triangular mesh after the movement of the listening position have a common side is also particularly referred to as the object-side selection condition.
[0452] In the case of performing three-point interpolation, if, among the viewpoint-side triangular meshes that satisfy the object-side selection condition, a triangular mesh that further satisfies the viewpoint-side selection condition is selected as the selected triangular mesh, that is sufficient. However, in the case where there is no viewpoint-side triangular mesh that satisfies the object-side selection condition, only the triangular mesh that satisfies the viewpoint-side selection condition is selected as the selected triangular mesh.
[0453] As described above, when the selected triangular mesh on the viewpoint side is selected such that it satisfies not only the viewpoint-side selection condition but also the object-side selection condition, the occurrence of discontinuous movement of the object position can be suppressed, and higher-quality acoustic reproduction can be achieved.
[0454] In this case, for example, in the example shown on the left side of Figure 16 , when the listening position moves from position P96 to position P96′, with respect to position P96′ which is the post-move listening position, the triangular mesh MS12 is selected as the selected triangular mesh on the viewpoint side.
[0455] For example, as shown on the right side of Figure 16 , the object-side triangular mesh MS21 corresponding to the viewpoint-side triangular mesh MS11 before the movement and the object-side triangular mesh MS22 corresponding to the viewpoint-side triangular mesh MS12 after the movement have a common side. Therefore, in this case, it can be seen that the object-side selection condition is satisfied.
[0456] Furthermore, position P101″ represents the object position when the listening position is at position P96′, and the object position is obtained by performing three-point interpolation using the triangular mesh MS22 as the selected triangular mesh on the object side.
[0457] Therefore, in this example, when the listening position moves from position P96 to position P96′, the position of the object corresponding to the listening position also moves from position P101 to position P101″.
[0458] In this case, since the triangle mesh MS21 and the triangle mesh MS22 have a common edge, discontinuous movement of the object position does not occur before and after the movement of the listening position.
[0459] For example, in this example, the positions of both ends of the common edge of triangular mesh MS21 and triangular mesh MS22 (i.e., the object position corresponding to position P92 as the reference viewpoint, and the object position corresponding to position P93 as the reference viewpoint) are the same positions before and after the listening position is moved.
[0460] As mentioned above, in Figure 16 In the example shown, even when the listening position is the same position at the position P96', the object position (ie, the position where the object is projected) changes depending on which of the triangle mesh MS12 and the triangle mesh MS13 is selected as the triangle mesh selected on the viewpoint side.
[0461] Therefore, by selecting a more appropriate triangular mesh from among the triangular meshes including the listening position, it is possible to suppress the occurrence of discontinuous movement of the object position (ie, the sound image position) and achieve higher quality acoustic reproduction.
[0462] Furthermore, by combining the use of three-point interpolation including a triangular mesh surrounding three reference viewpoints of the listening position and selecting the triangular mesh according to the selection condition, object arrangement considering reference viewpoints of an arbitrary listening position in a common absolute coordinate space can be achieved.
[0463] Note that also in the case of performing three-point interpolation, similarly to the case of performing two-point interpolation, interpolation processing weighted based on the bias coefficient α may be appropriately performed to obtain final object absolute coordinate position information and gain information.
[0464] <Configuration Example of Content Reproduction System>
[0465] Here, a more detailed embodiment of a content reproduction system to which the above-described present technology is applied will be described.
[0466] Figure 17 : is a diagram showing a configuration example of a content reproduction system to which the present technology is applied. Note that Figure 17 Zhongyu Figure 1 The corresponding parts are denoted by the same reference numerals, and the description is appropriately omitted.
[0467] Figure 17 The content reproduction system shown includes a server 11 that distributes content and a client 12 that receives content distribution from the server 11 .
[0468] Furthermore, the server 11 includes a configuration information recording unit 101 , a configuration information transmitting unit 21 , a recording unit 102 , and an encoded data transmitting unit 22 .
[0469] The configuration information recording unit 101 records, for example, Figure 4 the prepared system configuration information as shown, and provides the recorded system configuration information to the configuration information sending unit 21. Note that a part of the recording unit 102 can be the configuration information recording unit 101.
[0470] The recording unit 102 records, for example, the encoded audio data obtained by encoding the audio data of the objects that make up the content, the object polar coordinate encoding data of each object for each reference viewpoint, the encoding gain information, etc.
[0471] The recording unit 102 provides the encoded audio data, the object polar coordinate encoding data, the encoding gain information, etc. recorded in response to a request or the like to the encoded data sending unit 22.
[0472] In addition, the client 12 includes a listener position information acquisition unit 41, a viewpoint selection unit 42, a communication unit 111, a decoding unit 45, a position calculation unit 112, and a rendering processing unit 113.
[0473] The communication unit 111 corresponds to Figure 1 the configuration information acquisition unit 43 and the encoded data acquisition unit 44 as shown, and transmits and receives various data by communicating with the server 11.
[0474] For example, the communication unit 111 transmits the viewpoint selection information provided by the viewpoint selection unit 42 to the server 11, and receives the system configuration information and the bitstream transmitted from the server 11. That is, the communication unit 111 serves as a reference viewpoint information acquisition unit that acquires the system configuration information and the object polar coordinate encoding data and the encoding gain information included in the bitstream from the server 11.
[0475] The position calculation unit 112 generates polar coordinate position information indicating the object position based on the object polar coordinate position information provided by the decoding unit 45 and the system configuration information provided by the communication unit 111, and provides the polar coordinate position information to the rendering processing unit 113.
[0476] In addition, the position calculation unit 112 performs gain adjustment on the audio data of the object provided by the decoding unit 45, and provides the audio data after the gain adjustment to the rendering processing unit 113.
[0477] The position calculation unit 112 includes a coordinate transformation unit 46, an axis transformation processing unit 47, an object position calculation unit 48, and a polar coordinate transformation unit 49.
[0478] The rendering processing unit 113 performs rendering processing such as VBAP based on the polar coordinate position information and audio data provided by the polar coordinate transformation unit 49, and generates and outputs reproduction audio data for reproducing the sound of the content.
[0479] <Description of the provision processing and reproduction audio data generation processing>
[0480] Subsequently, the operation of the content reproduction system shown will be described. Figure 17 The operation of the content reproduction system shown will be described.
[0481] That is, the provision processing performed by the server 11 and the reproduction audio data generation processing performed by the client 12 will be described below with reference to Figure 18 the flowchart of.
[0482] For example, when a request for distribution of predetermined content is made from the client 12 to the server 11, the server 11 starts the provision processing and performs the processing of step S41.
[0483] That is, in step S41, the configuration information transmission unit 21 reads the system configuration information of the requested content from the configuration information recording unit 101, and transmits the read system configuration information to the client 12. For example, the system configuration information is prepared in advance, and immediately after starting the operation of the content reproduction system, that is, for example, immediately after establishing a connection between the server 11 and the client 12 and before transmitting the encoded audio data etc., the system configuration information is transmitted to the client 12 via a network or the like.
[0484] Then, in step S61, the communication unit 111 of the client 12 receives the system configuration information transmitted from the server 11 and provides the system configuration information to the viewpoint selection unit 42, the coordinate axis transformation processing unit 47, and the object position calculation unit 48.
[0485] Note that the timing at which the communication unit 111 obtains the system configuration information from the server 11 can be any timing as long as it is before the start of content reproduction.
[0486] In step S62, the listener position information acquisition unit 41 acquires listener position information based on the operation of the listener etc., and provides the listener position information to the viewpoint selection unit 42, the object position calculation unit 48, and the polar coordinate transformation unit 49.
[0487] In step S63, the viewpoint selection unit 42 selects two or more reference viewpoints based on the system configuration information provided by the communication unit 111 and the listener position information provided by the listener position information acquisition unit 41, and provides the viewpoint selection information indicating the selection result to the communication unit 111.
[0488] For example, in the case of selecting two reference viewpoints for the listening position represented by the listener position information, two reference viewpoints sandwiching the listening position are selected from the plurality of reference viewpoints represented by the system configuration information. That is, the reference viewpoints are selected such that the listening position lies on the line segment connecting the two selected reference viewpoints.
[0489] In addition, in the case where three-point interpolation is performed in the object position calculation unit 48, three or more reference viewpoints surrounding the listening position represented by the listener position information are selected from the plurality of reference viewpoints represented by the system configuration information.
[0490] In step S64, the communication unit 111 transmits the viewpoint selection information provided by the viewpoint selection unit 42 to the server 11.
[0491] Then, the process of step S42 is executed in the server 11. That is, in step S42, the configuration information transmission unit 21 receives the viewpoint selection information transmitted from the client 12 and provides the viewpoint selection information to the encoded data transmission unit 22.
[0492] The encoded data transmission unit 22 reads, for each object, the object polar coordinate encoded data and the encoding gain information of the reference viewpoint represented by the viewpoint selection information provided by the configuration information transmission unit 21 from the recording unit 102, and also reads the encoded audio data of each object of the content.
[0493] In step S43, the encoded data transmission unit 22 multiplexes the object polar coordinate encoded data, the encoding gain information, and the encoded audio data read from the recording unit 102 to generate a bitstream.
[0494] In step S44, the encoded data transmission unit 22 transmits the generated bitstream to the client 12 and provides the end of processing. Thus, the content is distributed to the client 12.
[0495] In addition, when transmitting the bitstream, the client 12 executes the process of step S65. That is, in step S65, the communication unit 111 receives the bitstream transmitted from the server 11 and provides the bitstream to the decoding unit 45.
[0496] In step S66, the decoding unit 45 extracts the object polar coordinate encoded data, the encoding gain information, and the encoded audio data from the bitstream provided by the communication unit 111, and decodes the object polar coordinate encoded data, the encoding gain information, and the encoded audio data.
[0497] The decoding unit 45 provides the object polar coordinate position information obtained by decoding to the coordinate transformation unit 46, provides the gain information obtained by decoding to the object position calculation unit 48, and provides the audio data obtained by decoding to the polar coordinate transformation unit 49.
[0498] In step S67, the coordinate transformation unit 46 performs coordinate transformation on the object polar coordinate position information of each object provided by the decoding unit 45, and provides the obtained object absolute coordinate position information to the coordinate axis transformation processing unit 47.
[0499] For example, in step S67, for each reference view point, the above formula (1) is calculated based on the object polar coordinate position information of each object, and the object absolute coordinate position information is calculated.
[0500] In step S68, the coordinate axis transformation processing unit 47 performs coordinate axis transformation processing on the object absolute coordinate position information provided by the coordinate transformation unit 46 based on the system configuration information provided by the communication unit 111.
[0501] The coordinate axis transformation processing unit 47 performs coordinate axis transformation processing on each object for each reference view point, and provides the resulting object absolute coordinate position information indicating the position of the object in the common absolute coordinate system to the object position calculation unit 48. For example, in step S68, a calculation similar to the above formula (3) is performed to calculate the object absolute coordinate position information.
[0502] In step S69, the object position calculation unit 48 performs interpolation processing based on the system configuration information provided by the communication unit 111, the listener position information provided by the listener position information acquisition unit 41, the object absolute coordinate position information provided by the coordinate axis transformation processing unit 47, and the gain information provided by the decoding unit 45.
[0503] In step S69, the above two-point interpolation or three-point interpolation is performed for each object as the interpolation processing, and the final object absolute coordinate position information and gain information are calculated.
[0504] For example, in the case of performing two-point interpolation, the object position calculation unit 48 obtains the ratio (m:n) by performing a calculation similar to the above formula (4) based on the reference view point position information and the listener position information included in the system configuration information.
[0505] Then, the object position calculation unit 48 performs the interpolation processing of two-point interpolation by performing a calculation similar to the above formula (5) based on the obtained ratio (m:n) and the object absolute coordinate position information and gain information of the two reference view points.
[0506] Note that by performing a calculation similar to formula (6) or (7) instead of formula (5), interpolation processing (two-point interpolation) can be performed by weighting the object absolute coordinate position information and gain information of the desired reference view point.
[0507] In addition, for example, in the case of performing three-point interpolation, the object position calculation unit 48 selects three reference viewpoints for forming (configuring) a triangular mesh that satisfies the viewpoint-side and object-side selection conditions based on the listener position information, the system configuration information, and the object absolute coordinate position information of each reference viewpoint. Then, the object position calculation unit 48 performs three-point interpolation based on the object absolute coordinate position information and the gain information of the three selected reference viewpoints.
[0508] That is, the object position calculation unit 48 obtains the internal division ratios (m, n) and the internal division ratios (k, l) by performing calculations similar to the above formulas (9) to (14) based on the reference viewpoint position information and the listener position information included in the system configuration information.
[0509] Then, the object position calculation unit 48 performs the interpolation process of three-point interpolation by performing calculations similar to the above formulas (15) to (24) based on the obtained internal division ratios (m, n) and the internal division ratios (k, l), the object absolute coordinate position information of each reference viewpoint, and the gain information. Note that also in the case of performing three-point interpolation, the interpolation process (three-point interpolation) can be performed by weighting the object absolute coordinate position information and the gain information of the desired reference viewpoint.
[0510] When the interpolation process is performed in this way and the final object absolute coordinate position information and the gain information are obtained, the object position calculation unit 48 provides the obtained object absolute coordinate position information and the gain information to the polar coordinate transformation unit 49.
[0511] In step S70, the polar coordinate transformation unit 49 performs a polar coordinate transformation on the object absolute coordinate position information provided from the object position calculation unit 48 based on the listener position information provided from the listener position information acquisition unit 41 to generate polar coordinate position information.
[0512] In addition, the polar coordinate transformation unit 49 performs a gain adjustment on the audio data of each object provided from the decoding unit 45 based on the gain information of each object provided from the object position calculation unit 48.
[0513] The polar coordinate transformation unit 49 provides the polar coordinate position information obtained by the polar coordinate transformation and the audio data of each object obtained by the gain adjustment to the rendering processing unit 113.
[0514] In step S71, the rendering processing unit 113 performs rendering processing such as VBAP based on the polar coordinate position information and the audio data of each object provided from the polar coordinate transformation unit 49, and outputs the resulting reproduced audio data.
[0515] For example, in a subsequent stage of the rendering processing unit 113, the sound of the content is reproduced based on the reproduced audio data by using a speaker or the like. When the reproduced audio data is generated and output in this way, the reproduced audio data generation processing ends.
[0516] Note that before the rendering process, the rendering processing unit 113 or the polar coordinate transformation unit 49 may perform processing corresponding to the reproduction mode on the audio data of the object based on the listener position information and the information indicating the reproduction mode included in the system configuration information.
[0517] In this case, for example, attenuation processing such as gain adjustment is performed on the audio data of the object located at the position overlapping the listening position, or the audio data is replaced with zero data and muted. In addition, for example, the sound of the audio data of the object located at the position overlapping the listening position is output from all channels (speakers).
[0518] In addition, the above-described provision processing and reproduced audio data generation processing are performed for each frame of the content.
[0519] However, the processing in steps S41 and S61 may be performed only when starting the reproduction of the content. In addition, it is not necessary to perform the processing of steps S42 and steps S62 to S64 for each frame.
[0520] As described above, the server 11 receives the viewpoint selection information, generates a bitstream including information on the reference viewpoint corresponding to the viewpoint selection information, and sends the bitstream to the client 12. In addition, the client 12 performs interpolation processing based on the information on each reference viewpoint included in the received bitstream, and obtains the object absolute coordinate position information and gain information of each object.
[0521] In this way, it is possible to realize the object layout based on the intention of the content creator according to the listening position, rather than the simple physical relationship between the listener and the object. Therefore, it is possible to realize the content reproduction based on the intention of the content creator, and the interest of the content can be fully conveyed to the listener.
[0522] <Description of viewpoint selection processing>
[0523] In addition, as described above, in the Figure 18 described reproduced audio data generation processing, when performing three-point interpolation in step S69, three reference viewpoints for performing three-point interpolation are selected.
[0524] Hereinafter, the viewpoint selection processing, which is the processing in which the client 12 selects three reference viewpoints when performing three-point interpolation, will be described with reference to the Figure 19 flowchart. This viewpoint selection processing corresponds to the Figure 18 processing of step S69.
[0525] In step S101, the object position calculation unit 48 calculates the distances from the listening position to each of a plurality of reference viewpoints based on the listener position information provided by the listener position information acquisition unit 41 and the system configuration information provided by the communication unit 111.
[0526] In step S102, the object position calculation unit 48 determines whether the frame of the audio data for which three-point interpolation is to be performed (hereinafter, also referred to as the current frame) is the first frame of the content.
[0527] When it is determined in step S102 that the frame is the first frame, the process proceeds to step S103.
[0528] In step S103, the object position calculation unit 48 selects a triangular mesh having the minimum total distance from a triangular mesh including any three reference viewpoints among the plurality of reference viewpoints. Here, the total distance is the sum of the distances from the listening position to the reference viewpoints constituting the triangular mesh.
[0529] In step S104, the object position calculation unit 48 determines whether the listening position is within (including within) the triangular mesh selected in step S103.
[0530] When it is determined in step S104 that the listening position is not within the triangular mesh, since the triangular mesh does not satisfy the viewpoint-side selection condition, thereafter, the process proceeds to step S105.
[0531] In step S105, the object position calculation unit 48 selects a triangular mesh having the minimum total distance from the viewpoint-side triangular meshes that have not been selected in the processing of steps S103 and S105 performed on the frame to be processed so far.
[0532] When a new viewpoint-side triangular mesh is selected in step S105, thereafter, the process returns to step S104, and the above processing is repeatedly executed until it is determined that the listening position is within the triangular mesh. That is, a triangular mesh that satisfies the viewpoint-side selection condition is searched for.
[0533] On the other hand, when it is determined in step S104 that the listening position is within the triangular mesh, the triangular mesh is selected as the triangular mesh for performing three-point interpolation, and thereafter, the process proceeds to step S110.
[0534] In addition, when it is determined in step S102 that the frame is not the first frame, thereafter, the processing of step S106 is executed.
[0535] In step S106, the object position calculation unit 48 determines whether the current listening position is within the viewpoint-side triangular mesh selected in the frame immediately preceding the current frame (hereinafter, also referred to as the previous frame).
[0536] When it is determined in step S106 that the listening position is within the triangular grid, the process then proceeds to step S107.
[0537] In step S107, the object position calculation unit 48 selects the same viewpoint-side triangular grid that was selected for three-point interpolation in the previous frame as the triangular grid for performing three-point interpolation also in the current frame. When the triangular grid for three-point interpolation (i.e., the three reference viewpoints) is selected in this way, the process then proceeds to step S110.
[0538] In addition, when it is determined in step S106 that the listening position is not within the viewpoint-side triangular grid selected in the previous frame, the process then proceeds to step S108.
[0539] In step S108, the object position calculation unit 48 determines whether there is a triangular grid in the object-side triangular grid in the current frame that has a common edge with the triangular grid (including) selected on the object side in the previous frame. The determination process in step S108 is performed based on the system configuration information and the object absolute coordinate position information.
[0540] When it is determined in step S108 that there is no triangular grid with a common edge, since there is no triangular grid that satisfies the object-side selection condition, the process then proceeds to step S103. In this case, only the triangular grid that satisfies the viewpoint-side selection condition is selected for three-point interpolation in the current frame.
[0541] In addition, when it is determined in step S108 that there is a triangular grid with a common edge, the process then proceeds to step S109.
[0542] In step S109, the object position calculation unit 48 selects the triangular grid that includes the listening position and has the minimum total distance from the viewpoint-side triangular grids in the current frame corresponding to the object-side triangular grids that have a common edge in step S108 as the triangular grid for three-point interpolation. In this case, the triangular grid that satisfies both the object-side selection condition and the viewpoint-side selection condition is selected. When the triangular grid for three-point interpolation is selected in this way, the process then proceeds to step S110.
[0543] When it is determined in step S104 that the listening position is within the triangular grid, the process of step S107 is executed, or the process of step S109 is executed, and then the process of step S110 is executed.
[0544] In step S110, the object position calculation unit 48 performs three-point interpolation based on the object absolute coordinate position information and the gain information of the triangular mesh selected for three-point interpolation (i.e., the three selected reference viewpoints), and generates the final object absolute coordinate position information and the gain information. The object position calculation unit 48 provides the obtained final object absolute coordinate position information and the gain information to the polar coordinate transformation unit 49.
[0545] In step S111, the object position calculation unit 48 determines whether there is a next frame to be processed, that is, whether the reproduction of the content has ended.
[0546] When it is determined in step S111 that there is a next frame, since the reproduction of the content has not ended, the process returns to step S101, and the above processing is repeated.
[0547] On the other hand, when it is determined in step S111 that there is no next frame, the reproduction of the content has ended, and the viewpoint selection process also ends.
[0548] As described above, the client 12 selects an appropriate triangular mesh based on the viewpoint side and the object side selection conditions, and performs three-point interpolation. In this way, the occurrence of discontinuous movement of the object position can be suppressed, and higher-quality acoustic reproduction can be achieved.
[0549] According to the present technology described above, during the movement of the listener in the free viewpoint space, reproduction can be achieved at each reference viewpoint according to the intention of the content creator, rather than using the reproduction based on the physical position relationship with respect to the conventional fixed object arrangement.
[0550] In addition, at any listening position sandwiched between multiple reference viewpoints, by performing interpolation processing based on the object arrangement of multiple reference viewpoints, the object position and gain suitable for any listening position can be generated. Therefore, the listener can move seamlessly between the reference viewpoints.
[0551] In addition, when the reference viewpoint overlaps with the object position, by reducing the signal level of the object or muting it, the listener can be given the feeling of as if the listener has become the object. Therefore, for example, a karaoke mode, a negative one performance mode, etc. can be realized, and the feeling of the listener's own participation in the content can be obtained.
[0552] In addition, in the interpolation processing of the reference viewpoint, when there is a reference viewpoint that the listener wants to get closer to, the feeling of movement is weighted by applying the bias coefficient α, so that even when the listener moves, the content can still be reproduced using the object arrangement closer to the listener's preferred viewpoint.
[0553] In addition, in the case where there are four or more reference viewpoints, a triangular mesh can be configured with three reference viewpoints, and three-point interpolation can be performed. In this case, since multiple triangular meshes can be configured, even when the listener freely moves in the area including the triangular meshes (i.e., the area surrounded by all the reference viewpoints), content reproduction can be achieved at an appropriate object position of an arbitrary position in the area as the listening position.
[0554] In addition, according to the present technology, in the case of using transmission in the polar coordinate system, audio reproduction of a free viewpoint space reflecting the intention of the content creator can be achieved only by adding system configuration information to a conventional MPEG-H encoding system.
[0555] <Configuration example of a computer>
[0556] Incidentally, the above-described series of processes can be executed by hardware and can also be executed by software. In the case where the series of processes are executed by software, a program constituting the software is installed in a computer. Here, the computer includes a computer installed in dedicated hardware, for example, a general-purpose personal computer or the like that can execute various functions by installing various programs.
[0557] Figure 20 FIG. is a block diagram showing a configuration example of the hardware of a computer that executes the above-described series of processes by a program.
[0558] In the computer, a central processing unit (CPU) 501, a read-only memory (ROM) 502, and a random access memory (RAM) 503 are interconnected via a bus 504.
[0559] An input / output interface 505 is further connected to the bus 504. An input unit 506, an output unit 507, a recording unit 508, a communication unit 509, and a drive 510 are connected to the input / output interface 505.
[0560] The input unit 506 includes a keyboard, a mouse, a microphone, an image sensor, etc. The output unit 507 includes a display, a speaker, etc. The recording unit 508 includes a hard disk, a non-volatile memory, etc. The communication unit 509 includes a network interface, etc. The drive 510 drives a removable recording medium 511 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0561] In the computer configured as described above, for example, the above-described series of processes are executed such that the CPU 501 loads a program stored in the recording unit 508 into the RAM 503 via the input / output interface 505 and the bus 504 and executes the program.
[0562] A program executed by a computer (CPU 501) can be provided by being recorded on a removable recording medium 511 such as an encapsulated medium, etc. In addition, the program can be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.
[0563] In the computer, when the removable recording medium 511 is installed on the drive 510, the program can be installed on the recording unit 508 via the input / output interface 505. In addition, the program can be received by the communication unit 509 via a wired or wireless transmission medium and installed on the recording unit 508. In addition, the program can be pre-installed on the ROM 502 or the recording unit 508.
[0564] Note that the program executed by the computer can be a program processed in chronological order according to the order described in this specification, or can be a program processed in parallel or at a required time (e.g., when a call is executed).
[0565] In addition, the embodiments of the present technology are not limited to the above embodiments, but various changes can be made without departing from the gist of the present technology.
[0566] For example, the present technology can adopt a cloud computing configuration in which a function is shared and jointly processed by multiple devices via a network.
[0567] In addition, each step described in the above flowchart can be executed by a single device or shared and executed by multiple devices.
[0568] In addition, in the case where a single step includes multiple processes, the multiple processes included in a single step can be executed by a single device or can be shared and executed by multiple devices.
[0569] In addition, the present technology can be configured as follows. (1)
[0571] An information processing device, comprising:
[0572] A listener position information acquisition unit that acquires listener position information of the viewpoint of the listener;
[0573] A reference viewpoint information acquisition unit that acquires the position information of the first reference viewpoint and the object position information of the object at the first reference viewpoint, and the position information of the second reference viewpoint and the object position information of the object at the second reference viewpoint; and
[0574] An object position calculation unit that calculates the position information of the object at the viewpoint of the listener based on the listener position information, the position information of the first reference viewpoint and the object position information of the object at the first reference viewpoint, and the position information of the second reference viewpoint and the object position information of the object at the second reference viewpoint. (2)
[0576] The information processing apparatus according to (1), wherein
[0577] the first reference viewpoint and the second reference viewpoint are viewpoints preset by the content creator. (3)
[0579] The information processing apparatus according to (1) or (2), wherein
[0580] the first reference viewpoint and the second reference viewpoint are viewpoints selected based on listener position information. (4)
[0582] The information processing apparatus according to any one of (1) to (3), wherein
[0583] the object position information is information indicating a position represented by polar coordinates or absolute coordinates; and
[0584] the reference viewpoint information acquisition unit acquires gain information of the object at the first reference viewpoint and gain information of the object at the second reference viewpoint. (5)
[0586] The information processing apparatus according to (4), wherein
[0587] the object position calculation unit calculates position information of the object at the listener's viewpoint by interpolation processing based on listener position information, position information of the first reference viewpoint and object position information at the first reference viewpoint, and position information of the second reference viewpoint and object position information at the second reference viewpoint. (6)
[0589] The information processing apparatus according to (4) or (5), wherein
[0590] the object position calculation unit calculates gain information of the object at the listener's viewpoint by interpolation processing based on listener position information, position information of the first reference viewpoint and gain information at the first reference viewpoint, and position information of the second reference viewpoint and gain information at the second reference viewpoint. (7)
[0592] The information processing apparatus according to (5) or (6), wherein
[0593] the object position calculation unit calculates position information or gain information of the object at the listener's viewpoint by performing interpolation processing with weighting on the object position information or gain information at the first reference viewpoint. (8)
[0595] The information processing apparatus according to any one of (1) to (4), wherein
[0596] The reference viewpoint information acquisition unit acquires the position information of the reference viewpoints and the object position information at the reference viewpoints for three or more multiple reference viewpoints including the first reference viewpoint and the second reference viewpoint; and
[0597] The object position calculation unit calculates the position information of the object at the listener's viewpoint through interpolation processing based on the listener position information, the position information of each of three reference viewpoints among the multiple reference viewpoints, and the object position information of each of the three reference viewpoints. (9)
[0599] The information processing apparatus according to (8), wherein
[0600] The object position calculation unit calculates the gain information of the object at the listener's viewpoint through interpolation processing based on the listener position information, the position information of each of the three reference viewpoints, and the gain information of each of the three reference viewpoints. (10)
[0602] The information processing apparatus according to (9), wherein
[0603] The object position calculation unit performs interpolation processing by weighting the object position information or gain information at a predetermined reference viewpoint among the three reference viewpoints to calculate the position information or gain information of the object at the listener's viewpoint. (11)
[0605] The information processing apparatus according to any one of (8) to (10), wherein
[0606] The object position calculation unit sets the region formed by any three reference viewpoints as a triangular grid, and selects three reference viewpoints that form a triangular grid satisfying a predetermined condition from among the multiple triangular grids as the three reference viewpoints to be used for interpolation processing. (12)
[0608] The information processing apparatus according to (11), wherein
[0609] In the case where the listener's viewpoint moves, the object position calculation unit
[0610] sets the region formed by each position of the object indicated by each object position information at the three reference viewpoints forming the triangular grid as an object triangular grid; and
[0611] selects three reference viewpoints for interpolation processing at the listener's viewpoint after movement based on the relationship between the object triangular grid corresponding to the triangular grid formed by the three reference viewpoints used for interpolation processing at the listener's viewpoint before movement and the object triangular grid corresponding to the triangular grid including the listener's viewpoint after movement. (13)
[0613] The information processing apparatus according to (12), wherein,
[0614] The object position calculation unit uses three reference viewpoints of a triangular mesh including a viewpoint after the movement of the listener that forms a triangular mesh corresponding to the object triangular mesh for interpolation processing of the viewpoint after the movement of the listener, the object triangular mesh having a common side with the object triangular mesh formed by the three reference viewpoints for interpolation processing of the viewpoint before the movement of the listener. (14)
[0616] The information processing apparatus according to any one of (1) to (13), wherein,
[0617] The object position calculation unit calculates the position information of the object at the viewpoint of the listener based on the listener position information, the position information of the first reference viewpoint, the object position information of the first reference viewpoint, the listener direction information indicating the direction of the face of the listener set at the first reference viewpoint, the position information of the second reference viewpoint, the object position information of the second reference viewpoint, and the listener direction information of the second reference viewpoint. (15)
[0619] The information processing apparatus according to (14), wherein,
[0620] The reference viewpoint information acquisition unit acquires configuration information including the position information and the listener direction information of each of a plurality of reference viewpoints, the plurality of reference viewpoints including a first reference viewpoint and a second reference viewpoint. (16)
[0622] The information processing apparatus according to (15), wherein,
[0623] The configuration information includes information indicating the number of the plurality of reference viewpoints and information indicating the number of objects. (17)
[0625] An information processing method, comprising: by an information processing apparatus,
[0626] acquiring listener position information of the viewpoint of the listener;
[0627] acquiring the position information of the first reference viewpoint and the object position information of the object at the first reference viewpoint, and the position information of the second reference viewpoint and the object position information of the object at the second reference viewpoint; and
[0628] Calculate the position information of an object at the listener's viewpoint based on the listener's position information, the position information of a first reference viewpoint and the object position information at the first reference viewpoint, and the position information of a second reference viewpoint and the object position information at the second reference viewpoint. (18)
[0630] A program that causes a computer to execute a process including the following steps:
[0631] Obtain the listener position information of the listener's viewpoint;
[0632] Obtain the position information of a first reference viewpoint and the object position information of the object at the first reference viewpoint, and the position information of a second reference viewpoint and the object position information of the object at the second reference viewpoint; and
[0633] Calculate the position information of the object at the listener's viewpoint based on the listener position information, the position information of the first reference viewpoint and the object position information at the first reference viewpoint, and the position information of the second reference viewpoint and the object position information at the second reference viewpoint.
[0634] Reference mark list
[0635] 11 Server
[0636] 12 Client
[0637] 21 Configuration information sending unit
[0638] 22 Encoded data sending unit
[0639] 41 Listener position information acquisition unit
[0640] 42 Viewpoint selection unit
[0641] 44 Encoded data acquisition unit
[0642] 46 Coordinate transformation unit
[0643] 47 Coordinate axis transformation processing unit
[0644] 48 Object position calculation unit
[0645] 49 Polar coordinate transformation unit
[0646] 111 Communication unit
[0647] 112 Position calculation unit
[0648] 113 Rendering processing unit.
Claims
1. An information processing apparatus, comprising: a listener position information acquisition unit that acquires listener position information of a viewpoint of a listener; a reference viewpoint information acquisition unit that acquires position information of a first reference viewpoint and object position information of an object at the first reference viewpoint, and position information of a second reference viewpoint and object position information of the object at the second reference viewpoint; and an object position calculation unit that calculates position information of the object at the viewpoint of the listener based on the listener position information, the position information of the first reference viewpoint and the object position information at the first reference viewpoint, and the position information of the second reference viewpoint and the object position information at the second reference viewpoint, wherein the object position calculation unit sets a region formed by any three of the reference viewpoints as a triangular mesh, and selects three of the reference viewpoints that form a triangular mesh satisfying a predetermined condition from among the plurality of triangular meshes as the three reference viewpoints to be used for interpolation processing, and wherein, when the viewpoint of the listener moves, the object position calculation unit, sets a region formed by each of the positions of the object indicated by each of the object position information at the three reference viewpoints forming the triangular mesh as an object triangular mesh; and selects three reference viewpoints for interpolation processing at the viewpoint after the movement of the listener based on a relationship between the object triangular mesh corresponding to the triangular mesh formed by the three reference viewpoints for interpolation processing at the viewpoint before the movement of the listener and the object triangular mesh corresponding to the triangular mesh including the viewpoint after the movement of the listener.
2. The information processing apparatus according to claim 1, wherein the first reference viewpoint and the second reference viewpoint are viewpoints preset by a content creator.
3. The information processing apparatus according to claim 1, wherein the first reference viewpoint and the second reference viewpoint are viewpoints selected based on the listener position information.
4. The information processing apparatus according to claim 1, wherein the object position information is information indicating a position represented by polar coordinates or absolute coordinates; and the reference viewpoint information acquisition unit acquires gain information of the object at the first reference viewpoint and gain information of the object at the second reference viewpoint.
5. The information processing apparatus according to claim 4, wherein the object position calculation unit calculates position information of the object at the viewpoint of the listener by interpolation processing based on the listener position information, the position information of the first reference viewpoint and the object position information at the first reference viewpoint, and the position information of the second reference viewpoint and the object position information at the second reference viewpoint.
6. The information processing apparatus according to claim 4, wherein The object position calculation unit calculates the gain information of the object at the viewpoint of the listener through interpolation processing based on the listener position information, the position information of the first reference viewpoint and the gain information at the first reference viewpoint, and the position information of the second reference viewpoint and the gain information at the second reference viewpoint.
7. The information processing apparatus according to claim 5, wherein the object position calculation unit calculates the position information or gain information of the object at the viewpoint of the listener by performing interpolation processing with weighting on the object position information or gain information at the first reference viewpoint.
8. The information processing apparatus according to claim 1, wherein the reference viewpoint information acquisition unit acquires the position information of the reference viewpoint and the object position information at the reference viewpoint for three or more reference viewpoints including the first reference viewpoint and the second reference viewpoint; and the object position calculation unit calculates the position information of the object at the viewpoint of the listener through interpolation processing based on the listener position information, the position information of each of three of the three or more reference viewpoints, and the object position information at each of the three reference viewpoints.
9. The information processing apparatus according to claim 8, wherein the object position calculation unit calculates the gain information of the object at the viewpoint of the listener through interpolation processing based on the listener position information, the position information of each of the three reference viewpoints, and the gain information of each of the three reference viewpoints.
10. The information processing apparatus according to claim 9, wherein the object position calculation unit calculates the position information or gain information of the object at the viewpoint of the listener by performing interpolation processing with weighting on the object position information or the gain information at a predetermined reference viewpoint among the three reference viewpoints.
11. The information processing apparatus according to claim 1, wherein the object position calculation unit uses three reference viewpoints forming a triangular mesh corresponding to the object triangular mesh and including the viewpoints after the movement of the listener for interpolation processing of the viewpoints after the movement of the listener, and the object triangular mesh has a common side with the object triangular mesh formed by the three reference viewpoints used for interpolation processing of the viewpoints before the movement of the listener.
12. The information processing apparatus according to claim 1, wherein the object position calculation unit calculates the position information of the object at the viewpoint of the listener based on the listener position information, the position information of the first reference viewpoint, the object position information at the first reference viewpoint, listener direction information indicating the direction of the face of the listener set at the first reference viewpoint, the position information of the second reference viewpoint, the object position information at the second reference viewpoint, and the listener direction information at the second reference viewpoint.
13. The information processing apparatus according to claim 12, wherein the reference viewpoint information acquisition unit acquires configuration information including the position information and the listener direction information of each of a plurality of reference viewpoints, the plurality of reference viewpoints including the first reference viewpoint and the second reference viewpoint.
14. The information processing apparatus according to claim 13, wherein the configuration information includes information indicating the number of the plurality of reference viewpoints and information indicating the number of the objects.
15. An information processing method, comprising: An information processing apparatus acquires listener position information of a listener's viewpoint; acquires position information of a first reference viewpoint and object position information of the object at the first reference viewpoint, and position information of a second reference viewpoint and object position information of the object at the second reference viewpoint; and and calculates position information of the object at the listener's viewpoint based on the listener position information, the position information of the first reference viewpoint and the object position information at the first reference viewpoint, and the position information of the second reference viewpoint and the object position information at the second reference viewpoint, wherein a region formed by any three of the reference viewpoints is set as a triangular mesh, and three of the reference viewpoints that form a triangular mesh satisfying a predetermined condition are selected from among the plurality of triangular meshes as three reference viewpoints to be used for interpolation processing, and wherein, when the listener's viewpoint moves, a region formed by each of the positions of the object indicated by each of the object position information at the three reference viewpoints forming the triangular mesh is set as an object triangular mesh; and three reference viewpoints for interpolation processing at the listener's viewpoint after the movement are selected based on a relationship between the object triangular mesh corresponding to the triangular mesh formed by the three reference viewpoints for interpolation processing at the listener's viewpoint before the movement and the object triangular mesh corresponding to the triangular mesh included in the listener's viewpoint after the movement.
16. A computer-readable storage medium storing a program which, when executed, causes a computer to execute a process including the following steps: acquire listener position information of a listener's viewpoint; acquire position information of a first reference viewpoint and object position information of the object at the first reference viewpoint, and position information of a second reference viewpoint and object position information of the object at the second reference viewpoint; and calculate position information of the object at the listener's viewpoint based on the listener position information, the position information of the first reference viewpoint and the object position information at the first reference viewpoint, and the position information of the second reference viewpoint and the object position information at the second reference viewpoint, Among them, set a region formed by any three of the reference viewpoints as a triangular mesh, and select three of the reference viewpoints that form a triangular mesh satisfying a predetermined condition from among the plurality of triangular meshes as three reference viewpoints to be used for interpolation processing, and wherein, in a case where the viewpoint of the listener moves, a region formed by each of the positions of the object indicated by each of the object position information at the three reference viewpoints forming the triangular grid is set as an object triangular grid; and based on a relationship between the object triangular grid corresponding to the triangular grid formed by the three reference viewpoints for interpolation processing of the viewpoint before the movement of the listener and the object triangular grid corresponding to the triangular grid including the viewpoint after the movement of the listener, three reference viewpoints for interpolation processing of the viewpoint after the movement of the listener are selected.
Citation Information
Patent Citations
Information processing device, method, and program
WO2019198540A1
Reproducing device, reproducing method, information processing device, information processing method, and program
CN109983786A