Acoustic metadata coordinate system transformation device and program

The acoustic metadata coordinate system conversion device addresses coordinate deviations by defining rectangular regions and line segment domains, enabling accurate conversions between polar and Cartesian coordinates for diverse speaker arrangements.

JP7866876B2Active Publication Date: 2026-05-28NIPPON HOSO KYOKAI
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NIPPON HOSO KYOKAI
Filing Date
2022-06-02
Publication Date
2026-05-28

AI Technical Summary

Technical Problem

Existing audio object coordinate system conversion methods fail to accurately convert between polar and Cartesian coordinates for speaker arrangements other than 5.1.4, leading to deviations in elevation and Z-axis directions, especially for 22.2ch and 7.1.4 configurations.

Method used

An acoustic metadata coordinate system conversion device that defines rectangular regions and line segment domains based on speaker positions, correcting coordinate values using endpoints and intersections to align with standard values, allowing conversions for any speaker arrangement.

Benefits of technology

Enables accurate conversion between polar and Cartesian coordinates, correcting deviations in elevation and Z-axis directions, ensuring alignment with standard values for various speaker configurations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007866876000028
    Figure 0007866876000028
  • Figure 0007866876000029
    Figure 0007866876000029
  • Figure 0007866876000030
    Figure 0007866876000030
Patent Text Reader

Abstract

To convert a polar coordinate value to an orthogonal coordinate value corresponding to a standard value or convert an orthogonal coordinate value to a polar coordinate value corresponding to a standard value, while correcting deviation of a coordinate value of a sound object.SOLUTION: An acoustic metadata coordinate system conversion apparatus 1 includes: a rectangular region specifying unit 11 which specifies a plurality of rectangular regions on a wall surface of a listening room, on the basis of a reference point of the listening room and a correspondence relation between a polar coordinate value and an orthogonal coordinate value in positions of speakers to be arranged; a line segment region specifying unit 12 which derives a Z-coordinate value z of a sound object, and corrects, when intersections of sides of rectangular regions and a horizontal plane in a height of the z are set as end points of the line segment regions and the z is not coincident with the Z-coordinate value in an intermediate layer, coordinate values of end points of line segment regions in consideration of misalignment in azimuth of a polar coordinate value corresponding to an orthogonal coordinate value; and a coordinate derivation unit 13 which searches for a line segment region that the sound object belongs to, and derives an orthogonal coordinate value of the sound object using the end points of the line segment regions.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an acoustic metadata coordinate system transformation device and program for transforming the positional information of audio objects assigned to acoustic metadata. [Background technology]

[0002] In object-based audio, the 3D positional information of audio objects is attached to the audio signal as acoustic metadata and transmitted, allowing the renderer to distribute the audio signal to each speaker in a manner adapted to the speaker placement. The positional information is defined by ITU-R and can be given in either a polar coordinate system (position defined by three variables: azimuth, elevation, and distance) or a Cartesian coordinate system (position defined by three variables: X, Y, and Z coordinate values). However, the playback method adopted by the renderer limits the interpretable coordinate system to one of these two, thus requiring a coordinate system conversion. Furthermore, some object-based audio renderers define the position of an audio object described in a Cartesian coordinate system as its relative positional relationship (ratio of distances) to speakers placed at the four corners of the listening room. Therefore, geometric conversions between polar and Cartesian coordinates in 3D Euclidean space cannot be applied to the edges, and it becomes necessary to apply a conversion rule based on the speaker placement.

[0003] Conventionally, coordinate system transformations in acoustic metadata have been standardized by the ITU-R, and Non-Patent Document 1 specifies a method for transforming coordinate systems based on the correspondence between polar coordinate values ​​and Cartesian coordinate values ​​in speaker arrangements described in 5.1.4. The speaker arrangement for this channel format is described, for example, in Non-Patent Document 2. [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] ITU-R BS.2127-0, “Audio Definition Model renderer for advanced sound systems”, 2019 [Non-Patent Document 2] ITU-R BS.2051-1, “Advanced sound system for programme production”, 2017 [Summary of the Invention] [Problems to be Solved by the Invention]

[0005] However, when converting the polar coordinate values of audio objects arranged at speaker positions such as 22.2ch and 7.1.4 to rectangular coordinate values using the technology defined in Non-Patent Document 1, which is a 5.1.4-based conversion, there is a problem that the values are different from the rectangular coordinate values of the speakers defined in Non-Patent Document 1 for 22.2ch and 7.1.4. Similarly, when converting the rectangular coordinate values of audio objects arranged at speaker positions such as 22.2ch and 7.1.4 to polar coordinate values using the technology defined in Non-Patent Document 1, which is a 5.1.4-based conversion, there is a problem that the values are different from the polar coordinate values of the speakers defined in Non-Patent Document 1 for 22.2ch and 7.1.4.

[0006] In addition, when the azimuth angles of the speakers arranged at the four corners of the middle, upper, and lower layers of a listening room such as 22.2ch or 7.1.4 are different, there is a problem that the deviation of the coordinate values of the audio object in the elevation direction or the Z-axis direction becomes large.

[0007] In view of such circumstances, an object of the present invention is to provide an acoustic metadata coordinate system conversion device and a program that can convert from polar coordinate values to rectangular coordinate values that match the standard values, or from rectangular coordinate values to polar coordinate values that match the standard values, and can correct the deviation of the coordinate values of the audio object in the elevation direction or the Z-axis direction even when rendered to speaker arrangements other than 5.1.4. [Means for Solving the Problems]

[0008] In order to solve the above problems, an acoustic metadata coordinate system conversion device according to the present invention is an acoustic metadata coordinate system conversion device that converts the position information of a voice object attached to acoustic metadata from polar coordinate values to orthogonal coordinate values. Based on the correspondence between polar coordinate values and orthogonal coordinate values at the reference point of the listening room and the positions of each speaker to be arranged, a rectangular region defining unit that defines a plurality of rectangular regions having as endpoints positions where the correspondence between coordinate systems is known in the upper, middle, and lower layers with respect to the wall surface of the listening room; a line segment region defining unit that derives the Z coordinate value z of the voice object, uses the intersections of the horizontal plane at the height of the z and the sides of the rectangular region as the endpoints of the line segment region, and corrects the coordinate values of the endpoints of the line segment region in consideration of the deviation of the azimuth angle of the polar coordinate value corresponding to the orthogonal coordinate value when the z does not match the Z coordinate value z = 0 of the middle layer; and a coordinate derivation unit that searches for the line segment region to which the voice object belongs and derives the orthogonal coordinate value of the voice object using the endpoints of the line segment region.

[0009] Also, in order to solve the above problems, an acoustic metadata coordinate system conversion device according to the present invention is an acoustic metadata coordinate system conversion device that converts the position information of a voice object attached to acoustic metadata from orthogonal coordinate values to polar coordinate values. Based on the correspondence between polar coordinate values and orthogonal coordinate values at the reference point of the listening room and the positions of each speaker to be arranged, a rectangular region defining unit that defines a plurality of rectangular regions having as endpoints positions where the correspondence between coordinate systems is known in the upper, middle, and lower layers with respect to the wall surface of the listening room; a line segment region defining unit that uses the intersections of the horizontal plane at the height z of the voice object and the sides of the rectangular region as the endpoints of the line segment region, and corrects the coordinate values of the endpoints of the line segment region in consideration of the deviation of the azimuth angle of the polar coordinate value corresponding to the orthogonal coordinate value when the z does not match the Z coordinate value z = 0 of the middle layer; and a coordinate derivation unit that searches for the line segment region to which the voice object belongs and derives the polar coordinate value of the voice object using the endpoints of the line segment region.

[0010] Furthermore, in the acoustic metadata coordinate system transformation device according to the present invention, the reference points of the listening room may be a total of 12 points at the four corners of each of the upper, middle, and lower levels of the listening room, or a total of 24 points including the four corners of each of the upper, middle, and lower levels of the listening room, as well as the front, rear, and sides of the listener.

[0011] Furthermore, in order to solve the above problems, the program according to the present invention causes a computer to function as the acoustic metadata coordinate system transformation device. [Effects of the Invention]

[0012] According to the present invention, for any speaker arrangement, it is possible to uniquely convert between polar coordinates and Cartesian coordinates while correcting for the correspondence between the polar coordinates and Cartesian coordinates of each speaker, as well as the shift in coordinate values ​​in the elevation direction or the Z-axis direction. [Brief explanation of the drawing]

[0013] [Figure 1] This is a block diagram showing an example configuration of an acoustic metadata coordinate system transformation device according to the first embodiment. [Figure 2] This flowchart shows the procedure for converting polar coordinate values ​​to Cartesian coordinate values ​​in the acoustic metadata coordinate system transformation device according to the first embodiment. [Figure 3] This figure shows an example of a reference point for defining the listening room and rectangular area in a Cartesian coordinate system in the acoustic metadata coordinate system transformation device according to the first embodiment. [Figure 4A] This figure shows images of multiple rectangular regions defined in the unfolded view of the listening room wall surface and line segment regions defined according to the height of the sound object, in the acoustic metadata coordinate system transformation device according to the first embodiment. [Figure 4B] This figure shows the listening room corresponding to Figure 4A. [Figure 5]This figure shows the results of performing a coordinate system transformation on the azimuth angle of an audio object in 1-degree increments from -180 degrees to 180 degrees, when the speaker arrangement is 22.2ch, in the acoustic metadata coordinate system transformation device according to the first embodiment. [Figure 6] This figure shows the results of coordinate system transformation performed on an audio object moving linearly from position (30,-30,1) through speaker M+030 to position (30,-30,1) in a polar coordinate system, and on an audio object moving linearly from position (-30,-30,1) through speaker M-030 to position (-30,-30,1), in the acoustic metadata coordinate system transformation device according to the first embodiment, when the speaker arrangement is 22.2ch. [Figure 7] This figure shows the results of performing a coordinate system transformation on the azimuth angle of an audio object in 1-degree increments from -180 degrees to 180 degrees, when the speaker arrangement is set to 7.1.4, in the acoustic metadata coordinate system transformation device according to the first embodiment. [Figure 8] This figure shows the results of coordinate system transformation performed on an audio object moving linearly from position (30,-30,1) through speaker M+030 to position (30,-30,1) in a polar coordinate system, and on an audio object moving linearly from position (-30,-30,1) through speaker M-030 to position (-30,-30,1), in the acoustic metadata coordinate system transformation device according to the first embodiment, when the speaker arrangement is 7.1.4. [Figure 9] This figure shows the results of performing a coordinate system transformation on the azimuth angle of an audio object in 1-degree increments from -180 degrees to 180 degrees, when the speaker arrangement is as described in 5.1.4, in the acoustic metadata coordinate system transformation device according to the first embodiment. [Figure 10]This figure shows the results of coordinate system transformation performed on an audio object moving linearly from position (30,-30,1) through speaker M+030 to position (30,-30,1) in a polar coordinate system, and on an audio object moving linearly from position (-30,-30,1) through speaker M-030 to position (-30,-30,1), in the acoustic metadata coordinate system transformation device according to the first embodiment, when the speaker arrangement is 5.1.4. [Figure 11] This figure shows the results of performing a coordinate system transformation on the azimuth angle of an audio object in 1-degree increments from -180 degrees to 180 degrees, when the speaker arrangement is 22.2ch and the virtual speakers are set to mid-level ±110 degrees and mid-level ±150 degrees, in the acoustic metadata coordinate system transformation device according to the first embodiment. [Figure 12] This is a block diagram showing an example configuration of an acoustic metadata coordinate system transformation device according to the second embodiment. [Figure 13] This flowchart shows the procedure for converting Cartesian coordinate values ​​to polar coordinate values ​​in an acoustic metadata coordinate system transformation device according to the second embodiment. [Figure 14] This figure shows the results of performing a coordinate system transformation on the azimuth angle of an audio object in 1-degree increments from -180 degrees to 180 degrees, when the speaker arrangement is 22.2ch, in the acoustic metadata coordinate system transformation device according to the second embodiment. [Figure 15] This figure shows the results of coordinate system transformation performed on an audio object moving linearly from position (-1,+1,-1) to position (-1,+1,+1) and an audio object moving linearly from position (+1,+1,-1) to position (+1,+1,+1) in a Cartesian coordinate system, when the speaker arrangement is 22.2ch, in the acoustic metadata coordinate system transformation device according to the second embodiment. [Figure 16] This figure shows the results of performing coordinate system transformations in the Cartesian coordinate system of the audio object in 1-degree increments from -180 degrees to 180 degrees, when the speaker arrangement is set to 7.1.4, in the acoustic metadata coordinate system transformation device according to the second embodiment. [Figure 17]This figure shows the results of performing a coordinate system transformation on an audio object moving linearly from position (-1,+1,-1) to position (-1,+1,+1) and an audio object moving linearly from position (+1,+1,-1) to position (+1,+1,+1) in a Cartesian coordinate system, when the speaker arrangement is set to 7.1.4, in the acoustic metadata coordinate system transformation device according to the second embodiment. [Figure 18] This figure shows the results of performing coordinate system transformations in the Cartesian coordinate system of the audio object in 1-degree increments from -180 degrees to 180 degrees, when the speaker arrangement is as described in 5.1.4, in the acoustic metadata coordinate system transformation device according to the second embodiment. [Figure 19] This figure shows the results of a coordinate system transformation performed on an audio object moving linearly from position (-1,+1,-1) to position (-1,+1,+1) and an audio object moving linearly from position (+1,+1,-1) to position (+1,+1,+1) in a Cartesian coordinate system, when the speaker arrangement is as described in 5.1.4, in the acoustic metadata coordinate system transformation device according to the second embodiment. [Figure 20] This figure shows the result of converting polar coordinate values ​​to Cartesian coordinate values ​​at positions in the polar coordinate system of an audio object in the second embodiment, where the speaker arrangement is 22.2ch / 7.1.4 / 5.1.4, and the elevation angle is converted from polar coordinate values ​​to Cartesian coordinate values ​​at positions in 5-degree increments from -180 degrees to 180 degrees in azimuth and in 15-degree increments from -90 degrees to 90 degrees in elevation, and then converting those Cartesian coordinate values ​​back to polar coordinate values ​​according to the present invention. [Figure 21] This figure shows the results of performing coordinate system transformations in 1-degree increments from -180 degrees to 180 degrees on the azimuth angle in the Cartesian coordinate system of the sound object, when the speaker arrangement is 22.2ch and the corresponding azimuth angles of the left front (x,y)=(-1,1) / right front (x,y)=(1,1) of the listening room are ±30 degrees regardless of the height z, in the acoustic metadata coordinate system transformation device according to the second embodiment. [Figure 22]This figure shows the results of coordinate system transformation performed on an audio object moving linearly from position (-1,+1,-1) to position (-1,+1,+1) and an audio object moving linearly from position (+1,+1,-1) to position (+1,+1,+1) in a Cartesian coordinate system, in the acoustic metadata coordinate system transformation device according to the second embodiment, where the speaker arrangement is 22.2ch and the corresponding azimuth angles of the left front (x,y)=(-1,1) and right front (x,y)=(1,1) of the listening room are ±30 degrees regardless of height z. [Modes for carrying out the invention]

[0014] Embodiments will be described in detail below with reference to the drawings. In this specification, Cartesian coordinate values ​​are represented by a set of three variables (x, y, z) consisting of the X, Y, and Z coordinates, and polar coordinate values ​​are represented by a set of three variables (φ, θ, d) consisting of the azimuth angle, elevation angle, and distance.

[0015] <First Embodiment> In the first embodiment, an acoustic metadata coordinate system transformation device is described that converts polar coordinate values ​​(φ,θ,d) in acoustic metadata attached to an audio signal to Cartesian coordinate values ​​(x,y,z).

[0016] Figure 1 is a block diagram showing an example configuration of the acoustic metadata coordinate system transformation device 1 according to the first embodiment. The acoustic metadata coordinate system transformation device 1 shown in Figure 1 comprises a rectangular area definition unit 11, a line segment area definition unit 12, and a coordinate derivation unit 13. The rectangular area definition unit 11 receives acoustic metadata in polar coordinates, listening room information, and speaker placement information as input. Here, "listening room information" refers to information indicating the polar coordinates and Cartesian coordinate values ​​of the listening room reference point. Also, "speaker placement information" refers to information indicating the polar coordinates and Cartesian coordinate values ​​of the speaker positions.

[0017] Figure 2 is a flowchart showing the procedure for converting the position information of an audio object from polar coordinates to Cartesian coordinates in the acoustic metadata coordinate system transformation device 1. This algorithm can be broadly divided into three steps, from step S1 to step S3.

[0018] In step S1, the rectangular area definition unit 11 defines multiple rectangular areas on the wall surface of the listening room, with endpoints at positions where the correspondence between the coordinate systems is known for the upper, middle, and lower layers, based on the correspondence between polar coordinate values ​​and Cartesian coordinate values ​​at the reference point of the listening room and the positions of each speaker to be placed.

[0019] In step S2, the line segment domain definition unit 12 derives the Z coordinate value of the sound object based on the elevation angle and distance. Then, considering the difference in the elevation direction between the polar coordinate value and the Cartesian coordinate value that occurs when the azimuth angle of the line segment domain endpoint differs between the middle layer and the upper / lower layers, the coordinate value of the line segment domain endpoint is corrected. Here, "line segment domain" means the intersection line between the horizontal plane at height z where the sound object is located and the rectangular domain, and "line segment domain endpoint (endpoint of the line segment domain)" means the intersection point between the horizontal plane at height z and the side of the rectangular domain.

[0020] In step S3, the coordinate derivation unit 13 searches for the line segment region to which the audio object belongs and derives the orthogonal coordinate values ​​of the audio object using the endpoints of the line segment region.

[0021] (Explanation of the operation in step S1) In step S11, in order to define the correspondence between polar coordinate values ​​and Cartesian coordinate values, the polar coordinates and Cartesian coordinates of the reference point of the listening room and the positions of each speaker are input.

[0022] The reference points in the listening room may consist of 12 points in total, located at the four corners of each level (upper, middle, and lower) of the listening room. Alternatively, the reference points in the listening room may consist of 24 points in total, located at the four corners of each level (upper, middle, and lower) of the listening room, as well as in front of, behind, and to the sides of the listener.

[0023] FIG. 3 is a diagram showing an example of a listening room in a rectangular coordinate system and a point serving as a reference for coordinate conversion. The listening room reference points refer to the vertices of the upper, middle, and lower layers shown in FIG. 3, i.e., (±1, ±1, ±1), (±1, ±1, 0) (arbitrary combination), and the midpoints of each side (±1, 0, ±1), (0, ±1, ±1), (±1, ±1, 0), (0, ±1, 0) (arbitrary combination). When it is not necessary to strictly specify the coordinate values on the side of the listening point, the input for the midpoints of each side may be omitted, and the following processing still holds true.

[0024] In step S12, a plurality of rectangular regions are defined on the wall surface of the listening room.

[0025] FIG. 4A is a diagram showing an image of a plurality of rectangular regions defined in the development view of the wall surface of the listening room and a line segment region defined according to the height of the audio object. FIG. 4B is a diagram showing the listening room corresponding to FIG. 4A. In this figure, the black circles indicate the positions where speakers or reference points exist among the endpoints of the rectangular regions. The dashed white circles indicate the positions where speakers or reference points do not exist among the endpoints of the rectangular regions, and virtual azimuth angles calculated are applied. The shaded circles indicate the positions of the audio objects. The dashed squares indicate the endpoints of the line segment regions, and the azimuth angles corrected according to the height at which the audio objects are located are applied.

[0026] In step S12, the wall surface of the listening room is divided into a plurality of rectangular regions such that the listening room reference points or speaker positions are taken as the endpoints of the rectangular regions, and the endpoints in the upper layer are (x l , y l , +1), (x r , y r , +1), the endpoints in the middle layer are (x l , y l , 0), (x r , y r , 0), and the endpoints in the lower layer are (x l , y l , -1), (x r , y r , -1). However, x l , y lThis represents the X,Y coordinates at the leftmost endpoint of the rectangular area as viewed from the listening point, and x r ,y r This represents the X,Y coordinates at the rightmost endpoint of the rectangular region as viewed from the listening point.

[0027] For the endpoints of the upper, middle, and lower layers of the obtained rectangular region, where neither speakers nor reference points exist and the correspondence between the coordinate systems is not defined, a virtual polar coordinate value is calculated using an arbitrary rule that does not contradict the correspondence between the coordinate systems at the listening room reference point and speaker positions. For the elevation angle in the virtual polar coordinate value, the elevation angle value of either the upper, middle, or lower layer is adopted, and for example, the azimuth angles located in front of and behind the listening point may be calculated using equations (1) and (2). Here, φ and x represent the virtual azimuth angles to be calculated and their x-coordinates. l ,φ~ r ,x~ l ,x~ r This shows the azimuth angles and x-coordinates of the points to the left and right of the listening point, as viewed from the listening point, of the listening room reference point that encloses the virtually calculated azimuth angle. The azimuth angles of points located to the sides of the listening point can be similarly calculated by substituting x with y in equations (1) and (2).

[0028]

number

number

[0029] (Explanation of the operation in step S2) In step S21, the Z-coordinate value of the sound object is derived, and the line segment region is defined based on that Z-coordinate value. However, when deriving the Z-coordinate value, if the sound object is located above the upper layer or below the lower layer, it is transformed as a sound object on the ceiling or floor of the listening room, and if it is located between the upper and lower layers, it is transformed as a sound object on or inside the wall of the listening room.

[0030] In step S22, it is determined whether the audio object is located above the upper layer or below the lower layer. If the audio object is located above the upper layer or below the lower layer, the Z coordinate of the audio object is obtained as the Z coordinate of the upper or lower layer of the listening room.

[0031] If the audio object is located between the upper and lower layers, its Z-coordinate is calculated using an arbitrary transformation rule that is consistent with the elevation angles of the upper and lower layer speakers. The method for calculating the Z-coordinate of the audio object may be, for example, given by equation (3). However, the variable θ... T ,θ B ,θ' T ,θ' B θ' represents the elevation angles of the upper and lower layers in polar coordinates and Cartesian coordinates, respectively. T ,θ' B Considering the ratio of the horizontal and vertical distances between the listening point and the upper front position, we decided to use 45 degrees, θ T ,θ B For example, the elevation angle of the upper and lower speakers is 30 degrees as defined in Non-Patent Document 2. Also, instead of equation (3), |θ as defined in Non-Patent Document 1 is used. T |=|θ B You may also use a calculation rule for the Z coordinate that assumes |.

[0032]

number

[0033] In step S23, if the sound object is located on the wall surface of the listening room, the intersection of the horizontal plane and the rectangular region at the height z of the sound object is newly defined as a line segment region (see Figure 4). The rectangular region defines endpoints at the height of the upper layer (z=1), the height of the middle layer (z=0), and the height of the lower layer (z=-1), respectively, while the line segment region defines an endpoint at the height z where the sound object is located. If the Z coordinate value z of the sound object does not coincide with the Z coordinate value z=0 of the middle layer, and the azimuth angles of the rectangular region endpoints differ between the middle layer and the upper and lower layers, the azimuth angle φ of the polar coordinate value corresponding to the Cartesian coordinate value (x, y) of the line segment region endpoint is corrected, taking into account the difference in the azimuth angle of the polar coordinate value corresponding to the Cartesian coordinate value. For example, equation (4) may be used as the correction rule for the line segment region endpoints. However, φ(x,y,z) is the azimuth angle of the endpoint of the line segment region at coordinate (x,y,z), and φ(x,y,0), φ(x,y,1), and φ(x,y,-1) are the azimuth angles of the endpoints of the rectangular region in the middle, upper, and lower layers, respectively.

[0034]

number

[0035] (Explanation of the operation in step S3) Finally, in step S31, the position vector (x) of the endpoint of the line segment region to which the audio object belongs on the XY plane is obtained. l ,y l ,0), (x r ,y r The X and Y coordinate values ​​are calculated using (0), and the previously derived Z coordinate value is added to derive the Cartesian coordinate value. The search for the line segment region to which the sound object belongs is determined based on the relationship between the azimuth angle of the sound object and the azimuth angles of the endpoints of the line segment region. However, if the sound object is located on the ceiling or floor, the endpoints in the upper or lower layer of the rectangular region are considered as the endpoints of the line segment region, and the determination is made based on the relationship between the azimuth angle of those endpoints and the azimuth angle of the sound object. The derivation of the Cartesian coordinate value is, for example, based on the internal division ratio p∈[0,1] of the endpoints of the line segment region obtained by equation (6), and the quantity r that controls the magnitude of the position vector obtained by equation (7). xyUsing this, the position vector that divides any endpoint of a line segment region located on the same wall surface internally in the ratio p:(1-p) may be calculated by equation (5). Alternatively, instead of equation (6), the calculation rule for the internal division ratio using the tangent function specified in Non-Patent Literature 1 may be used, and instead of equation (7), |θ specified in Non-Patent Literature 1 may be used. T |=|θ B You may also use a calculation rule that assumes |.

[0036]

number

number

number

[0037] (Examples) To confirm the effectiveness of the acoustic metadata coordinate system transformation device 1, numerical experiments were conducted comparing coordinate system transformations performed using the conventional method and the present invention with standard values. "Conventional method" refers to the case where the coordinate system is transformed based on the correspondence between polar coordinate values ​​and Cartesian coordinate values ​​in 5.1.4 (Non-Patent Literature 1). "The present invention" refers to the case where the coordinate system is transformed using the acoustic metadata coordinate system transformation device 1. "Standard values" refer to the case where the coordinate system is transformed based on the correspondence between polar coordinate values ​​and Cartesian coordinate values ​​according to the standard at the speaker position (however, transformation is only possible at the position where the speaker is located).

[0038] (Example: 22.2ch) Assuming the correspondence between the coordinate systems of the listening room reference point is as shown in Table 1, and the speaker arrangement is 22.2ch, i.e., the correspondence between the coordinate systems at the speaker positions is as shown in Table 2, the listening room wall is divided into 10 rectangular regions as shown in Table 3. The speaker labels in the speaker arrangement information in Table 2 are described with reference to Non-Patent Literature 2, where U (Upper) represents the upper layer, M (Middle) represents the middle layer, B (Bottom) represents the lower layer, and T (Top) represents overhead. Table 3 shows only the azimuth angles of the endpoints of the rectangular regions, and the calculated virtual azimuth angles are enclosed in parentheses.

[0039] [Table 1] [Table 2] [Table 3]

[0040] Figure 5 shows the results of coordinate system transformations performed on the azimuth angles of audio objects in the upper, middle, and lower layers, in 1-degree increments from -180 degrees to 180 degrees, when the speaker configuration is 22.2ch. The graph on the left of this figure shows the relationship between the azimuth angle before transformation (horizontal axis) and the X-coordinate after transformation (vertical axis), and the graph on the right shows the relationship between the azimuth angle before transformation (horizontal axis) and the Y-coordinate after transformation (vertical axis). Referring to Figure 5, it can be confirmed that the transformation results using the conventional method do not partially match the transformation results based on standard values ​​in the upper and middle layers, whereas the transformation results according to the present invention completely match the transformation results based on standard values.

[0041] Figure 6 shows the results of coordinate system transformation for an audio object moving linearly from position (30,-30,1) through speaker M+030 to position (30,-30,1) in a polar coordinate system, and for an audio object moving linearly from position (-30,-30,1) through speaker M-030 to position (-30,-30,1), assuming a speaker configuration of 22.2ch. The graph in this figure shows the relationship between the X coordinate (horizontal axis) and Z coordinate (vertical axis) after transformation. Referring to Figure 6, in the transformation results using the conventional method, as the audio object moves from the middle layer to the upper and lower layers, it approaches the speaker with an azimuth angle of ±45 degrees. However, in the transformation results according to the present invention, the coordinate values ​​are transformed to approach a position between the speaker with an azimuth angle of ±45 degrees and the speaker with an azimuth angle of 0 degrees, and the effect of correcting the coordinate shift in the elevation direction can be confirmed.

[0042] (Example: 7.1.4) If the correspondence between the coordinate systems of the listening room reference point is the same as in the 22.2ch case as shown in Table 3, and the speaker arrangement is as in 7.1.4, i.e., the correspondence between the coordinate systems at the speaker positions is as shown in Table 4, then the listening room wall is divided into seven rectangular regions as shown in Table 5.

[0043] [Table 4] [Table 5]

[0044] Figure 7 shows the results of coordinate system transformations performed on the azimuth angles of audio objects in the upper, middle, and lower layers, in 1-degree increments from -180 degrees to 180 degrees, when the speaker arrangement is set to 7.1.4. The graph on the left of this figure shows the relationship between the azimuth angle before transformation (horizontal axis) and the X-coordinate after transformation (vertical axis), and the graph on the right shows the relationship between the azimuth angle before transformation (horizontal axis) and the Y-coordinate after transformation (vertical axis). Referring to Figure 7, it can be confirmed that the transformation results using the conventional method do not partially match the transformation results based on the standard values ​​in the upper and middle layers, whereas the transformation results according to the present invention completely match the transformation results based on the standard values.

[0045] Figure 8 shows the results of coordinate system transformation for an audio object moving linearly from position (30,-30,1) through speaker M+030 to position (30,-30,1) in a polar coordinate system, and for an audio object moving linearly from position (-30,-30,1) through speaker M-030 to position (-30,-30,1), assuming speaker placement 7.1.4. The graph in this figure shows the relationship between the X coordinate (horizontal axis) and Z coordinate (vertical axis) after transformation. Referring to Figure 8, in the transformation results of the conventional method, as the audio object moves from the middle layer to the upper and lower layers, it approaches the speaker with an azimuth angle of ±45 degrees. On the other hand, in the proposed method, the coordinate values ​​are transformed to approach a position between the speaker with an azimuth angle of ±45 degrees and the speaker with an azimuth angle of 0 degrees, and the effect of correcting the coordinate shift in the elevation direction can be confirmed.

[0046] (Example: 5.1.4) Assuming the correspondence between the coordinate systems of the listening room reference point is as shown in Table 6, and the speaker arrangement is as described in 5.1.4, i.e., the correspondence between the coordinate systems at the speaker positions is as shown in Table 7, the listening room wall is divided into five rectangular regions as shown in Table 8. In this example, the specification of polar coordinates for the midpoints of each side of the listening room is omitted.

[0047] [Table 6] [Table 7] [Table 8]

[0048] Figure 9 shows the results of coordinate system transformations performed on the azimuth angle of audio objects in the upper, middle, and lower layers, in 1-degree increments from -180 degrees to 180 degrees, when the speaker arrangement is set to 5.1.4. The graph on the left of this figure shows the relationship between the azimuth angle before transformation (horizontal axis) and the X-coordinate after transformation (vertical axis), and the graph on the right shows the relationship between the azimuth angle before transformation (horizontal axis) and the Y-coordinate after transformation (vertical axis). Referring to Figure 9, it can be confirmed that the transformation results of the conventional method and the present invention all match the transformation results based on the standard values ​​at the speaker position, and that there is no significant difference between the two transformation results at other positions as well.

[0049] Figure 10 shows the results of coordinate system transformation for an audio object moving linearly from position (30,-30,1) through speaker M+030 to speaker U+030 in a polar coordinate system, and for an audio object moving linearly from position (-30,-30,1) through speaker M-030 to speaker U-030, when the speaker arrangement is 5.1.4. The graph in this figure shows the relationship between the X coordinate (horizontal axis) and Z coordinate (vertical axis) after the transformation. Referring to Figure 10, it can be confirmed that in the speaker arrangement of 5.1.4, the azimuth angles of the middle and upper speakers coincide, so no correction is made according to the height of the endpoints of the line segment domain, and a transformation result equivalent to that of the conventional method is obtained.

[0050] The above experiments confirmed that the acoustic metadata coordinate system transformation device 1 has a transformation performance comparable to that of conventional methods based on 5.1.4 for the speaker arrangement in 5.1.4, and, unlike conventional methods, for 7.1.4 and 22.2ch, it is possible to obtain transformation results that match the coordinate values ​​specified in Non-Patent Document 1 at all speaker positions, and it is also possible to confirm that it has an effect of correcting coordinate deviations in the elevation direction.

[0051] (Variations) The acoustic metadata coordinate system transformation device 1 has been described above. The speaker placement information does not necessarily have to conform to a standard value as long as it is inconsistent with the input listening room information. In addition to the placed speakers, it is also possible to control the amount of movement of the sound object per unit angle within the rectangular area by adding virtual speakers.

[0052] For example, if the correspondence between the coordinate systems of the listening room reference point is as shown in Table 1, and the speaker placement information is specified as shown in Table 9, with the placement of 22.2ch speakers plus the placement of virtual speakers at ±110 degrees and ±150 degrees of the middle layer using orthogonal coordinate values, then the listening room wall will be divided into 14 rectangular regions as shown in Table 10.

[0053] [Table 9] [Table 10]

[0054] Figure 11 shows the results of coordinate system transformations performed on the azimuth angle of the audio object in 1-degree increments from -180 degrees to 180 degrees, assuming a speaker arrangement of 22.2ch and virtual speakers at mid-level ±110 degrees and mid-level ±150 degrees. The graph on the left of this figure shows the relationship between the azimuth angle before transformation (horizontal axis) and the X-coordinate after transformation (vertical axis), and the graph on the right shows the relationship between the azimuth angle before transformation (horizontal axis) and the Y-coordinate after transformation (vertical axis). From Figure 11, it can be confirmed that the amount of movement per unit angle of the audio object changes in the rectangular regions from azimuth angle ±90 degrees to azimuth angle ±135 degrees (combined same order) and from azimuth angle ±135 degrees to azimuth angle ±180 degrees (combined same order).

[0055] Furthermore, the amount of movement of the audio object per unit angle within this rectangular region can also be controlled by employing a calculation rule for the internal division ratio of the endpoints of the line segment region that differs from that of equation (6). For example, a function whose domain is the azimuth angle of the endpoints of the line segment region and whose range is constrained to [0,1] (e.g., a polynomial function, a spline function, or a sigmoid function), or a method using the tangent rule used in Non-Patent Document 1 may be used.

[0056] <Second Embodiment> Next, a second embodiment will be described. In the second embodiment, an acoustic metadata coordinate system transformation device will be described that transforms Cartesian coordinate values ​​(x, y, z) in acoustic metadata attached to an audio signal into polar coordinate values ​​(φ, θ, d).

[0057] Figure 12 is a block diagram showing an example configuration of the acoustic metadata coordinate system transformation device 2 according to the second embodiment. The acoustic metadata coordinate system transformation device 2 shown in Figure 12 comprises a rectangular area definition unit 21, a line segment area definition unit 22, and a coordinate derivation unit 23. Note that the acoustic metadata coordinate system transformation device 1 according to the first embodiment may also have the functions of the acoustic metadata coordinate system transformation device 2.

[0058] Figure 12 is a flowchart showing the procedure for converting the position information of an audio object from Cartesian coordinate values ​​to polar coordinate values ​​in the acoustic metadata coordinate system transformation device 2. This algorithm can be broadly divided into three steps, from step S4 to step S6.

[0059] In step S4, the rectangular area definition unit 21 defines multiple rectangular areas on the wall surface of the listening room, with endpoints at positions where the correspondence between the coordinate systems is known for the upper, middle, and lower layers, based on the correspondence between polar coordinate values ​​and Cartesian coordinate values ​​at the reference point of the listening room and the positions of each speaker to be placed.

[0060] In step S5, the line segment domain definition unit 22 defines the endpoints of the line segment domain by correcting the correspondence between the orthogonal coordinates and polar coordinates that occur in the Z-axis direction using the orthogonal coordinate values ​​of the sound object.

[0061] In step S6, the coordinate derivation unit 23 searches for the line segment region to which the audio object belongs, and derives the polar coordinate value of the audio object based on the correspondence between the Cartesian coordinate values ​​and polar coordinates of the endpoints of the line segment region.

[0062] (Explanation of the operation in step S4) In step S41, polar and Cartesian coordinates are input for the reference point of the listening room and each speaker position in order to define the correspondence between Cartesian and polar coordinate values. In step S42, multiple rectangular areas are defined on the wall surface of the listening room. The operation of step S4 is the same as the operation of step S1 described in the first embodiment, so a detailed explanation is omitted.

[0063] (Explanation of Step S5) In step S51, it is determined whether the audio object is located above the upper layer or below the lower layer.

[0064] In step S52, the endpoints of the line segment domain are defined based on the Z coordinate value of the sound object. Specifically, if the sound object is located on the ceiling or floor of the listening room, the endpoints of the rectangular domain defined in the upper or lower layer are used as the endpoints of the line segment domain. If the sound object is located below the upper layer or above the lower layer, the intersection line of the horizontal plane and the rectangular domain at the height z of the sound object is newly defined as the line segment domain. If the Z coordinate value z of the sound object does not coincide with the Z coordinate value z=0 of the middle layer, and the azimuth angles of the rectangular domain endpoints differ between the middle layer and the upper / lower layers, the azimuth angle φ of the polar coordinate value corresponding to the Cartesian coordinate value (x, y) of the line segment domain endpoint is corrected, taking into account the difference in the azimuth angle of the polar coordinate value corresponding to the Cartesian coordinate value. As a correction rule for the line segment domain endpoints, for example, equation (4) as described above may be used.

[0065] (Explanation of the operation in step S6) Finally, in step S61, polar coordinate values ​​are derived based on the information of the endpoints of the line segment region to which the sound object belongs. The search for the line segment region to which the sound object belongs is determined based on the relative magnitudes of the azimuth angles φ' in the Cartesian coordinate system of the sound object and the endpoints of the line segment region. The azimuth angle in the Cartesian coordinate system is calculated, for example, using equation (8) specified in Non-Patent Document 1. However, if the sound object is located on the ceiling or floor, the endpoints of the rectangular region in the upper or lower layer are considered as the endpoints of the line segment region, and the determination is made based on the relative magnitudes of their azimuth angles and the azimuth angle of the sound object.

number

[0066] Next, the azimuth angle of the sound object is derived using the endpoints of the line segment region to which the sound object belongs. For example, the internal division ratio g of the line segment region endpoints of the sound object position in the Cartesian coordinate system obtained by equation (9) specified in Non-Patent Document 1. r ,g l It may also be calculated using equation (10). Equation (10) corresponds to the inverse function of the function used to derive the internal division ratio of a line segment region in a Cartesian coordinate system from polar coordinate values ​​as exemplified in the first embodiment. However, g mThis is the internal division ratio at the position corresponding to the average of the azimuth angles of the left and right endpoints of the line segment region, and is calculated, for example, by equation (11). Alternatively, instead of equation (10), the azimuth angle calculation rule using the tangent function specified in Non-Patent Document 1 may be used.

number

number

number

[0067] Next, the elevation angle θ of the audio object is derived from the Z coordinate. For example, the elevation angle may be calculated using equation (13) with the elevation angle θ' in the Cartesian coordinate system calculated by equation (12) specified in Non-Patent Document 1. However, the variable θ T ,θ B ,θ' T ,θ' B θ' represents the elevation angles of the upper and lower layers in polar coordinates and Cartesian coordinates, respectively. T ,θ' B Considering the ratio of the horizontal and vertical distances between the listening point and the upper front position, we decided to use 45 degrees, θ T ,θ B For example, the elevation angle of the upper and lower speakers is 30 degrees as defined in Non-Patent Document 2. Also, instead of equation (13), |θ' as defined in Non-Patent Document 1 is used. T |=|θ' B You may also use a calculation rule that assumes |.

number

number

[0068] Finally, we derive the distance d of the audio object. For example, it may be calculated using equation (14). Alternatively, instead of equation (14), |θ' as defined in Non-Patent Document 1 may be used. T |=|θ'B You may also use a calculation rule that assumes |.

number

[0069] (Examples) To confirm the effectiveness of the Acoustic Metadata Coordinate System Transformation Device 2, numerical experiments were conducted comparing coordinate system transformations performed using the conventional method and the present invention with standard values. In the figure, "Conventional Method" refers to the case where the coordinate system is transformed based on the correspondence between polar coordinate values ​​and Cartesian coordinate values ​​in 5.1.4 (Non-Patent Literature 1). "The Present Invention" refers to the case where the coordinate system is transformed using the Acoustic Metadata Coordinate System Transformation Device 2. "Standard Value" refers to the case where the coordinate system is transformed based on the correspondence between polar coordinate values ​​and Cartesian coordinate values ​​according to the standard at the speaker position (however, transformation is only possible at the location where the speaker is present).

[0070] (Example: 22.2ch) Assuming the correspondence between the coordinate systems of the listening room reference point is as shown in Table 1 above, and the speaker arrangement is 22.2ch, i.e., the correspondence between the coordinate systems at the speaker positions is as shown in Table 2 above, the listening room wall surface is divided into 10 rectangular regions as shown in Table 3 above.

[0071] Figure 14 shows the results of coordinate system transformations performed in 1-degree increments from -180 degrees to 180 degrees in the Cartesian coordinate system of audio objects in the upper, middle, and lower layers, assuming a speaker configuration of 22.2ch. The graph in this figure shows the relationship between the Cartesian coordinates (x,y) (horizontal axis) before transformation and the azimuth angle after transformation (vertical axis). Referring to Figure 14, it can be confirmed that while the transformation results using the conventional method do not partially match the transformation results based on standard values ​​in the upper and middle layers, the transformation results according to the present invention completely match the transformation results based on standard values.

[0072] Figure 15 shows the results of coordinate system transformation for audio objects moving linearly from position (-1,+1,-1) to position (-1,+1,+1) and audio objects moving linearly from position (+1,+1,-1) to position (+1,+1,+1) in a Cartesian coordinate system, assuming a speaker configuration of 22.2ch. The graph in this figure shows the relationship between the azimuth angle (horizontal axis) and elevation angle (vertical axis) after the transformation. Referring to Figure 15, if we adhere to the standard values ​​for speaker positions, when moving linearly from position (-1,+1,-1) to position (-1,+1,+1) in a Cartesian coordinate system, the movement should be from speaker B±045, via speaker M±030, to speaker U±045. While conventional methods consistently maintain an azimuth angle of ±30 degrees, the proposed method transforms the coordinate values ​​so that the audio object moves as expected, demonstrating the effectiveness of correcting the coordinate shift in the Z-axis direction.

[0073] (Example: 7.1.4) If the correspondence between the coordinate systems of the listening room reference point is the same as in the 22.2ch case as shown in Table 3 above, and the speaker arrangement is as in 7.1.4, i.e., the correspondence between the coordinate systems at the speaker positions is as shown in Table 4 above, then the listening room wall is divided into seven rectangular regions as shown in Table 5 above.

[0074] Figure 16 shows the results of coordinate system transformations performed on the azimuth angle of the audio object in the Cartesian coordinate system in 1-degree increments from -180 degrees to 180 degrees in the upper, middle, and lower layers, assuming a speaker arrangement of 7.1.4. The graph in this figure shows the relationship between the Cartesian coordinates (x,y) (horizontal axis) before transformation and the azimuth angle after transformation (vertical axis). Referring to Figure 16, it can be confirmed that while the transformation results using the conventional method do not fully match the transformation results based on the standard values ​​in the upper and middle layers, the transformation results according to the present invention fully match the transformation results based on the standard values.

[0075] Figure 17 shows the results of coordinate system transformation for audio objects moving linearly from position (-1,+1,-1) to position (-1,+1,+1) and from position (+1,+1,-1) to position (+1,+1,+1) in a Cartesian coordinate system, assuming speaker placement 7.1.4. The graph in this figure shows the relationship between the azimuth angle (horizontal axis) and elevation angle (vertical axis) after transformation. Referring to Figure 17, while the transformation results using the conventional method always maintain an azimuth angle of ±30 degrees, the proposed method transforms the coordinate values ​​so that the audio object moves as expected, and the effect of correcting the coordinate shift in the Z-axis direction can be confirmed.

[0076] (Example: 5.1.4) Assuming the correspondence between the coordinate systems of the listening room reference point is as shown in Table 6 above, and the speaker arrangement is as described in 5.1.4, i.e., the correspondence between the coordinate systems at the speaker positions is as shown in Table 7 above, the listening room wall is divided into five rectangular regions as shown in Table 8 above.

[0077] Figure 18 shows the results of coordinate system transformations performed in 1-degree increments from -180 degrees to 180 degrees in the Cartesian coordinate system of audio objects in the upper, middle, and lower layers, assuming speaker placement 5.1.4. The graph in this figure shows the relationship between the Cartesian coordinates (x,y) (horizontal axis) before transformation and the azimuth angle after transformation (vertical axis). Referring to Figure 18, it can be confirmed that the transformation results of the conventional method and the present invention completely match the transformation results based on standard values ​​at speaker positions, and that there are no significant differences between the two transformation results at other positions.

[0078] Figure 19 shows the results of coordinate system transformation for audio objects moving linearly from position (-1,+1,-1) to position (-1,+1,+1) and audio objects moving linearly from position (+1,+1,-1) to position (+1,+1,+1) in a Cartesian coordinate system, when the speaker arrangement is as described in 5.1.4. The graph in this figure shows the relationship between the azimuth angle (horizontal axis) and elevation angle (vertical axis) after the transformation. Referring to Figure 19, it can be confirmed that in the speaker arrangement of 5.1.4, the azimuth angles of the middle and upper speakers coincide, so no correction is made according to the height of the endpoints of the line segment domain, and a transformation result equivalent to that of the conventional method is obtained.

[0079] Therefore, this experiment confirms that the speaker arrangement in 5.1.4 has a conversion performance comparable to that of conventional methods based on 5.1.4, and that, unlike conventional methods, for 7.1.4 and 22.2ch, conversion results matching the coordinate values ​​specified in Non-Patent Document 1 can be obtained at all speaker positions, and that there is also a correction effect on coordinate deviations in the Z-axis direction.

[0080] Figure 20 shows the results of converting polar coordinate values ​​to Cartesian coordinate values ​​using the acoustic metadata coordinate system conversion device 1 according to the first embodiment, with listening room information and speaker placement information as described in the above embodiment (22.2ch / 7.1.4 / 5.1.4), at positions in the polar coordinate system of the sound object in 5-degree increments from -180 degrees to 180 degrees in azimuth angle and in 15-degree increments from -90 degrees to 90 degrees in elevation angle, and then converting those Cartesian coordinate values ​​back to polar coordinate values ​​using the acoustic metadata coordinate system conversion device 2. The graph in this figure shows the relationship between the azimuth angle in polar coordinates before conversion (horizontal axis) and the azimuth angle after conversion (vertical axis). Referring to Figure 20, it can be confirmed that, except when the azimuth angle φ before conversion is ±180 degrees, the relationship with the azimuth angle φ^ after conversion lies on the straight line φ^=φ, and the values ​​before and after conversion are the same. Furthermore, even at positions with azimuth angles of +180 degrees and -180 degrees, the two azimuth angles indicate the same position, and therefore, they can be said to take the same value before and after the transformation. Thus, it can be said that the acoustic metadata coordinate system transformation device 1 and the acoustic metadata coordinate system transformation device 2 according to the first embodiment can uniquely perform the inverse transformation.

[0081] (Variations) The acoustic metadata coordinate system transformation device 2 has been described above. The speaker placement information does not necessarily have to conform to a standard value as long as it does not contradict the input listening room information. In addition to the placed speakers, it is also possible to control the amount of movement of the sound object per unit angle within the rectangular area by adding virtual speakers.

[0082] For example, in the above embodiment, for an audio object moving linearly from position (±1,+1,-1) to position (±1,+1,+1), the present invention confirmed that it moves in accordance with the standard value as shown in Figure 15. However, in a Cartesian coordinate system, the audio object moves as if rising vertically, but in polar coordinates, it moves while changing its azimuth angle. For example, if it is desired that the azimuth angle be maintained in the polar coordinate system as in the Cartesian coordinate system, the correspondence between each coordinate system of the listening room reference point can be specified as shown in Table 11, with the azimuth angles at the upper and lower left front and right front of the upper and lower layers set to 30 degrees, the same as the middle layer, and the speaker placement information for 22.2ch can be specified as shown in Table 12. In this case, the listening room wall surface is divided into 10 rectangular regions as shown in Table 13. For speakers whose placement is not specified, the azimuth angle is derived virtually, similar to the point where the correspondence between coordinate systems where neither the speaker nor the reference point exists is not defined. [Table 11] [Table 12] [Table 13]

[0083] Figure 21 shows the results of coordinate system transformations performed in 1-degree increments from -180 degrees to 180 degrees in the Cartesian coordinate system of audio objects in the upper, middle, and lower layers, assuming a speaker configuration of 22.2ch and the corresponding azimuth angles of the left front (x,y)=(-1,1) and right front (x,y)=(1,1) of the listening room are ±30 degrees regardless of height z. The graph in this figure shows the relationship between the Cartesian coordinates (x,y) before transformation (horizontal axis) and the azimuth angle after transformation (vertical axis).

[0084] Figure 22 shows the results of a coordinate system transformation for audio objects moving linearly from position (-1,+1,-1) to position (-1,+1,+1) and from position (+1,+1,+1) to position (+1,+1,+1) in a Cartesian coordinate system, assuming a speaker configuration of 22.2ch and the corresponding azimuth angles of the left front (x,y)=(-1,1) and right front (x,y)=(1,1) in the listening room are ±30 degrees regardless of height z. The graph in this figure shows the relationship between the azimuth angle (horizontal axis) and elevation angle (vertical axis) after the transformation.

[0085] Referring to Figures 21 and 22, the result shows that only the azimuth angle of ±45 degrees for the upper and lower layers does not match the standard value, but it can be confirmed that in the present invention as well, the sound object moves while maintaining an azimuth angle of ±30 degrees.

[0086] Furthermore, the unit movement amount of the audio object within this rectangular region can also be controlled by adding a virtual speaker or by adopting a calculation rule different from that in equation (10). For example, a function whose domain is [0,1] and whose range is constrained by the azimuth angle of the endpoints of the rectangular region (e.g., a polynomial function or a spline function), or a method using the tangent rule used in Non-Patent Document 1 may be used.

[0087] (program) To function as the acoustic metadata coordinate system transformation devices 1 and 2 described above, it is also possible to use a computer capable of executing program instructions. Here, the computer may be a general-purpose computer, a dedicated computer, a workstation, a PC (Personal Computer), an electronic notepad, etc. The program instructions may be program code, code segments, etc., for executing the required tasks.

[0088] A computer comprises a processor, a memory unit, an input unit, an output unit, and a communication interface. The processor may be a CPU (Central Processing Unit), MPU (Micro Processing Unit), GPU (Graphics Processing Unit), DSP (Digital Signal Processor), SoC (System on a Chip), etc., and may be composed of multiple processors of the same or different types. The processor controls each of the above components and performs various calculations by reading and executing programs from the memory unit. At least a part of these processes may be implemented in hardware. The input unit is an input interface that receives user input operations and acquires information based on user operations, such as a pointing device, keyboard, or microphone. The output unit is an output interface that outputs information, such as a display or speaker. The communication interface is an interface for communicating with external devices, such as a LAN (Local Area Network) interface.

[0089] The program may be recorded on a computer-readable recording medium. Using such a medium, the program can be installed on the computer. The recording medium on which the program is recorded may be a non-transitory recording medium. Non-transitory recording media are not particularly limited, but may include, for example, CD-ROMs, DVD-ROMs, or USB (Universal Serial Bus) memory. Alternatively, the program may be downloaded from an external device via a network.

[0090] Although the embodiments described above are representative examples, it will be apparent to those skilled in the art that many modifications and substitutions are possible within the spirit and scope of the present invention. Therefore, the present invention should not be interpreted as being limited by the embodiments described above, and various modifications or changes are possible without departing from the scope of the claims. For example, it is possible to integrate multiple component blocks shown in the configuration diagram of the embodiments, or to divide a single component block. [Explanation of symbols]

[0091] 1,2 Acoustic metadata coordinate system transformation device 11,21 Rectangular area definition part 12,22 Line segment region definition section 13,23 Coordinate Derivation Section

Claims

1. An acoustic metadata coordinate system transformation device that converts the positional information of an audio object attached to acoustic metadata from polar coordinate values ​​to orthogonal coordinate values, A rectangular area defining unit defines multiple rectangular areas on the wall surface of the listening room, with endpoints at positions where the coordinate system correspondence is known for the upper, middle, and lower layers, based on the correspondence between polar coordinate values ​​and Cartesian coordinate values ​​at the reference point of the listening room and the positions of each speaker to be placed. A line segment region defining unit derives the Z coordinate value z of the sound object, defines the intersection point of the horizontal plane and the edge of the rectangular region at the height of z as the endpoint of the line segment region, and corrects the coordinate value of the endpoint of the line segment region by considering the difference in the azimuth angle of the polar coordinate value corresponding to the orthogonal coordinate value if z does not match the Z coordinate value of the middle layer. A coordinate derivation unit that searches for the line segment region to which the sound object belongs and derives the orthogonal coordinate values ​​of the sound object using the endpoints of the line segment region, An acoustic metadata coordinate system transformation device equipped with the following features.

2. An acoustic metadata coordinate system transformation device that converts the positional information of an audio object attached to acoustic metadata from orthogonal coordinate values ​​to polar coordinate values, A rectangular area defining unit defines multiple rectangular areas on the wall surface of the listening room, with endpoints at positions where the coordinate system correspondence is known for the upper, middle, and lower layers, based on the correspondence between polar coordinate values ​​and Cartesian coordinate values ​​at the reference point of the listening room and the positions of each speaker to be placed. A line segment domain defining unit defines the endpoints of the line segment domain at the intersection of the horizontal plane and the edges of the rectangular region at the height z of the sound object, and corrects the coordinate values ​​of the endpoints of the line segment domain by considering the difference in the azimuth angle of the polar coordinate value corresponding to the orthogonal coordinate value if z does not match the Z coordinate value of the middle layer, A coordinate derivation unit searches for the line segment region to which the sound object belongs and derives the polar coordinate value of the sound object using the endpoints of the line segment region. An acoustic metadata coordinate system transformation device equipped with the following features.

3. The reference point of the aforementioned listening room is, A total of 12 points at the four corners of each of the upper, middle, and lower levels of the aforementioned listening room, or The acoustic metadata coordinate system transformation device according to claim 1 or 2, wherein the four corners of each of the upper, middle, and lower levels of the listening room, as well as a total of 24 points in front of, behind, and to the side of the listener.

4. A program for causing a computer to function as an acoustic metadata coordinate system transformation device according to claim 1 or 2.

Citation Information

Patent Citations

  • ITRBS.2051-1,

  • ITRBS.2127-0,

  • Audio localization setting apparatus, method, and program

    JP2015179986A

  • Apparatus for converting object position of an audio object, audio stream provider, audio content production system, audio playback apparatus, method and computer program

    JP2021513775A