Information processing method, information processing system, and program

The method addresses sound localization issues in spatial audio format conversions by using a conversion device with position and gain adjustments, ensuring consistent audio playback across formats like MPEG-H 3D Audio and AC-4.

WO2026005047A1PCT designated stage Publication Date: 2026-01-02SONY GROUP CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/023327
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-27
Filing Date
2025-06-27
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing methods for converting between egocentric and allocentric spatial audio formats risk losing sound localization due to differences in rendering methods, leading to inconsistent audio playback experiences.

Method used

An information processing method and system that convert position data of sound sources between different spatial audio formats, such as MPEG-H 3D Audio and AC-4, by using a conversion device with units like speaker layout input, position correction, gain ratio calculation, and panning ratio calculation to maintain sound localization.

Benefits of technology

Enables seamless data sharing and consistent sound localization across different spatial audio formats, ensuring accurate audio playback by adjusting position and gain ratios during format conversion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025023327_02012026_PF_FP_ABST
    Figure JP2025023327_02012026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to an information processing method, an information processing system, and a program that enable data sharing between different spatial audio schemes. In the method of the present invention: first position data of a first sound source in a first data format are acquired; an audio signal of the first sound source is converted from object-based audio to channel-based audio on the basis of first position data of the first sound source; and second position data of a second sound source other than the first sound source in the first data format are converted into third position data in a second data format. The present disclosure is applicable, for example, to information processing methods, conversion methods, information processing systems, programs, information processing devices, conversion devices, and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing method, information processing system, and program

[0001] The present disclosure relates to an information processing method, an information processing system, and a program, and more particularly to an information processing method, an information processing system, and a program that enable data sharing between different spatial audio formats.

[0002] In recent years, object-based audio technology has been attracting attention. There are several data formats and rendering methods for object-based audio data (see, for example, Patent Documents 1 and 2). For example, there are data formats known as egocentric and allocentric. To improve the convenience of distribution services and playback devices, it has been desirable to share waveform signals and metadata between these formats.

[0003] For example, when converting from an egocentric system to an allocentric system, the values ​​of Azimuth (horizontal angle), Elevation (vertical angle), and Radius (radius) in polar coordinate format based on the listening position contained in the metadata can be converted simply based on the laws of physics into x-coordinate, y-coordinate, and z-coordinate values, which are positioning information in absolute coordinate format.

[0004] International Publication No. WO 2016 / 203994 International Publication No. WO 2024 / 228269

[0005] However, because the rendering methods used for the audio data before and after conversion are different, there is a risk that the same sense of sound localization may not be achieved during playback simply by converting the localization information as described above.

[0006] The present disclosure has been made in light of such circumstances, and aims to enable data sharing between different spatial audio formats.

[0007] An information processing method according to one aspect of the present technology is an information processing method that acquires first position data of a first sound source in a first data format, converts an audio signal of the first sound source from object-based audio to channel-based audio based on the first position data of the first sound source, and converts second position data of a second sound source other than the first sound source in the first data format into third position data in a second data format.

[0008] An information processing system according to one aspect of the present technology is an information processing system that acquires first position data of a first sound source in a first data format, converts an audio signal of the first sound source from object-based audio to channel-based audio based on the first position data of the first sound source, and converts second position data of a second sound source other than the first sound source in the first data format into third position data in a second data format.

[0009] A program according to another aspect of the present technology is a program for causing a computer to execute a process of acquiring first position data of a first sound source in a first data format, converting an audio signal of the first sound source from object-based audio to channel-based audio based on the first position data of the first sound source, and converting second position data of a second sound source other than the first sound source in the first data format into third position data in a second data format.

[0010] In an information processing method, an information processing system, and a program according to one aspect of the present technology, first position data of a first sound source in a first data format is acquired, and based on the first position data of the first sound source, the audio signal of the first sound source is converted from object-based audio to channel-based audio, and second position data of a second sound source other than the first sound source in the first data format is converted into third position data in a second data format.

[0011] FIG. 1 is a diagram showing an example of a gain ratio of each channel in an allocentric system. FIG. 2 is a block diagram showing an example of the main configuration of a conversion device. FIG. 3 is a diagram showing an example of a gain ratio of each speaker. FIG. 4 is a diagram showing an example of a polar coordinate format. FIG. 5 is a flowchart showing an example of the flow of a conversion process. FIG. 6 is a diagram showing an example of an audio signal. FIG. 7 is a diagram showing an example of an audio signal. FIG. 8 is a block diagram showing an example of the main configuration of a conversion device. FIG. 9 is a flowchart showing an example of the flow of a conversion process. FIG. 10 is a block diagram showing an example of the main configuration of a conversion device. FIG. 11 is a flowchart showing an example of the flow of a conversion process. FIG. 12 is a diagram showing an example of a user interface of an application. FIG. 13 is a diagram showing an example of a user interface of an application. FIG. 14 is a diagram showing an example of the main configuration of a conversion system. FIG. 15 is a block diagram showing an example of the main configuration of a computer.

[0012] Hereinafter, modes for carrying out the present disclosure (hereinafter referred to as embodiments) will be described. The description will be made in the following order: 1. Literature supporting technical content and technical terminology 2. Object-based audio 3. Format conversion 4. Supplementary notes

[0013] <1. Literature, etc. supporting technical content and technical terminology> The scope of disclosure of the present technology includes not only the content described in the embodiments, but also the content described in the following patent documents and non-patent documents that were publicly known at the time of filing, as well as the content of other documents referenced in the following patent documents and non-patent documents.

[0014] Patent document 1: (described above) Patent document 2: (described above) Non-patent document 1: ISO / IEC 23008-3 Information technology - High efficiency coding and media delivery in heterogeneous environments - Part 3: 3D audio Non-patent document 2: Ville Pulkki, "Virtual Sound Source Positioning Using Vector Base Amplitude Panning", Journal of AES, vol.45, no.6, pp.456-466, 1997 Non-patent document 3: ETSI TS 103 448 V1.1.1(2016-09) "AC-4 Object Audio Renderer for Consumer Use" Non-patent document 4: Recommendation ITU-R BS.2051-3(05 / 2022) "Advanced sound system for program production" Non-patent document 5: Recommendation ITU-R BS.2076-2(10 / 2019) "Audio definition model" Non-patent document 6: EBU Tech 3285 - Specification of the Broadcast Wave Format(BWF) version 2.0

[0015] In other words, the contents of any of the above-mentioned patent documents and non-patent documents, as well as the contents of other documents referenced in the above-mentioned patent documents and non-patent documents, are also used as the basis for determining the support requirement. For example, even if a syntax or term described in any of the above-mentioned patent documents or non-patent documents is not directly defined in this disclosure, it is still within the scope of this disclosure and satisfies the support requirement of the claims.

[0016] 2. Object-Based Audio Object-Based Audio Data Format Object-based audio technology has been attracting attention in recent years.

[0017] In object-based audio, audio data is composed of waveform data for an object and metadata that indicates the sound positioning information of the object based on that waveform data, and during playback, rendering is performed to the desired number of channels based on the metadata.

[0018] There are several methods for data formatting and rendering audio data for object-based audio.

[0019] One is a data format called MPEG (Moving Picture Experts Group)-H 3D Audio (see, for example, Non-Patent Document 1).

[0020] In MPEG-H 3D Audio, the positioning information is position information in the form of polar coordinates that indicates the direction of an object relative to the listening point (listening position), and rendering is performed based on this position information using a rendering method called VBAP (Vector Based Amplitude Panning) (see, for example, Non-Patent Document 2).

[0021] On the other hand, there is a data format called AC-4 (see, for example, Non-Patent Document 3).

[0022] In AC-4, the positioning information is in the form of absolute coordinates that indicate the position in xyz space, and rendering is performed by calculating the gain ratio of each speaker from the ratio of the distance between the object and the speaker in each of the x-axis, y-axis, and z-axis directions based on that position information.

[0023] These data formats, more specifically rendering methods, are called the egocentric method and the allocentric method, based on their respective characteristics.

[0024] However, in order to improve the convenience of distribution services and playback devices, it has been desired to share waveform signals and metadata between the two formats.

[0025] Here, for example, consider converting egocentric audio data into allocentric audio data.

[0026] In this case, the Azimuth (horizontal angle), Elevation (vertical angle), and Radius values ​​contained in the metadata, which are in polar coordinate format based on the listening position, can be converted into x-coordinate, y-coordinate, and z-coordinate values, which are positioning information in absolute coordinate format, simply based on the laws of physics.

[0027] However, because the rendering methods used for the audio data before and after conversion are different, there is a risk that the same sense of sound localization may not be achieved during playback simply by converting the localization information as described above.

[0028] A specific example will be explained. For example, in the egocentric system, an object is located at the polar coordinates (Az, El, Rad) = (+30, 0, 1), and the sound of that object is to be played back using a 7.1.4ch speaker layout. Note that in the following, the polar coordinate Az indicates the Azimuth (horizontal angle). Furthermore, the polar coordinate El indicates the Elevation (vertical angle). Furthermore, the polar coordinate Rad indicates the Radius.

[0029] In this case, VBAP positions the sound image of the object at +30 degrees, regardless of the exact position of the speaker. Now, let's assume that this waveform signal and metadata are converted and played back in an allocentric format. If the position information is simply converted based on the laws of physics, the polar coordinates (Az, El, Rad) = (+30, 0, 1) become coordinates (x, y, z) = (-0.5, 0.87, 0) according to the following equation (1).

[0030] ...(1)

[0031] When playback is performed using the same 7.1.4ch speaker layout in the Allocentric format based on this position information, signals are distributed to each channel with the gain ratios shown in the table of FIG.

[0032] If this 7.1.4ch speaker layout were arranged in the standard positions specified in ITU-R BS.2051 (see, for example, Non-Patent Document 4) in the table, not only would the sound image be localized approximately +20 degrees between the L and C speakers, but the gain would also be distributed to the Ls and Rs speakers, resulting in a somewhat broad sound image, or in other words, an unclear localization. Therefore, if something originally produced using the egocentric method with a localization of +30 degrees were simply converted from its polar coordinate positioning information to absolute coordinate format based on the laws of physics, it would end up being localized approximately +20 degrees and unclear, which could significantly change the sound localization and impression.

[0033] <3. Format Conversion> <Conversion Device 1> Therefore, data conversion is performed while maintaining sound localization in different spatial audio formats. For example, first position data of a first sound source in a first data format is acquired, and based on the first position data of the first sound source, the audio signal of the first sound source is converted from object-based audio to channel-based audio, and second position data of a second sound source other than the first sound source in the first data format is converted into third position data in a second data format. In this way, spatial audio data can be shared between both the egocentric and allocentric formats.

[0034] Fig. 2 is a block diagram showing an example of the configuration of a conversion device, which is one aspect of an information processing device to which the present technology is applied. The conversion device 100 shown in Fig. 2 is a device that converts object metadata in a data format in which the rendering method is egocentric, such as MPEG-H 3D Audio, into object metadata in a data format in which the rendering method is allocentric, such as AC-4. In other words, the first data format may be an egocentric data format, and the second data format may be an allocentric data format.

[0035] Audio data in an egocentric data format (first data format), i.e., a data format in which the rendering method is egocentric, includes waveform data for one or more objects and object metadata for each object. The audio data may also include waveform data for each channel.

[0036] Egocentric object metadata includes at least the localization position of the object's sound, that is, object position data that indicates the object's position.

[0037] In particular, in egocentric object metadata, object position data is in the form of polar coordinates that indicate the relative position of an object as seen from a listening position (listening point) that serves as a reference in virtual space. Specifically, the object position data is in the form of polar coordinates (Az, El, Rad) that indicate the relative position of an object as seen from the listening position.

[0038] Furthermore, egocentric object metadata may include, as appropriate, name information indicating the name of an object (object name), type information indicating the type of object (e.g., vocals or guitar), size information indicating the size of the object, priority information indicating the priority of the object, spread information indicating the degree of spread of the object, and gain information indicating the gain value of the object. For example, each piece of object metadata may have a preset name or value, or the object metadata may be changeable by the user as appropriate. For example, priority information may use a value from 0 to 7, with the highest priority object being assigned a priority of 7 and the lowest priority object being assigned a priority of 0. Furthermore, in real-time processing, the name and value of each piece of object metadata may be dynamically changed.

[0039] In contrast, audio data in an Allocentric data format (second data format), i.e., a data format using the Allocentric rendering method, includes waveform data, which is an audio signal for reproducing the sound of one or more objects, and object metadata for each object. Note that the audio data may include waveform data (audio signals) for each channel.

[0040] The object metadata includes at least the localization position of the sound of the object, that is, object position data that indicates the position of the object.

[0041] In particular, in Allocentric object metadata, object position data is data in the form of absolute coordinates that indicate the absolute position of an object in a virtual space. Specifically, the object position data is absolute coordinates (x, y, z) that indicate the position in absolute coordinate space (xyz space) where the object is placed.

[0042] The object metadata also includes object name information, type information, size information, priority information, spread information, gain information, and the like, as appropriate.

[0043] As described above, the object waveform data is the same between egocentric audio data and allocentric audio data, but the rendering method and the coordinate format of the object position data are different. Furthermore, as described above, simply converting the coordinate format of the object position data based on the laws of physics could significantly change the sound localization and impression. Therefore, the converting device 100 generates allocentric object position data by appropriately converting the coordinates of the egocentric object position data so as to maintain the sound localization, thereby enabling spatial audio data to be shared between both the egocentric and allocentric methods.

[0044] Note that Fig. 2 shows the main processing units, data flows, etc., and is not necessarily all that is shown in Fig. 2. In other words, in the conversion device 100, there may be processing units that are not shown as blocks in Fig. 2, or there may be processing or data flows that are not shown as arrows, etc. in Fig. 2.

[0045] 2, the conversion device 100 includes an object metadata conversion unit 101. The object metadata conversion unit 101 converts object metadata from an egocentric data format to an allocentric data format. The object metadata conversion unit 101 includes a speaker layout information input unit 111, a speaker selection unit 112, an object metadata input unit 113, a position information correction unit 114, a gain ratio calculation unit 115, a panning ratio calculation unit 116, a position data determination unit 117, and an object metadata output unit 118.

[0046] Speaker layout information to be used for playback in the Allocentric format after conversion is input to the conversion device 100, and acquired by the speaker layout information input unit 111. This speaker layout information includes speaker configuration information indicating the number of speaker channels, etc., and speaker position information indicating the positions of the speakers. In other words, the speaker layout information input unit 111 acquires speaker layout information for playback of audio data in the second data format. Based on the speaker layout information, the object metadata conversion unit 101 converts the second position data of the second sound source in the first data format into third position data in the second data format. The speaker layout information input unit 111 supplies the speaker configuration information and speaker position information to the speaker selection unit 112.

[0047] The speaker selection unit 112 acquires speaker configuration information and speaker position information supplied from the speaker layout information input unit 111. From the acquired speaker configuration information, the speaker selection unit 112 selects speakers for which gain ratios to be used in subsequent panning ratio calculations (i.e., speakers to be processed) are to be calculated. For example, in allocentric playback, the speaker selection unit 112 selects speakers corresponding to positions (-1,1,0)(1,1,0)(-1,-1,0)(1,-1,0)(-1,1,1)(1,1,1)(-1,-1,1)(1,-1,1), extracts position information of the selected speakers from the speaker position information, and supplies this information to the gain ratio calculation unit 115 as selected speaker position information indicating the positions of the selected speakers. For example, if the speaker layout used in playback is 7.1.4ch, the position information of each speaker is as shown in the table in FIG. 3. 3 are positions defined in Non-Patent Document 3 and Non-Patent Document 4, but the actual speaker layout may differ slightly, in which case the speaker selection unit 112 may supply selected speaker position information indicating the actual speaker positions to the gain ratio calculation unit 115. When the actual speaker positions are supplied to the gain ratio calculation unit 115, conversion is performed assuming playback at those speaker positions, making it possible to reproduce localization closer to that before conversion.

[0048] Meanwhile, object metadata is input to the conversion device 100, and acquired by the object metadata input unit 113. Here, the object metadata is assumed to include object position information as well as gain information, size information, name information, etc. The object metadata input unit 113 extracts object position information from the acquired object metadata and supplies it to the position information correction unit 114. The object metadata input unit 113 also supplies other meta information to the object metadata output unit 118.

[0049] Here, the object position information is configured in polar coordinate format, and for example, the listening position shown in Figure 4 is used as the origin, and the horizontal angle Az indicating the direction and distance of the object, the vertical angle El, and the radius Rad value.

[0050] The position information correcting unit 114 acquires object position information supplied from the object metadata input unit 113. The position information correcting unit 114 corrects the acquired object position information. For example, in the case of the egocentric method, an object may exist below the horizontal plane of the user position, that is, on the spherical surface of the lower hemisphere. In contrast, the allocentric method after conversion does not support object placement below the horizontal plane of the user position. In such cases, the position information correcting unit 114 corrects the object position information.

[0051] More specifically, for example, if El<0.0, the position information correction unit 114 performs a correction such as replacing El with El=0.0. Then, the position information correction unit 114 supplies the corrected object position information to the gain ratio calculation unit 115. On the other hand, if such a correction is not performed, the position information correction unit 114 supplies the object position information supplied from the object metadata input unit 113 to the gain ratio calculation unit 115 as is.

[0052] The gain ratio calculation unit 115 acquires selected speaker position information supplied from the speaker selection unit 112. The gain ratio calculation unit 115 also acquires object position information supplied from the position information correction unit 114. The gain ratio calculation unit 115 calculates, using VBAP, gain ratios to be assigned to each speaker for the object position information given by the object position information in the speaker layout given by the selected speaker position information. For example, when object position information (Az, El, Rad) = (-33, 20, 1) is input for the above speaker layout, the gain ratios of each speaker are as shown in FIG. 3. The gain ratio calculation unit 115 supplies the gain ratio information obtained in this manner to the panning ratio calculation unit 116.

[0053] The panning ratio calculation unit 116 acquires the gain ratio information supplied from the gain ratio calculation unit 115. From the gain ratio of each speaker indicated in the gain ratio information, the panning ratio calculation unit 116 calculates the panning ratio that results in that gain ratio, using the following equations (2) and (3).

[0054] ...(2) ...(3)

[0055] Here, gainL, gainR, ..., gainTbr respectively indicate the gain ratios of speakers L, R, ..., Tbr in Figure 3, ratioX indicates the panning ratio in the X-axis direction, ratioY indicates the panning ratio in the Y-axis direction, and ratioZ indicates the panning ratio in the Z-axis direction. In the above example, ratioX = 0.973217010, ratioY = 1.0, and ratioZ = 0.415914526. The panning ratio calculation unit 116 supplies panning ratio information indicating the panning ratios calculated in this way to the position data determination unit 117.

[0056] The position data determination unit 117 acquires the panning ratio information supplied from the panning ratio calculation unit 116. The position data determination unit 117 determines position data conforming to the post-conversion format from the panning ratios in the X-axis direction, Y-axis direction, and Z-axis direction indicated by the panning ratio information. For example, if the range of position data x is -1.0≦x≦+1.0, the range of position data y is -1.0≦y≦+1.0, and the range of position data z is 0.0≦z≦+1.0, the position data determination unit 117 calculates the position data (x, y, z) of the object using the following equation (4):

[0057] ...(4)

[0058] In the above example, (x, y, z) = (0.946, 1.00, 0.416). The position data determination unit 117 supplies object position data (absolute coordinate format) indicating the position of the object determined in this way in absolute coordinate format to the object metadata output unit 118.

[0059] The object metadata output unit 118 acquires object position data (in absolute coordinate format) supplied from the position data determination unit 117. The object metadata output unit 118 also acquires meta information other than position information supplied from the object metadata input unit 113. The object metadata output unit 118 outputs the acquired object position data and the acquired meta information (other than position information) to the outside of the conversion device 100 as object metadata in the Allocentric format.

[0060] With this configuration, the converting device 100 can convert egocentric object metadata into allocentric metadata while maintaining the sound localization. Therefore, the converting device 100 can share data between different spatial audio formats.

[0061] <Conversion Process Flow 1> An example of the flow of the conversion process executed by the conversion device 100 will be described with reference to the flowchart of FIG.

[0062] When the conversion process starts, in step S101, the speaker selection unit 112 selects speakers to be processed based on the speaker layout information (speaker configuration information and speaker position information) input to the speaker layout information input unit 111, and sets the speaker positions in the Allocentric format after conversion.

[0063] In step S102, the object metadata input unit 113 acquires metadata for the egocentric object audio to be converted.

[0064] In step S103, the object metadata input unit 113 extracts object position information, which is position data, from the acquired metadata.

[0065] In step S104, the position information correction unit 114 determines whether the vertical angle El is less than 0 (El<0). If it is determined that the vertical angle El is less than 0, that is, that the object is located in an area below the user's listening position, the process proceeds to step S105.

[0066] In step S105, the position information correction unit 114 sets the vertical angle El to "0" (El=0). That is, the position of the object is corrected to the same horizontal plane as the user's listening position. When the processing of step S105 ends, the processing proceeds to step S106.

[0067] Also, if it is determined in step S104 that the vertical angle El is equal to or greater than "0", that is, that the object is located at the same height as the user's listening position or in an area above the user's listening position, processing proceeds to step S106.

[0068] In step S106, the gain ratio calculation unit 115 calculates the gains to be assigned to each speaker for the object signal based on the egocentric rendering method (VBAP). That is, the gain ratio calculation unit 115 calculates the gain ratios for each speaker in the second data format based on the speaker layout information.

[0069] In step S107, the panning ratio calculation unit 116 calculates the panning ratios in the X-axis, Y-axis, and Z-axis directions based on the Allocentric system from the gains assigned to each speaker. That is, the panning ratio calculation unit 116 calculates the panning ratios that result in the gain ratios.

[0070] In step S108, the position data determination unit 117 converts the obtained panning ratio into absolute coordinate format, that is, the position data determination unit 117 derives third position data based on the panning ratio.

[0071] In step S109 , the object metadata output unit 118 generates allocentric object metadata including the obtained position data in absolute coordinate format, and outputs the generated object metadata to the outside of the conversion device 100 .

[0072] In step S110, the object metadata conversion unit 101 determines whether all object metadata has been processed. If it is determined that unprocessed object metadata exists, the process returns to step S102. In other words, if multiple object metadata exist in chronological order for one object, the processes from step S102 to step S110 are executed for each object metadata. If it is determined in step S110 that all object metadata has been processed, the process proceeds to step S111.

[0073] In step S111, the object metadata conversion unit 101 determines whether all objects have been processed. If it is determined that an unprocessed object exists, the process returns to step S102. In other words, if there are multiple objects, the processes from step S102 to step S111 are executed for each object. Then, if it is determined in step S111 that all objects have been processed, the conversion process ends.

[0074] By performing the processes described above, the converting device 100 can convert egocentric object metadata into allocentric metadata while maintaining the sound localization. Therefore, the converting device 100 can share data between different spatial audio formats.

[0075] The above example is an example in which all metadata is available and can be executed offline. However, the present technology is not limited to this example and can be applied. For example, if real-time conversion is required in chronological order, the process may be repeated for all objects in chronological order.

[0076] As described above, egocentric object metadata is converted into allocentric metadata. 6 to 8 show examples of rendering output in the egocentric method before conversion, rendering output in the allocentric method with conversion based on simple physical laws, and rendering output in the allocentric method with conversion using this technology.

[0077] 6 shows the speaker playback signal obtained by egocentric rendering of a sweep signal with object position information (Az, El, Rad) = (-33, 20, 1). For comparison, the speaker layout is the same as the allocentric method described below, and is based on the standard positions for 7.1.4ch specified in ITU-R BS.2051 (see Non-Patent Document 4).

[0078] When object position information is converted using a conversion based on simple physical laws, the converted values ​​are (x, y, z) = (0.512, 0.788, 0.342), but when this is rendered using the Allocentric method in 7.1.4ch, the speaker playback signal looks like Figure 7. It is very different from the signal before conversion in Figure 6, and it is easy to imagine that the auditory positioning of the sound will also be very different.

[0079] In contrast, when the object position information is converted using this technology, the converted values ​​become (x, y, z) = (0.946, 1.00, 0.416) as described above, but when this is rendered in 7.1.4ch using the Allocentric method, the speaker playback signal looks like Figure 8. This is very close to the signal before conversion in Figure 6, and it can be expected that the auditory positioning of the sound will also be roughly the same.

[0080] <Conversion Device No. 2> In the above examples, gain information of an object is output as object metadata, but instead, the waveform signal of the corresponding object may be multiplied by the gain value. Fig. 9 is a block diagram showing an example of the main configuration of a conversion device in this case. Similar to the conversion device 100 of Fig. 2, the conversion device 200 shown in Fig. 9 is a device that applies the present technology to convert object metadata in a data format in which the rendering method is egocentric, such as MPEG-H 3D Audio, into object metadata in a data format in which the rendering method is allocentric, such as AC-4.

[0081] Note that Fig. 9 shows the main processing units, data flows, etc., and is not necessarily all that is shown in Fig. 9. In other words, in conversion device 200, there may be processing units that are not shown as blocks in Fig. 9, or there may be processing or data flows that are not shown as arrows, etc. in Fig. 9.

[0082] 9, the conversion device 200 includes an object metadata conversion unit 201 and a multiplication unit 202. The object metadata conversion unit 201 converts object metadata from an egocentric data format to an allocentric data format, similar to the object metadata conversion unit 101. In addition to the configuration of the object metadata conversion unit 101 described with reference to FIG. 2, the object metadata conversion unit 201 further includes a gain information extraction unit 211.

[0083] In the case of the object metadata conversion unit 201 , the object metadata input unit 113 supplies the meta information other than the object position information from the acquired object metadata to the gain information extraction unit 211 .

[0084] The gain information extraction unit 211 acquires meta information other than the position information supplied from the object metadata input unit 113. The gain information extraction unit 211 extracts gain information from the meta information. This gain information indicates the gain of the object. Therefore, it is also referred to as object gain information. The gain information extraction unit 211 supplies the extracted object gain information to the multiplication unit 202. The gain information extraction unit 211 also supplies the extracted meta information, i.e., the meta information other than the position information and gain information, to the object metadata output unit 118.

[0085] The multiplication unit 202 acquires the object gain information supplied from the gain information extraction unit 211. The multiplication unit 202 also acquires the object audio signal input to the conversion device 200. The multiplication unit 202 multiplies the object audio signal by a gain value of the object to be processed, which is indicated by the object gain information, and outputs the object audio signal after multiplication to the outside of the conversion device 200. In other words, it can be said that the object gain information is multiplied by the object audio signal and then output. In other words, in this case, the object gain information is not output as object metadata in the Allocentric format.

[0086] With this configuration, the converting device 200 can convert egocentric object metadata into allocentric metadata while maintaining sound localization, similar to the case of the converting device 100. Therefore, the converting device 200 can share data between different spatial audio formats.

[0087] <Conversion Process Flow Part 2> An example of the flow of the conversion process executed by the conversion device 200 will be described with reference to the flowchart in Fig. 10. In this case, too, the processes from step S201 to step S208 are executed in the same manner as the processes from step S101 to step S108 in Fig. 5.

[0088] In step S209, the gain information extraction unit 211 extracts gain information of the object. Furthermore, the multiplication unit 202 multiplies the object audio signal by the extracted gain information. That is, the multiplication unit 202 reflects the gain value of the second sound source in the first data format in the audio signal of the second sound source. After step S209 is completed, the process proceeds to step S210.

[0089] The processes from step S210 to step S212 are executed in the same manner as the processes from step S109 to step S111 in Fig. 5. If it is determined in step S212 that all objects have been processed, the conversion process ends.

[0090] By performing each process in this manner, the converting device 200 can convert egocentric object metadata into allocentric metadata while maintaining sound localization, similar to the case of the converting device 100. Therefore, the converting device 200 can share data between different spatial audio formats.

[0091] <Conversion Device No. 3> In the above, the egocentric data of all objects is converted into allocentric data, but the present technology is not limited to this example. For example, the egocentric data of some or all objects may be converted into allocentric data for each channel (channel-based format).

[0092] Fig. 11 is a block diagram showing an example of the main configuration of a conversion device in this case. Similar to the conversion device 100 in Fig. 2, the conversion device 300 shown in Fig. 11 is a device that applies the present technology to convert object metadata in a data format in which the rendering method is egocentric, such as MPEG-H 3D Audio, into object metadata in a data format in which the rendering method is allocentric, such as AC-4. However, the conversion device 300 can also convert some objects into data for each channel in the allocentric format (channel-based format).

[0093] Note that Fig. 11 shows the main processing units, data flows, etc., and does not necessarily show everything. In other words, in the conversion device 300, there may be processing units that are not shown as blocks in Fig. 11, or there may be processing or data flows that are not shown as arrows, etc. in Fig. 11.

[0094] As shown in FIG. 11, the conversion device 300 includes a pre-rendering object selection unit 301 , an object metadata conversion unit 302 , a multiplication unit 303 , and a pre-rendering processing unit 304 .

[0095] The pre-rendering object selection unit 301 acquires object metadata and object audio signals input to the conversion device 300. These are egocentric audio data. The pre-rendering object selection unit 301 selects objects to be pre-rendered from all objects. Hereinafter, an object selected as an object to be pre-rendered will also be referred to as a "selected object," and an object not selected will also be referred to as a "non-selected object." Note that, although the example described here is that of selecting an "object to be pre-rendered," it is of course also possible to select an "object not to be pre-rendered" (in which case, the data of the "non-selected object" will be pre-rendered).

[0096] Any method for selecting the object may be used. For example, the pre-rendering object selection unit 301 may select the object to be pre-rendered in accordance with an external selection by a user, an application, or the like. Alternatively, the pre-rendering object selection unit 301 may select the object to be pre-rendered based on the position of the object. That is, the pre-rendering object selection unit 301 may select the object as the first sound source to be converted into channel-based audio based on the area in which the object is located, and convert the audio signal of the object into channel-based audio. For example, the pre-rendering object selection unit 301 may select, as the object to be pre-rendered, an object located in an area not supported by the Allocentric method (e.g., an area below the horizontal plane including the user's listening position). That is, the pre-rendering object selection unit 301 may select, as the first sound source to be converted into channel-based audio, an object located in an area not supported by the second data format. For example, the pre-rendering object selection unit 301 may select, as the object to be pre-rendered, an object whose vertical angle El is less than "0" (El<0) based on object position information included in the object metadata. In other words, the area incompatible with the second data format may be an area below the user's position.

[0097] The pre-rendering object selection unit 301 supplies selected object metadata, which is metadata of the selected object, and selected object audio signals, which are audio signals of the selected objects, to the pre-rendering processing unit 304. The pre-rendering object selection unit 301 also supplies non-selected object metadata, which is metadata of the non-selected objects, to the object metadata conversion unit 302. The pre-rendering object selection unit 301 also supplies the non-selected object audio signals, which are audio signals of the non-selected objects, to the multiplication unit 303.

[0098] Note that all objects can also be subject to pre-rendering. In this case, the pre-rendering object selection unit 301 sets the metadata of all objects as selected object metadata and supplies the audio signals of all objects as selected object audio signals to the pre-rendering processing unit 304. That is, in this case, the pre-rendering object selection unit 301 omits supplying non-selected object metadata and non-selected object audio signals. Therefore, the conversion device 300 does not output object metadata and audio signals.

[0099] The object metadata conversion unit 302 has the same configuration as the object metadata conversion unit 201 ( FIG. 9 ) and performs the same processing. For example, the object metadata conversion unit 302 acquires speaker layout information input to the conversion device 300. The object metadata conversion unit 302 also acquires egocentric non-selected object metadata supplied from the pre-rendering object selection unit 301. Based on the speaker layout information, the object metadata conversion unit 302 converts the non-selected object metadata from an egocentric data format to an allocentric data format. The object metadata conversion unit 302 outputs the converted object metadata, i.e., the allocentric object metadata, to the outside of the conversion device 300. Since this object metadata is metadata of non-selected objects, it can also be referred to as (allocentric) non-selected object metadata. The object metadata conversion unit 302 also extracts gain information (also referred to as non-selected object gain information) from the object metadata and supplies it to the multiplication unit 303.

[0100] The object metadata conversion unit 302 may have the same configuration as the object metadata conversion unit 101 (FIG. 2) and may execute the same processing. In this case, the gain information is output as metadata, and the multiplication unit 303 is omitted. In this case, the non-selected object audio signals are output to the outside of the conversion device 300 without being subjected to any gain.

[0101] The multiplication unit 303 has the same configuration as the multiplication unit 202 ( FIG. 9 ) and executes the same processing. For example, the multiplication unit 303 acquires a non-selected object audio signal supplied from the pre-rendering object selection unit 301. The multiplication unit 303 acquires non-selected object gain information supplied from the object metadata conversion unit 302. The multiplication unit 303 multiplies the non-selected object audio signal by a gain value indicated by the non-selected object gain information, and outputs the object audio signal after the multiplication to the outside of the conversion device 300. Since this object audio signal is an audio signal of a non-selected object, it can also be called a non-selected object audio signal (in the allocentric format).

[0102] The pre-rendering processing unit 304 acquires the selected object metadata and selected object audio signal supplied from the pre-rendering object selection unit 301. The pre-rendering processing unit 304 performs pre-rendering processing using the acquired selected object metadata and selected object audio signal, and renders the audio signal into a predetermined speaker layout. In other words, the pre-rendering processing unit 304 generates audio signals for each channel corresponding to the speaker layout (also referred to as channel audio signals) and metadata (also referred to as channel metadata) including channel configuration information indicating the channel configuration, etc.

[0103] More specifically, the pre-rendering processing unit 304 performs pre-rendering processing to convert the object selected by the pre-rendering object selection unit 301 into a channel-based format in the allocentric format after conversion. For example, if the channel-based format in the allocentric format after conversion is 7.1.2ch, the pre-rendering processing unit 304 performs rendering processing for the 7.1.2ch speaker layout using VBAP.

[0104] The pre-rendering processing unit 304 outputs the channel audio signals and channel metadata generated by the above-described pre-rendering processing to the outside of the converting device 300. For example, the pre-rendering processing unit 304 converts the 7.1.2ch audio signals generated as described above into Allocentric channel audio signals, and outputs channel configuration information and the like as channel metadata to the outside of the converting device 300.

[0105] With this configuration, the converting device 300 can convert egocentric object metadata into allocentric metadata while maintaining sound localization, similar to the converting devices 100 and 200. Therefore, the converting device 300 can share data between different spatial audio formats.

[0106] Furthermore, the conversion device 300 can convert some or all of the object data in the Egocentric format into channel-based data in the Allocentric format, and convert the remaining object data into object data in the Allocentric format. In this case, the data converted into channel-based data loses its editability, but it becomes possible to maintain appropriate playback sound even for objects in positions that are not supported by the Allocentric format without correcting their positions.

[0107] <Conversion Process Flow No. 3> An example of the flow of the conversion process (information processing method) executed by the conversion device 300 will be described with reference to the flowchart of FIG.

[0108] When the conversion process starts, the pre-rendering object selection unit 301 acquires object metadata (position data) and object audio signals in the egocentric format (first data format) in step S301. That is, the pre-rendering object selection unit 301 acquires first position data of a first sound source in the first data format.

[0109] In step S302, the pre-rendering object selection unit 301 selects an object (first sound source) to be pre-rendered from the objects.

[0110] In step S303, the object metadata conversion unit 302 converts the egocentric object metadata (second position data) for the non-selected object (second sound source) to generate allocentric object metadata (third position data) (second data format). That is, the object metadata conversion unit 302 converts the second position data of the second sound source other than the first sound source in the first data format into third position data in the second data format.

[0111] In step S304, the multiplier 303 multiplies the object audio signal by the gain information for the non-selected object.

[0112] In step S305, the pre-rendering processing unit 304 performs pre-rendering processing on the selected object (first sound source) to convert it into the channel-based format in the allocentric format after conversion, and generates channel metadata and channel audio signals. That is, the pre-rendering processing unit 304 converts the audio signal of the first sound source from object-based audio to channel-based audio based on the first position data of the first sound source.

[0113] When the process of step S305 is completed, the conversion process ends.

[0114] By performing each process in this manner, the converting device 300 can convert egocentric object metadata into allocentric metadata while maintaining sound localization, similar to the converting devices 100 and 200. Therefore, the converting device 300 can share data between different spatial audio formats.

[0115] Furthermore, the conversion device 300 can convert some or all of the object data in the Egocentric format into Allocentric channel-based data (Channel-based format), and convert the remaining objects into Allocentric object data. In this case, the data converted into channel-based data loses its editability, but it becomes possible to maintain appropriate playback sound even for objects in positions that are not supported by Allocentric objects, without correcting their positions.

[0116] <Example of Application to Applications> The present technology as described above can also be applied to applications, for example. In this case, for example, instructions from a user or the like may be accepted, and settings related to conversion may be made based on the instructions. FIG. 13 is a diagram showing an example of a user interface of an application. This setting screen 410 is a user interface for accepting settings related to the data format conversion of audio data as described above, and is displayed on a monitor or the like as a GUI (Graphical User Interface). User instructions input based on this display are then accepted and used for setting.

[0117] As shown in FIG. 13, the setting screen 410 includes an input field 411, an input field 412, a conversion button 413, a speaker layout setting area 414, a pull-down menu 415, a text box 416, a check box 417, an object selection button 418, and a check box 419.

[0118] An input field 411 is an interface for specifying an input file. Here, as an example, an ADM BWF file format is used, but a file or a group of files (folder) containing PCM data and metadata may also be specified.

[0119] An input field 412 is an interface for specifying an output file. Here, an ADM BWF file format is used as an example, but a file or a group of files (folder) containing PCM data and metadata may also be specified.

[0120] It should be noted that the ADM BWF file here is a file in Broadcast Wave Format (see Non-Patent Document 6) that is composed of chunks that store metadata in accordance with the Audio definition model standard (see Non-Patent Document 5) and PCM data.

[0121] The convert button 413 is a button for executing the conversion. After specifying input / output files using the input fields 411 and 412 and setting the speaker layout using the speaker layout setting area 414, pressing the convert button 413 outputs (generates) a file containing metadata and PCM data converted from the egocentric format to the allocentric format.

[0122] The speaker layout setting area 414 indicates an area for setting the speaker layout used in allocentric rendering playback.

[0123] The pull-down menu 415 is a pull-down menu for selecting the number of speaker layout channels and the configuration. By operating this pull-down menu 415, it is possible to switch between channel configurations such as 5.1.2ch, 7.1.4ch, and 9.1.6ch.

[0124] Text box 416 is a text box for setting the direction in which each speaker is installed. For each speaker, enter the horizontal angle (Az) in the left text box and the vertical angle (El) in the right text box. Here, an example is shown in which the angle from the listening point is entered, but it can also be entered in absolute coordinate format.

[0125] A check box 417 is a check box for selecting whether or not to perform conversion and output in Channel-based format by pre-rendering.

[0126] The object selection button 418 is a button that opens a setting screen for performing conversion and output in a channel-based format by pre-rendering. Pressing this object selection button 418 displays a setting screen 420, such as that shown in FIG. 14 , allowing the user to select an object to be subjected to pre-rendering. This setting screen 420 is a user interface for selecting a track (object) to be converted and output in a channel-based format by pre-rendering. In other words, this setting screen 420 accepts the selection of an object to be converted to channel-based audio. As shown in FIG. 14 , this setting screen 420 includes a list 421 containing, for example, numbers and track names read from the metadata of the input file. The user checks an object to be subjected to pre-rendering from this list 421 and presses the "OK" button. This closes the setting screen 420 (returning to the setting screen 410 of FIG. 13 ), and the checked object is set as the selected object. This means that the selected object is selected as the first sound source to be converted to channel-based audio, and its audio signal is converted to channel-based audio. Note that, by analyzing the metadata, for example, objects that exist below a horizontal plane that includes the user's listening position may be automatically selected as targets for pre-rendering processing.

[0127] Returning to Figure 13, for example, if check box 417 is checked, pre-rendering is performed on the object set on setting screen 420, and the track (object) is converted and output in Allocentric Channel-based format.

[0128] Checkbox 419 is a checkbox for selecting whether or not to apply the gain of a track (object) that is not pre-rendered to the audio signal and output it, rather than as metadata. In other words, if this checkbox 419 is checked, the audio signal is multiplied by the gain value, as in the object metadata conversion unit 201. If this checkbox 419 is not checked, the gain information is output as metadata, as in the object metadata conversion unit 101.

[0129] By applying the present technology to an application in this way, it is possible to convert egocentric object metadata into allocentric metadata while maintaining sound localization, as in the case of the conversion device 300. Therefore, the application can share data between different spatial audio formats.

[0130] <Application Example to Audio Reproduction System> The present technology can also be applied to an information processing system, for example, in addition to an information processing device, an information processing method, a program, etc. For example, by applying the present technology to an audio reproduction system that reproduces audio data, it is possible to reproduce egocentric audio data using an allocentric reproduction device.

[0131] Fig. 15 is a block diagram showing an example of the configuration of an audio reproduction system, which is one aspect of an information processing system to which the present technology is applied. The audio reproduction system 500 shown in Fig. 15 is a system that converts egocentric audio data into allocentric audio data and reproduces the audio data.

[0132] Note that Fig. 15 shows the main devices, data flows, etc., and is not necessarily all that is shown in Fig. 15. In other words, audio playback system 500 may include devices not shown as blocks in Fig. 15, and processes and data flows not shown as arrows, etc. in Fig. 15.

[0133] As shown in FIG. 15 , the audio playback system 500 includes a metadata separation device 511, a pre-rendering object selection device 512, an object metadata conversion device 513, a pre-rendering device 514, an Allocentric audio data generation device 515, an Allocentric audio playback device 516, a speaker 517, and a headphone 518.

[0134] The metadata separation device 511 acquires egocentric object audio data (also referred to as egocentric audio data) input to the audio playback system 500. Here, the egocentric audio data is assumed to be bitstream data that has already been decoded and is in an uncompressed state. The metadata separation device 511 separates this egocentric audio data into metadata and an audio signal, and supplies them to the pre-rendering object selection device 512.

[0135] The pre-rendering object selection device 512 receives the metadata and audio signal supplied from the metadata separation device 511. That is, the pre-rendering object selection device 512 receives first position data of a first sound source in a first data format. The pre-rendering object selection device 512 has a configuration similar to that of the pre-rendering object selection unit 301 ( FIG. 11 ) and performs similar processing. That is, the pre-rendering object selection device 512 selects objects to be pre-rendered from all objects. For example, the pre-rendering object selection device 512 may automatically select objects located in an area not supported by the Allocentric method (an area below the horizontal plane including the user's listening position (El<0)), may select all objects, or may not select any objects at all. That is, the pre-rendering object selection device 512 may select an object as a first sound source to be converted into channel-based audio based on the area in which the object is located, and convert the audio signal into channel-based audio. For example, the pre-rendering object selection device 512 may select an object located in an area not supported by the second data format as a first sound source to be converted into channel-based audio. The area incompatible with the second data format may be an area below the user's position. The pre-rendering object selection device 512 may also externally accept a selection of an object to be pre-rendered. For example, the pre-rendering object selection device 512 may accept a selection of an object to be converted to channel-based audio, select the selected object as a first sound source to be converted to channel-based audio, and convert the audio signal into channel-based audio. The pre-rendering object selection device 512 supplies non-selected object metadata to the object metadata conversion device 513. The pre-rendering object selection device 512 supplies the non-selected object audio signal to the Allocentric audio data generation device 515.The pre-rendering object selector 512 provides the selected object metadata and the selected object audio signal to the pre-rendering device 514 .

[0136] The object metadata conversion device 513 has the same configuration as the object metadata conversion unit 302 ( FIG. 11 ) (i.e., the object metadata conversion unit 101 ( FIG. 2 ) or the object metadata conversion unit 201 ( FIG. 9 )) and executes the same processing. For example, the object metadata conversion device 513 acquires speaker layout information for Allocentric playback input to the audio playback system 500. That is, the object metadata conversion device 513 acquires speaker layout information for playback of audio data in the second data format. The object metadata conversion device 513 acquires non-selected object metadata supplied from the pre-rendering object selection device 512. Based on the speaker layout information, the object metadata conversion device 513 converts the non-selected object metadata into Allocentric metadata. That is, the object metadata conversion device 513 converts second position data of a second sound source other than the first sound source in the first data format into third position data in the second data format. For example, the object metadata conversion device 513 converts second position data of a second sound source in the first data format into third position data in the second data format based on speaker layout information during playback of audio data in the second data format. More specifically, the object metadata conversion device 513 calculates a gain ratio for each speaker in the second data format based on the speaker layout information, calculates a panning ratio that achieves the gain ratio, and derives the third position data based on the panning ratio. The object metadata conversion device 513 supplies the converted metadata (also referred to as converted metadata) to the Allocentric audio data generation device 515. This converted metadata can also be referred to as non-selected object metadata.

[0137] The pre-rendering device 514 has a similar configuration to the pre-rendering processing unit 304 ( FIG. 11 ) and performs similar processing. For example, the pre-rendering device 514 acquires selected object metadata and selected object audio signals supplied from the pre-rendering object selection device 512. The pre-rendering device 514 performs pre-rendering processing on them and converts them into an Allocentric Channel-based format. That is, the pre-rendering device 514 converts the audio signal of the first sound source from object-based audio to channel-based audio based on the first position data of the first sound source. The pre-rendering device 514 supplies the converted metadata and audio signals (channel metadata and channel audio signals) to the Allocentric audio data generation device 515.

[0138] The Allocentric audio data generator 515 receives converted metadata (Allocentric non-selected object metadata) supplied from the object metadata converter 513. The Allocentric audio data generator 515 receives non-selected object audio signals supplied from the pre-rendering object selector 512. The Allocentric audio data generator 515 receives channel metadata and channel audio signals supplied from the pre-rendering device 514. The Allocentric audio data generator 515 uses these to generate Allocentric audio data (also referred to as Allocentric audio data). That is, the Allocentric audio data generator 515 generates an audio bitstream based on the audio signal of the first sound source converted into channel-based audio and the audio signal of the second sound source having third position data. The Allocentric audio data generator 515 supplies the generated Allocentric audio data to the Allocentric audio playback device 516.

[0139] The Allocentric audio playback device 516 is a device that plays back Allocentric audio data. The Allocentric audio playback device 516 acquires Allocentric audio data supplied from the Allocentric audio data generation device 515. The Allocentric audio playback device 516 performs a predetermined rendering process on the Allocentric audio data and outputs an audio signal from a speaker 517 or headphones 518. In other words, the Allocentric audio playback device 516 performs playback processing of the audio signal of a first sound source that has been converted into channel-based audio and the audio signal of a second sound source that has third position data.

[0140] By doing as described above, the audio playback system 500 can convert metadata while maintaining the sense of sound localization when converting between different object-based audio formats, specifically between egocentric formats. Therefore, the audio playback system 500 enables spatial audio data to be shared between different spatial audio formats.

[0141] <4. Supplementary Notes> <Computer> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs that make up the software are installed on a computer. Here, the term "computer" includes computers built into dedicated hardware, and general-purpose personal computers, for example, that can execute various functions by installing various programs.

[0142] FIG. 16 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program.

[0143] In a computer 1900 shown in FIG. 16, a CPU (Central Processing Unit) 1901, a ROM (Read Only Memory) 1902, and a RAM (Random Access Memory) 1903 are interconnected via a bus 1904.

[0144] An input / output interface 1910 is also connected to the bus 1904. An input unit 1911, an output unit 1912, a storage unit 1913, a communication unit 1914, and a drive 1915 are connected to the input / output interface 1910.

[0145] The input unit 1911 includes, for example, a keyboard, a mouse, a microphone, a touch panel, and an input terminal. The output unit 1912 includes, for example, a display, a speaker, and an output terminal. The storage unit 1913 includes, for example, a hard disk, a RAM disk, and a non-volatile memory. The communication unit 1914 includes, for example, a network interface. The drive 1915 drives removable media 1921 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.

[0146] In a computer configured as described above, the CPU 1901 performs the above-described series of processes by, for example, loading a program stored in the storage unit 1913 into the RAM 1903 via the input / output interface 1910 and the bus 1904 and executing the program. The RAM 1903 also stores data and the like necessary for the CPU 1901 to execute various processes.

[0147] The program executed by the computer can be applied by recording it on, for example, removable media 1921 such as package media. In this case, the program can be installed in storage unit 1913 via input / output interface 1910 by attaching removable media 1921 to drive 1915.

[0148] This program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, digital satellite broadcasting, etc. In this case, the program can be received by the communication unit 1914 and installed in the storage unit 1913.

[0149] Alternatively, this program can be installed in advance in the ROM 1902 or the storage unit 1913 .

[0150] <Applicable Targets of the Present Technology> The present technology can be applied to any data format.

[0151] Furthermore, the present technology can be applied to any configuration, for example, various electronic devices.

[0152] Furthermore, for example, the present technology can be implemented as a part of a device, such as a processor (e.g., a video processor) as a system LSI (Large Scale Integration), a module (e.g., a video module) using multiple processors, a unit (e.g., a video unit) using multiple modules, or a set (e.g., a video set) in which other functions are further added to a unit. In particular, each of the units in FIG. 34 , such as the preprocessing unit, encoding unit, and file generation unit, and each of the units in FIG. 38 , such as the parse processing unit, media acquisition unit, and decoding unit, may be configured to be independent modules or processors, or one or more modules or processors may be configured to perform the processing of multiple units.

[0153] Furthermore, for example, the present technology can also be applied to a network system configured with multiple devices. For example, the present technology may be implemented as cloud computing in which multiple devices share and collaborate on processing via a network. For example, the present technology may be implemented in a cloud service that provides image (video)-related services to any terminal, such as a computer, an AV (Audio Visual) device, a portable information processing terminal, or an IoT (Internet of Things) device.

[0154] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are housed in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.

[0155] <Fields and uses to which this technology can be applied> Systems, devices, processing units, etc. to which this technology is applied can be used in any field, for example, transportation, medical care, crime prevention, agriculture, livestock farming, mining, beauty, factories, home appliances, weather, nature monitoring, etc. In addition, the uses thereof are also arbitrary.

[0156] For example, the present technology can be applied to systems and devices used to provide decorative content, etc. Furthermore, for example, the present technology can also be applied to systems and devices used for transportation, such as monitoring traffic conditions and automatic driving control. Furthermore, for example, the present technology can also be applied to systems and devices used for security. Furthermore, for example, the present technology can also be applied to systems and devices used for automatic control of machines, etc. Furthermore, for example, the present technology can also be applied to systems and devices used for agriculture and livestock farming. Furthermore, for example, the present technology can also be applied to systems and devices used to monitor natural conditions, such as volcanoes, forests, and oceans, and wildlife. Furthermore, for example, the present technology can also be applied to systems and devices used for sports.

[0157] <Others> In this specification, a "flag" refers to information for identifying multiple states, and includes not only information used to identify two states, true (1) or false (0), but also information capable of identifying three or more states. Therefore, the value that this "flag" can take may be, for example, two values, 1 / 0, or three or more values. That is, the number of bits constituting this "flag" is arbitrary, and may be one bit or multiple bits. Furthermore, identification information (including flags) can be included not only in a bitstream, but also in a bitstream that includes differential information of the identification information relative to certain reference information. Therefore, in this specification, "flag" and "identification information" encompass not only the information itself, but also differential information relative to the reference information.

[0158] Furthermore, various information (e.g., metadata) related to the coded data (bitstream) may be transmitted or recorded in any form as long as it is associated with the coded data. Here, the term "associate" means, for example, making one piece of data available (linked) when processing the other piece of data. That is, data associated with each other may be combined into one piece of data or may be stored as separate pieces of data. For example, information associated with coded data (image) may be transmitted over a transmission path separate from that of the coded data (image). Furthermore, for example, information associated with coded data (image) may be recorded on a recording medium separate from that of the coded data (image) (or on a different recording area of ​​the same recording medium). Note that this "association" may refer not to the entire data, but to only a portion of the data. For example, an image and information corresponding to that image may be associated with each other in any unit, such as multiple frames, one frame, or a portion of a frame.

[0159] In this specification, terms such as "composite," "multiplex," "add," "integrate," "include," "store," "embed," "insert," and the like refer to combining multiple items into one, such as combining encoded data and metadata into one piece of data, and refer to one method of "associating" as described above.

[0160] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present technology.

[0161] For example, a configuration described as one device (or processing unit) may be divided and configured as multiple devices (or processing units). Conversely, configurations described above as multiple devices (or processing units) may be combined and configured as one device (or processing unit). Of course, configurations other than those described above may be added to the configuration of each device (or each processing unit). Furthermore, as long as the configuration and operation of the entire system are substantially the same, part of the configuration of one device (or processing unit) may be included in the configuration of another device (or other processing unit).

[0162] Furthermore, for example, the above-described program may be executed in any device, as long as the device has the necessary functions (functional blocks, etc.) and is able to obtain the necessary information.

[0163] Also, for example, each step of a single flowchart may be executed by a single device, or may be shared and executed by multiple devices. Furthermore, when a single step includes multiple processes, the multiple processes may be executed by a single device, or may be shared and executed by multiple devices. In other words, multiple processes included in a single step can be executed as multiple step processes. Conversely, processes described as multiple steps can be executed collectively as a single step.

[0164] For example, the steps of a program executed by a computer may be executed in chronological order in the order described herein, or may be executed in parallel or individually at the required timing, such as when a call is made. In other words, as long as no contradiction occurs, the steps may be executed in an order different from the order described above. Furthermore, the steps of this program may be executed in parallel with the processing of another program, or may be executed in combination with the processing of another program.

[0165] Furthermore, for example, multiple technologies related to the present technology can be implemented independently and independently, as long as no contradiction occurs. Of course, any multiple technologies can also be implemented in combination. For example, part or all of the present technology described in any embodiment can be implemented in combination with part or all of the present technology described in another embodiment. Furthermore, part or all of any of the above-described present technologies can be implemented in combination with other technologies not described above.

[0166] Note that the present technology can also be configured as follows. (1) An information processing method comprising: acquiring first position data of a first sound source in a first data format; converting an audio signal of the first sound source from object-based audio to channel-based audio based on the first position data of the first sound source; and converting second position data of a second sound source other than the first sound source in the first data format into third position data in a second data format. (2) The information processing method described in (1), acquiring speaker layout information during playback of audio data in the second data format; and converting the second position data of the second sound source in the first data format into third position data in the second data format based on the speaker layout information. (3) The information processing method described in (2), calculating a gain ratio of each speaker in the second data format based on the speaker layout information; calculating a panning ratio that results in the gain ratio; and deriving the third position data based on the panning ratio. (4) The information processing method according to any one of (1) to (3), comprising accepting a selection of an object to be converted into channel-based audio, selecting the selected object as the first sound source to be converted into channel-based audio, and converting the audio signal into the channel-based audio. (5) The information processing method according to (4), comprising displaying a user interface for accepting the selection of the object, and accepting the selection of the object input based on the user interface. (6) The information processing method according to any one of (1) to (5), comprising selecting the object as the first sound source to be converted into channel-based audio based on an area in which the object is placed, and converting the audio signal into the channel-based audio. (7) The information processing method according to (6), comprising selecting the object placed in an area incompatible with the second data format as the first sound source to be converted into the channel-based audio.(8) The information processing method according to (7), wherein the region incompatible with the second data format is a region below a user's position. (9) The information processing method according to any one of (1) to (8), wherein a gain value of the second sound source in the first data format is reflected in the audio signal of the second sound source. (10) The information processing method according to any one of (1) to (9), wherein the first data format is an egocentric data format, and the second data format is an allocentric data format. (11) The information processing method according to any one of (1) to (10), wherein an audio bitstream is generated based on the audio signal of the first sound source converted to the channel-based audio and the audio signal of the second sound source having the third position data. (12) The information processing method according to any one of (1) to (11), wherein a playback process is performed on the audio signal of the first sound source converted to the channel-based audio and the audio signal of the second sound source having the third position data. (13) An information processing system that acquires first position data of a first sound source in a first data format, converts an audio signal of the first sound source from object-based audio to channel-based audio based on the first position data of the first sound source, and converts second position data of a second sound source other than the first sound source in the first data format into third position data in a second data format. (14) The information processing system according to (13), that acquires speaker layout information during playback of the audio data in the second data format, and converts the second position data of the second sound source in the first data format into the third position data in the second data format based on the speaker layout information.(15) The information processing system according to (14), which calculates a gain ratio for each speaker in the second data format based on the speaker layout information, calculates a panning ratio that results in the gain ratio, and derives the third position data based on the panning ratio. (16) The information processing system according to any one of (13) to (15), which accepts selection of an object to be converted to channel-based audio, selects the selected object as the first sound source to be converted to channel-based audio, and converts the audio signal to channel-based audio. (17) The information processing system according to any one of (13) to (16), which selects the object as the first sound source to be converted to channel-based audio based on an area in which the object is located, and converts the audio signal to channel-based audio. (18) The information processing system according to (17), which selects the object located in an area incompatible with the second data format as the first sound source to be converted to channel-based audio. (19) The information processing system according to (18), wherein the area incompatible with the second data format is an area below a user's position. (20) A program for causing a computer to execute a process of acquiring first position data of a first sound source in a first data format, converting an audio signal of the first sound source from object-based audio to channel-based audio based on the first position data of the first sound source, and converting second position data of a second sound source other than the first sound source in the first data format into third position data in a second data format.

[0167] REFERENCE SIGNS LIST 100 Conversion device, 101 Object metadata conversion unit, 111 Speaker layout information input unit, 112 Speaker selection unit, 113 Object metadata input unit, 114 Position information correction unit, 115 Gain ratio calculation unit, 116 Panning ratio calculation unit, 117 Position data determination unit, 118 Object metadata output unit, 200 Conversion device, 201 Object metadata conversion unit, 202 Multiplication unit, 211 Gain information extraction unit, 300 Conversion device, 301 Pre-rendering object selection unit, 302 Object metadata conversion unit, 303 Multiplication unit, 304 Pre-rendering processing unit, 500 Audio playback system, 511 Metadata separation device, 512 Pre-rendering object selection device, 513 Object metadata conversion device, 514 Pre-rendering device, 515 Allocentric audio data generation device, 516 Allocentric audio playback device, 517 speakers, 518 headphones, 1900 computers

Claims

1. An information processing method comprising: acquiring first position data of a first sound source in a first data format; converting an audio signal of the first sound source from object-based audio to channel-based audio based on the first position data of the first sound source; and converting second position data of a second sound source other than the first sound source in the first data format into third position data in a second data format.

2. The information processing method according to claim 1, further comprising: acquiring speaker layout information when playing back audio data in the second data format; and converting the second position data of the second sound source in the first data format into the third position data in the second data format based on the speaker layout information.

3. The information processing method according to claim 2, further comprising: calculating a gain ratio for each speaker in the second data format based on the speaker layout information; calculating a panning ratio that yields the gain ratio; and deriving the third position data based on the panning ratio.

4. The information processing method according to claim 1, further comprising: accepting a selection of an object to be converted into the channel-based audio; selecting the selected object as the first sound source to be converted into the channel-based audio; and converting the audio signal into the channel-based audio.

5. The information processing method according to claim 4, further comprising: displaying a user interface for accepting the selection of the object; and accepting the selection of the object input based on the user interface.

6. The information processing method according to claim 1, further comprising: selecting an object as the first sound source to be converted into the channel-based audio based on an area in which the object is located; and converting the audio signal into the channel-based audio.

7. The information processing method according to claim 6, wherein the object located in an area incompatible with the second data format is selected as the first sound source to be converted into the channel-based audio.

8. The information processing method according to claim 7, wherein the area incompatible with the second data format is an area below the user's position.

9. The information processing method according to claim 1, wherein a gain value of the second sound source in the first data format is reflected in the audio signal of the second sound source.

10. The information processing method according to claim 1, wherein the first data format is an egocentric data format, and the second data format is an allocentric data format.

11. The information processing method according to claim 1, wherein an audio bitstream is generated based on the audio signal of the first sound source converted into the channel-based audio and the audio signal of the second sound source having the third position data.

12. The information processing method according to claim 1, further comprising the step of: performing a playback process on the audio signal of the first sound source converted into the channel-based audio and the audio signal of the second sound source having the third position data.

13. An information processing system that acquires first position data of a first sound source in a first data format; converts an audio signal of the first sound source from object-based audio to channel-based audio based on the first position data of the first sound source; and converts second position data of a second sound source other than the first sound source in the first data format into third position data in a second data format.

14. The information processing system of claim 13, further comprising: acquiring speaker layout information when playing back audio data in the second data format; and converting the second position data of the second sound source in the first data format into the third position data in the second data format based on the speaker layout information.

15. The information processing system according to claim 14, further comprising: calculating a gain ratio for each speaker in the second data format based on the speaker layout information; calculating a panning ratio that yields the gain ratio; and deriving the third position data based on the panning ratio.

16. The information processing system of claim 13, further comprising: accepting a selection of an object to be converted into the channel-based audio; selecting the selected object as the first sound source to be converted into the channel-based audio; and converting the audio signal into the channel-based audio.

17. The information processing system of claim 13, wherein the object is selected as the first sound source to be converted into the channel-based audio based on an area in which the object is located, and the audio signal is converted into the channel-based audio.

18. The information processing system according to claim 17, wherein the object located in an area incompatible with the second data format is selected as the first sound source to be converted into the channel-based audio.

19. The information processing system according to claim 18, wherein the area incompatible with the second data format is an area below the user's position.

20. A program for causing a computer to execute the following processes: acquiring first position data of a first sound source in a first data format; converting the audio signal of the first sound source from object-based audio to channel-based audio based on the first position data of the first sound source; and converting second position data of a second sound source other than the first sound source in the first data format into third position data in a second data format.

Citation Information

Patent Citations

  • Acoustic signal conversion device, acoustic signal conversion method, and acoustic signal conversion program

    JP2016019041A

  • Audio Signal Rendering Method and Apparatus

    US20230179941A1

  • Information processing device and method, and program

    WO2024228269A1