Information processing device and method, and program

By grouping data into common and unique parts and modeling separate parameters, the method improves recording efficiency and accuracy of directional data restoration, addressing the degradation issues in existing technologies.

WO2025182579A1PCT designated stage Publication Date: 2025-09-04SONY GROUP CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/004704
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-29
Filing Date
2025-02-13
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Existing methods for recording directional data using a mixture model degrade the accuracy of directivity data during restoration due to collective recording of parameters, failing to capture unique shapes of individual frequency bins.

Method used

The method involves grouping data into common and unique parts, modeling and recording separate model parameters for each, allowing for more accurate restoration by distinguishing between similar and unique parts.

Benefits of technology

This approach enhances recording efficiency while maintaining high precision and flexibility in data restoration, enabling detailed shape representation and adaptability to different usage scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025004704_04092025_PF_FP_ABST
    Figure JP2025004704_04092025_PF_FP_ABST
Patent Text Reader

Abstract

The present technology relates to an information processing device and method, and a program which make it possible to obtain higher-accuracy data at the time of restoration while improving recording efficiency. This information processing device comprises: an acquisition unit that acquires matching portion data for restoring matching portions among a plurality of pieces of object data, and unique portion data of each of the plurality of pieces of object data for restoring a unique portion different from the matching portion in the object data; and a data restoration unit that restores the object data on the basis of the matching portion data and the unique portion data. The present technology can be applied to an information processing device.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, method, and program

[0001] The present technology relates to an information processing device, method, and program, and more particularly to an information processing device, method, and program that improves recording efficiency and enables more accurate data to be obtained during restoration.

[0002] When recording or transmitting various types of data, such as data relating to directivity, it is sometimes desirable to reduce the amount of data.

[0003] For example, a technique has been proposed for improving data recording efficiency by recording directional data using a vMF (von Mises Fisher) distribution or a Kent distribution, which are distribution functions on a spherical surface (see, for example, Patent Document 1).

[0004] In this technology, multiple frequency bins are grouped into bands, a mixture model representing the directivity is assigned to each band, and the parameters of the mixture model are recorded. When the directivity data is used, the directivity data for each of the multiple frequency bins belonging to that band is restored from the parameters of the mixture model for the band.

[0005] International Publication No. 2023 / 074800

[0006] However, with the above-mentioned technology, although recording efficiency can be improved by collectively recording directivity data of multiple frequency bins with similar directivity shapes using parameters of one band, degradation of the directivity data may occur. In other words, when only the band parameters are recorded, the directivity data for each frequency bin obtained by restoration may not be able to fully express the unique shape of each frequency bin.

[0007] The present technology has been made in view of such circumstances, and aims to improve recording efficiency while enabling more accurate data to be obtained during restoration.

[0008] An information processing device according to a first aspect of the present technology includes an acquisition unit that acquires common part data for restoring common parts of a plurality of target data and unique part data of each of the plurality of target data for restoring unique parts that are different from the common parts of the target data, and a data restoration unit that restores the target data based on the common part data and the unique part data.

[0009] An information processing method or program according to a first aspect of the present technology includes a step of acquiring common part data for restoring a common part of a plurality of target data and unique part data of each of the plurality of target data for restoring a unique part that is different from the common part of the target data, and restoring the target data based on the common part data and the unique part data.

[0010] In a first aspect of the present technology, common part data for restoring the common part of multiple target data and unique part data for each of the multiple target data for restoring a unique part that is different from the common part of the target data are obtained, and the target data is restored based on the common part data and the unique part data.

[0011] An information processing device according to a second aspect of the present technology includes a generation unit that generates, based on a plurality of target data, common part data for restoring the common part of the plurality of target data, and unique part data for each of the plurality of target data for restoring a unique part that is different from the common part of the target data.

[0012] An information processing method according to a second aspect of the present technology includes a step of generating, based on a plurality of target data, common portion data for restoring a common portion of the plurality of target data, and unique portion data for each of the plurality of target data for restoring a unique portion of the target data that is different from the common portion.

[0013] In a second aspect of the present technology, based on a plurality of target data, common part data for restoring the common part of the plurality of target data and unique part data of each of the plurality of target data for restoring a unique part different from the common part of the target data are generated.

[0014] FIG. 1 is a diagram illustrating data modeling. FIG. 2 is a diagram illustrating the present technology. FIG. 3 is a diagram illustrating grouping. FIG. 4 is a diagram illustrating a change in directivity shape. FIG. 5 is a diagram illustrating an example of a distribution function. FIG. 6 is a diagram illustrating modeling of data to be recorded. FIG. 7 is a diagram illustrating an individual recording method. FIG. 8 is a diagram illustrating an individual designation method. FIG. 9 is a diagram illustrating a combination designation method. FIG. 10 is a diagram illustrating an example of data to be recorded. FIG. 11 is a diagram illustrating an example of data syntax. FIG. 12 is a diagram illustrating an example of data syntax. FIG. 13 is a diagram illustrating an example of a server configuration. A flowchart illustrating encoding processing. A diagram illustrating an example of a configuration of an information processing device. A flowchart illustrating directivity data generation processing. A flowchart illustrating output audio data generation processing. A diagram illustrating limiting the range of use. A diagram illustrating an example of a computer configuration.

[0015] Hereinafter, embodiments to which the present technology is applied will be described with reference to the drawings.

[0016] First Embodiment About the Present Technology In the present technology, data having values ​​at each position on a surface such as a spherical surface, a flat surface, a curved surface, etc. is regarded as data to be recorded (hereinafter also referred to as data to be recorded). In other words, data indicating a distribution on the surface is regarded as data to be recorded.

[0017] As a specific example, consider an example in which shape data expressed on a spherical surface is represented and recorded using a mixture model consisting of a plurality of distribution functions (models).

[0018] In this case, the shape data, which is the data to be recorded, is data having a value at each position on the spherical surface, that is, data representing the distribution (shape) on the spherical surface.

[0019] For example, as shown in Figure 1, the distribution function F org The shape data D11 expressed by (x) is assumed to be modeled.

[0020] In this case, in the modeling, a desired shape, that is, the shape data D11, is expressed by superimposing a plurality of unimodal distribution functions (models).

[0021] That is, the original shape data D11 before modeling is represented by shape data D12 represented by distribution function f(x, Θ1), shape data D13 represented by distribution function f(x, Θ2), and shape data D14 represented by distribution function f(x, Θ3).

[0022] Here, weights w1, w2, and w3 are assigned to distribution functions f(x,Θ1), f(x,Θ2), and f(x,Θ3), and the distribution functions are superimposed to form a single mixture model (mixture distribution). Note that Θ1 to Θ3 in each distribution function are parameters that represent shape data.

[0023] During modeling, the mixture model is used to generate the original shape data D11 before modeling, i.e., the distribution function F org The parameters Θ1 to Θ3 and weights w1 to w3 that can best represent the shape represented by (x) are determined (estimated).

[0024] The determined parameters Θ1 to Θ3 and weights w1 to w3 are then recorded or transmitted as parameters of the mixture model representing the shape data D11.

[0025] Here, a model consisting of multiple models is also referred to as a mixed model, and parameters representing the models that make up the mixed model and the weights of the models are also referred to as model parameters. In particular, a parameter group consisting of model parameters of each model for restoring the mixed model is also simply referred to as the model parameters of the mixed model.

[0026] When the shape data D11 is used, such as when playing back content, the shape data D11 representing the original shape is restored from the model parameters. More specifically, shape data represented by a mixed model corresponding to the shape data D11 is generated (restored).

[0027] Let us consider the case where modeling using the above-described mixed model is performed collectively on a plurality of data series, that is, a plurality of pieces of shape data (data to be recorded).

[0028] Generally, each piece of shape data to be recorded or transmitted has a slightly different shape, so if you want to accurately record the shape (distribution) represented by the shape data in detail, you need to model each piece of shape data. This increases the total amount of model parameters for the multiple shape data, i.e., the amount of shape data to be recorded.

[0029] However, when focusing on a part of the shape data, it is often the case that the shapes (rough shapes) of the shape data are similar to each other in that part.

[0030] For example, when there is shape data for each frequency or time (frame) and an index indicating the frequency or time is assigned, there is often little change (fluctuation) in the overall shape between shape data with adjacent indexes, i.e., shape data with similar frequencies or times.

[0031] In this way, the shape data can be divided into portions whose shapes are similar among a plurality of pieces of shape data and portions whose shapes are unique to each piece of shape data.

[0032] Therefore, in this technology, similar parts of the data to be recorded (shape data) to other data to be recorded and unique parts are modeled and recorded separately, thereby improving recording efficiency and enabling more accurate data to be obtained during restoration.

[0033] In this technology, a data series consisting of multiple pieces of data to be recorded is grouped together and modeled.

[0034] At this time, models such as distribution functions constituting a mixture model representing the data to be recorded are classified into models that represent parts whose shapes are common among a plurality of data to be recorded (hereinafter also referred to as common parts), and models that represent parts that are unique to each data to be recorded (hereinafter also referred to as unique parts).The model parameters obtained by modeling the common parts and the unique parts are then recorded separately.

[0035] In other words, the data to be recorded is divided into a common part that is similar to other data to be recorded and a unique part that is not similar to other data to be recorded, and model parameters are recorded separately for the common part and the unique part.

[0036] Specifically, for the common parts, groups are formed of recording target data whose shapes (general shapes) in a predetermined area are similar to each other. The areas of the recording target data whose shapes are similar to each other are then treated as common parts, and model parameters are recorded for each group.

[0037] For example, a mixture model representing a group is given, and the model parameters of the mixture model are recorded as one common model parameter for the common portion of all recording target data belonging to the group.

[0038] On the other hand, for the unique part, model parameters are generated individually for each piece of data to be recorded.

[0039] In the specific part, whether or not to record or use the model parameters of the models (distribution functions) that make up the mixed model can be individually specified (defined) for each data item to be recorded. For example, whether or not to record or use the model parameters may be specified for each model that makes up the mixed model of the specific part.

[0040] An example of recording the data to be recorded will be described with reference to FIG.

[0041] In the example shown in Figure 2, J pieces of recording target data including recording target data D31-1 with index 1, recording target data D31-j with index j, and recording target data D31-J with index J are grouped together.

[0042] Hereinafter, when there is no need to particularly distinguish between record target data D31-1, record target data D31-j, and other record target data belonging to a group, the data will be simply referred to as record target data D31. In Figure 2, the distribution on the spherical surface is expanded onto a plane as the record target data D31.

[0043] In this example, similar areas in all of the recording target data D31 belonging to the group are regarded as common areas of the group.

[0044] Here, the common part is a part in each piece of recording target data D31 that has a similar shape (value distribution), such as region R11 in recording target data D31-1, region R12 in recording target data D31-j, and region R13 in recording target data D31-J. The common parts in each piece of recording target data D31 are regions that are in the same position and size.

[0045] When the record target data D31 is recorded, a representative mixture model that represents these common parts is determined, and model parameters of the mixture model are generated.

[0046] On the other hand, for example, when focusing on the data to be recorded D31-1, areas R21, R22, R23, etc. in the data to be recorded D31-1 are not similar to the same areas in other data to be recorded D31, and are therefore considered to be unique parts.

[0047] When the record target data D31 is recorded, a mixture model of data consisting of these unique parts is determined, and model parameters of the mixture model are generated.

[0048] When the record target data D31 is used, such as when playing back content, each record target data D31 is restored based on the model parameters of the common part and the unique part.

[0049] For example, for the data to be recorded D31-1, a mixed model representing the data of the common parts of the data to be recorded D31-1 is reconstructed based on the model parameters of the common parts for each group. Also, a mixed model representing the data of the unique parts of the data to be recorded D31-1 is reconstructed based on the model parameters of the unique parts.

[0050] Then, the mixture model of the common part and the mixture model of the unique part are integrated to generate one mixture model, and the data represented by the obtained mixture model is set as the recording target data D31-1 obtained by restoration.

[0051] It is to be noted that among the parts of the recording target data D31 that are different from the common parts, there may be parts that are similar to a plurality of other recording target data D31 that belong to the same group as the recording target data D31.

[0052] In such a case, a group consisting of the recording target data D31 having the similar parts may be treated as a subgroup, and model parameters of a mixed model representing the common parts of the multiple recording target data D31 belonging to the subgroup may be generated.

[0053] When one or more subgroups are formed within a group, the recording target data D31 belonging to the subgroups will be restored based on the model parameters of the common part of the group, the model parameters of the common part of the subgroups, and the model parameters of the unique part.

[0054] In this technology, a plurality of pieces of recording target data are grouped, and model parameters are recorded by dividing the pieces of data into a common part common to the group and a unique part for each piece of recording target data.

[0055] This makes it possible to reduce the overall recording volume, i.e., the overall data volume, when recording multiple pieces of recording data, while maintaining the high expressiveness and precision (quality) of the recording data during restoration. In other words, it is possible to improve recording efficiency and obtain more precise data during restoration.

[0056] In particular, the present technology can improve recording efficiency by collectively modeling the common parts and thereby generating model parameters for one common part for each group, thereby improving transmission efficiency when transmitting data to be recorded.

[0057] Furthermore, the unique parts are modeled for each piece of recording data, and model parameters are generated, so that it is possible to record detailed shapes that differ for each piece of recording data.

[0058] In other words, compared to providing a mixture model representing multiple pieces of recording target data and recording only the model parameters of that mixture model, it is possible to record the shape (distribution) indicated by each piece of recording target data in more detail, thereby suppressing data degradation due to modeling.

[0059] Furthermore, by treating the common part and the unique part as separate data series, it becomes possible to give flexibility to shape representation.

[0060] For example, multiple unique parts may be prepared in advance for one piece of recording target data, and when restoring the recording target data, the unique part to be used for restoration may be selected from the multiple unique parts depending on the situation, such as the usage scene.

[0061] As a specific example, for example, model parameters for indoor use and model parameters for outdoor use may be prepared for the characteristic part, and model parameters to be used for restoration may be selected from among these model parameters.

[0062] In this way, by switching the model parameters of the unique parts depending on the situation, etc., it is possible to add deformation to the shape indicated by the original data to be recorded during restoration.

[0063] Furthermore, when restoring the data to be recorded, it may be possible to switch whether or not to use the model parameters of each model constituting the mixed model that expresses the unique portion, depending on the usage scene, the environment such as playback, the communication status, the computing power of the device, the remaining battery power of the device, etc. In this case, whether or not to use the model parameters may be switched on a mixed model basis, or on a model basis that constitutes the mixed model.

[0064] In this way, it is possible to realize restoration of the data to be recorded taking into consideration the usage scenario, etc. For example, if it is desired to reduce the processing load on the device side, it is possible to restore the data to be recorded using only the model parameters of the common parts, without using the model parameters of the unique parts.

[0065] <Specific Examples of Recording Data> Specific examples of recording data will be described.

[0066] A specific example of the data to be recorded is directivity data (acoustic directivity data) in the frequency domain for each sound source.

[0067] In this case, directivity data indicating the directivity of a sound source is prepared for each frequency (frequency bin) of the sound output from the sound source, as shown in Fig. 3, for example. That is, the directivity data for each frequency bin prepared for one sound source is set as the data to be recorded. Here, the directivity data D51, directivity data D52, directivity data D53, etc. are set as the directivity data for each frequency bin (data to be recorded).

[0068] In such an example, since the directivity shapes between adjacent frequency bins are often similar, the entire frequency range can be divided into multiple frequency bands according to the similarity of the directivity shapes, and these frequency bands can be grouped together. In this case, for example, multiple frequency bins belong to a frequency band corresponding to one group, and for each group (frequency band), model parameters of the common part and model parameters of the unique part of each frequency bin are generated and recorded. For example, in Figure 3, directivity data D51 to D53 are grouped together.

[0069] When the data to be recorded is directional data (acoustic directional data) for each frequency bin of a sound source, the following models can be selected as models constituting a mixed model that expresses common parts and unique parts.

[0070] That is, when only the amplitude characteristics of the directivity are recorded as the directivity data, the directivity data becomes real number data, so a distribution function (model) expressed in real numbers, such as the vMF distribution or the Kent distribution, can be used.

[0071] When the directivity data includes not only the amplitude characteristics of the directivity but also the phase characteristics, the directivity data becomes complex number data.

[0072] Therefore, for example, it is conceivable to separate the complex amplitude at each position indicated by the directivity data into a real part and an imaginary part, and model the real part and the imaginary part separately. In such a case, a distribution function (model) expressed in real numbers, such as a vMF distribution or a Kent distribution, can be used to model the real part and the imaginary part.

[0073] Furthermore, the complex amplitude at each position indicated by the directional data can be modeled as a whole, without separating it into real and imaginary parts. In such cases, the directional data can be modeled using a distribution function (model) expressed in complex numbers, such as the complex Bingham distribution or the complex Watson distribution.

[0074] A specific example of directional data (acoustic directional data) indicating the directivity of a sound source to be recorded is radiation directional data indicating the radiation directivity of a sound source of a specified sound source type, such as a violin.

[0075] The radiation directivity data is data for each frequency that represents the sound pressure intensity of sound waves emitted from a sound source in each direction as seen from the sound source. In this case, grouping is performed on the radiation directivity data prepared for each frequency.

[0076] Other examples of data related to directivity that may be recorded include sound pickup directivity data that indicates the sound pickup directivity for each frequency of a microphone, a head-related transfer function (HRTF) that has a value for each frequency, etc. For example, sound pickup directivity data is data that indicates the sound pickup sensitivity for each frequency in each direction as seen from a sound source.

[0077] The data to be recorded is not limited to data in the frequency domain (for each frequency), but may also be acoustic data in the time domain (for each time period).

[0078] That is, the data to be recorded may be directional data indicating the directionality for each time interval to which a time index indicating the time interval is assigned. For example, directional data that can take different values ​​for each time index is set as the data to be recorded.

[0079] In such a case, the directivity data of adjacent time indexes can be grouped because at least a part of the directivity shape can be considered to be similar. Furthermore, the directivity data for each time index to be modeled is, for example, real number data.

[0080] Specific examples of data to be recorded for each time index include acoustic data indicating the direction of sound wave arrival at each time at a specific listening point in a specific space such as a room, a room impulse response (RIR), etc. The room transfer function can also be applied to recording for each direction.

[0081] The acoustic data indicating the direction of sound wave arrival is data indicating the direction and amplitude of sound waves arriving at a listening point in space at each time indicated by a time index, i.e., data indicating the amplitude of sound waves from each direction as seen from the listening point, etc. In this case, data for several consecutive time indexes is treated as one group.

[0082] The data to be recorded may be data in a series in the frequency direction or time direction, or data for each of a plurality of types, or data for each size or posture of an object related to a sound source.

[0083] For example, when recording acoustic directivity data for different types of sound sources together, the data to be recorded can be directional data indicating the acoustic directivity of multiple sound sources (musical instruments) that may have similar acoustic directivities, such as guitars, violins, and violas.

[0084] Since the acoustic directivity (directivity) of each instrument is thought to have both common parts and parts that are unique to each instrument, if the directivity data for each instrument is used as the data to be recorded, it is possible to expect improved recording efficiency.

[0085] Even for the same instrument, the shape of the directivity is thought to change if the environment during performance, such as the space in which the instrument is played, or the environment in which the instrument is played, is different.

[0086] As a specific example, when expressing the directionality of the sound of a musical instrument as a sound source (object), the directionality differs depending on whether the instrument is far from the ceiling of the space in which it is played or whether there is no ceiling, as shown by arrow Q11 in Figure 4, or whether the instrument is close to the ceiling, as shown by arrow Q12.

[0087] In particular, it is believed that whether or not the musical instrument (sound source) is located close to the ceiling does not make a significant difference in the common part of the directional pattern, but the unique part of the directional pattern changes.

[0088] Furthermore, the directivity shape of a musical instrument as a sound source, particularly the specific part of the directivity, is thought to change depending on the physique of the performer playing the instrument, the posture of the performer when playing, and other individual habits of the performer.

[0089] Therefore, it is possible to improve recording efficiency by preparing data to be recorded not only by type of sound source such as a musical instrument, but also by the environment surrounding the sound source, such as the space in which the sound source is placed or the performer of the instrument that serves as the sound source, and grouping the data appropriately.

[0090] <Mixed Model> A mixed model is obtained by mixing (combining) one or more models.

[0091] For example, when a distribution function is used as a model, one mixture distribution obtained by superimposing (mixing) one or more distribution functions is regarded as a mixture model.

[0092] In the following, the description will be continued assuming that a mixture distribution obtained by superposing a plurality of unimodal distribution functions is used as the mixture model.

[0093] In this case, a desired shape, that is, the shape of the common part or unique part of the data to be recorded, is expressed by superimposing a plurality of unimodal distribution functions.

[0094] Any distribution function may be used as a component of the mixture model, but for example, when the data to be recorded is data indicating a distribution (shape) on a spherical surface, it is possible to use a vMF distribution, a Kent distribution, a spherical Laplace distribution, etc. Furthermore, the distribution functions (models) that make up the mixture model may be of different types, such as a vMF distribution and a Kent distribution.

[0095] Each distribution function constituting the mixture model has a parameter, or in other words, each distribution function is expressed by a parameter.

[0096] For example, the vMF distribution, which is a distribution function, is a function that represents the distribution on the surface of a sphere, as shown in Figure 5. The vMF distribution can represent an isotropic distribution centered on the position indicated by the vector γ (mean vector) represented by the arrow in the figure.

[0097] Such a vMF distribution is expressed by the following equation (1).

[0098]

[0099] Equation (1) represents the value of the vMF distribution f(x, Θ) at position x on the surface of the sphere, i.e., at the position indicated by vector x. In equation (1), c(κ) is a normalization constant, κ represents the concentration (parameter concentration), and γ represents the vector (mean vector) that defines the center of the mean direction distribution.

[0100] Θ is a set of parameters that represent the vMF distribution f(x,Θ). Specifically, the set of parameters Θ, i.e., the parameter Θ, consists of the concentration factor κ and the vector γ. In particular, the vector γ can be expressed by the azimuth and elevation angles of a polar coordinate system with the center of the sphere as the origin.

[0101] Therefore, if the set of parameters Θ, i.e., the concentration degree κ and the azimuth and elevation angles of the vector γ, are recorded as model parameters, the vMF distribution f(x, Θ), which is a distribution function (model), can be reconstructed from the model parameters.

[0102] In the following, we will explain the case where the data to be recorded is mainly shape data representing a distribution on a spherical surface, but the data to be recorded may also be, for example, data representing a distribution (shape) on a two-dimensional plane or data representing a distribution on an arbitrary curved surface.

[0103] For example, when the data to be recorded is data that represents a distribution on a two-dimensional plane, it is conceivable to use a mixed Gaussian distribution made up of a plurality of two-dimensional Gaussian distributions as the mixture model.

[0104] A mixture model can be obtained by superimposing a plurality of distribution functions (models) as described above, and when recording a common portion of the data to be recorded, a mixture model representing the group is given as described above.

[0105] For example, the distribution function F org The mixture model representing (x) (shape data D11), i.e., the mixture model obtained by superimposing the distribution functions f(x, Θ1), f(x, Θ2), and f(x, Θ3), is assumed to be the mixture model representing the common part of the recording target data belonging to the group (hereinafter also referred to as the common mixture model).

[0106] The common mixed model may be one of the mixed models representing the common parts of each recording target data belonging to the group, or it may be one mixed model calculated from the mixed model of the common parts of each recording target data, such as the average value of those mixed models.

[0107] Alternatively, the common mixed model may be a mixed model that represents a single piece of data obtained from the common part of each recording target data, such as the average value of the common part of each recording target data belonging to a group, or the common mixed model may be a predetermined or specified one.

[0108] The common mixture model can be reconstructed using the parameters Θ1 to Θ3 of each model (distribution function) and the weights w1 to w3 of each model. These parameters and weights are considered to be the model parameters of the common mixture model, i.e., the model parameters of the common part of the groups.

[0109] In order to restore each original record target data with higher accuracy using the model parameters of the common mixture model, parameters for adjusting the difference between the common mixture model and the common part of the record target data are also required.

[0110] In this technology, a scale factor and an offset value are used as parameters for adjusting the difference between the common mixture model and the common portion of the data to be recorded.

[0111] The scale factor is a parameter related to the dynamic range of the data to be recorded, i.e., a parameter that determines the magnification of the entire mixture model, and adjusts for differences in the dynamic range of each piece of data to be recorded. For example, the scale factor is a parameter based on the ratio between the common mixture model and the common part of the data to be recorded or the entire data to be recorded.

[0112] The offset value is a parameter related to the shift amount of the data to be recorded, i.e., a parameter that determines the shift amount of the entire mixture model, and adjusts the difference in the lower limit value for each piece of data to be recorded. For example, the offset value is a parameter based on the difference in the lower limit value between the common mixture model and the common part of the data to be recorded or the entire data to be recorded.

[0113] For example, if a common mixture model has K distribution functions f(x,Θ K ) In this case, the model parameters of the common mixture model are the parameters Θ1 to Θ K and weights w1 to w K It is also assumed that a scale factor S and an offset value C are given to one piece of data to be recorded.

[0114] In such a case, the mixture distribution F rep (x) is the parameter Θ1 to Θ K and weights w1 to w K and the scale factor S and the offset value C, the following equation (2) can be used:

[0115]

[0116] The mixture distribution F obtained by this equation (2) rep (x) is the common part (mixture model of the common part) of the data to be recorded obtained by the restoration.

[0117] As described above, for each piece of recording target data belonging to a group, the model parameters of the common part, the scale factor and offset value of each piece of recording target data, and the model parameters of the unique part of each piece of recording target data are obtained as modeling results.

[0118] For example, as shown in FIG. 6, it is assumed that a total of J pieces of recording target data from the recording target data with index 1 to the recording target data with index J are grouped into one group.

[0119] In this example, as indicated by arrow Q31, a common mixture model for the common part of J pieces of recording target data is expressed by K models (distribution functions), and parameters Θ1 to Θ2 are used as model parameters of the common mixture model. K and weights w1 to w K These common part model parameters are common part data for restoring the common part of all recording target data belonging to the group.

[0120] Furthermore, for each of the J pieces of data to be recorded, a scale factor, an offset value, and model parameters of the characteristic part are obtained.

[0121] For example, in the portion indicated by the arrow Q32, for the data to be recorded with an index of 1, the parameter Θ, which is a model parameter of the specific portion, is 11 ~Θ 1K(1) and weight w 11 ~w 1K(1) , a scale factor S1 and an offset value C1 are shown.

[0122] In particular, it can be seen that for this recording target data, the mixture model of the eigenpart is composed of K(1) models (distribution functions). The model parameters of the eigenpart are eigenpart data for restoring the eigenpart that is different from the common part of the recording target data.

[0123] Similarly, for example, in the portion indicated by the arrow Q33, for the data to be recorded having the index J, the parameter Θ, which is a model parameter of the inherent portion, is J1 ~Θ JK(J) and weight w J1 ~w JK(J) and the scale factor S J and offset value C J In particular, for this recording target data, the mixture model of the eigenpart is composed of K(J) models (distribution functions).

[0124] Here, for any j-th record target data, the model parameters of the specific part are set as parameters Θ j1 ~Θ jK(j) and weight w j1 ~w jK(j) and the scale factor and offset value are written as S j and C j It will be written as follows.

[0125] In the example of Figure 6, there are K distribution functions for the common part, and for each record target data, there are K(j) distribution functions for the unique part, and the shape of each record target data is expressed by a mixture of these distribution functions.

[0126] For example, the jth data to be recorded is the common part parameters Θ1 to Θ K and weights w1 to w K and the parameter Θ of the eigenpart j1 ~Θ jK(j) and weight w j1 ~w jK(j) and the scale factor S j and offset value C j From this, it can be obtained by the following formula (3).

[0127]

[0128] The mixture distribution F obtained by such equation (3) j (x) is the jth data to be recorded obtained by the restoration. jIt can be said that (x) is obtained by overlaying (adding) a mixture model of the common part (common mixture model) and a mixture model of the unique part to generate one mixture model, and then adjusting the scale and lower limit value of that mixture model using a scale factor and offset value.

[0129] During recording, it is possible to record only some of the model parameters of the unique part for each piece of data to be recorded, or not record any model parameters of the unique part.Similarly, during restoration, it is possible to use only some of the model parameters of the unique part for restoring each piece of data to be recorded, or not use any model parameters of the unique part.

[0130] Incidentally, when recording or transmitting model parameters, etc., the model parameters of the unique part defined (or specified) for each data to be recorded can be handled using, for example, the individual recording method or the specification method shown below.

[0131] The individual recording method is a method in which model parameters of a unique part are defined for each piece of recording target data belonging to a group, and all of the model parameters of the unique part are recorded.

[0132] In contrast, the specification method involves defining a set of distribution functions (models) that may be used as unique parts for the entire group, and then specifying one or more distribution functions from among these distribution functions to be used to restore the unique parts for each piece of data to be recorded.

[0133] The individual recording method and the designation method will be described in more detail below with reference to Figures 7 to 9. Note that in Figures 7 to 9, one group is formed (configured) by J pieces of data to be recorded. Also, in Figures 7 to 9, parameters that are the same as those in Figure 6 are indicated by the same letters (symbols), etc., and their explanation will be omitted as appropriate.

[0134] FIG. 7 shows an example of the individual recording method.

[0135] In the individual recording method, model parameters of the unique parts of each piece of data to be recorded belonging to a group are defined as subparameters, and the model parameters of these unique parts are recorded.

[0136] Specifically, for example, for the first record target data whose index is 1, the parameter Θ is used as the model parameter of the unique part as shown by the arrow Q41. 11 ~Θ 1K(1) and weight w 11 ~w 1K(1) is generated and recorded.

[0137] Similarly, for the data to be recorded with index j, the parameter Θ is used as the model parameter of the inherent part, as shown by arrow Q42. j1 ~Θ jK(j) and weight w j1 ~w jK(j) is generated and recorded. For the data to be recorded with index J, the parameter Θ is generated and recorded as shown by arrow Q43. J1 ~Θ JK(J) and weight w J1 ~w JK(J) is generated and recorded.

[0138] On the other hand, possible specification methods include a method of individually specifying the distribution functions that make up the mixture model of the eigenpart (hereinafter also referred to as the individual specification method), and a method of specifying a combination of distribution functions that make up the mixture model of the eigenpart (hereinafter also referred to as the combination specification method).

[0139] FIG. 8 shows an example of an individual designation method, which is an example of the designation method.

[0140] In the individual specification method, for all J pieces of data to be recorded belonging to a group, a set of K(A) predetermined distribution functions and weights is prepared (defined) in advance as candidates (hereinafter also referred to as candidate parameters) for the model (distribution function) of the eigenpart, as shown by arrow Q51, for example.

[0141] Here, the candidate parameters Θ1 to Θ2 are used as model parameters for the distribution function (model) that constitutes the mixture model of the eigenpart. K(A) and weights w1 to wK(A) These candidate parameters can also be generated from each piece of recording target data that belongs to the group.

[0142] In the following, an arbitrary j-th candidate parameter among the K(A) candidate parameters is specifically referred to as the candidate parameter Θ j , w j The candidate parameter Θ j , w j is the parameter Θ j and weight w j This is a set of:

[0143] During recording, for each data to be recorded, for each of one or more distribution functions (models) that constitute the mixture model of the unique part of the data to be recorded, a candidate parameter representing the distribution function is selected (specified) from among K(A) candidate parameters, and an index indicating the selected candidate parameter is recorded.

[0144] In FIG. 8, for example, the part indicated by the arrow Q52 has an index i 11 ~i 1K(1) is shown.

[0145] Therefore, it can be seen that for the recording target data with an index of 1, the mixture model of the eigenpart is made up of K(1) distribution functions (models).

[0146] For example, index i 11 indicates any one of the K(A) candidate parameters such as the candidate parameters Θ3 and w3.

[0147] Similarly, for example, the part indicated by the arrow Q53 has an index i indicating each of the K(J) candidate parameters for the specific part selected for the recording target data with index J. J1 ~i JK(J) is shown.

[0148] In the individual specification method, index information indicating the distribution function (model) that constitutes the mixture model, more specifically the model parameters (candidate parameters) of the distribution function, is recorded as information for obtaining a mixture model of the unique part of the data to be recorded by reconstruction.

[0149] In this way, by recording the index instead of the model parameters themselves as information for obtaining the unique portion, it is possible to further improve the recording efficiency compared to the individual recording method.

[0150] The number of recording bits per index representing a candidate parameter depends on the total number of candidate parameters (distribution functions), K(A). For example, if the total number of candidate parameters, K(A), is 16, one index can be represented by 4 bits.

[0151] Similarly, the number of bits required to record all the indices for one mixture model representing the eigenpart depends on the number of distribution functions (models) that make up the mixture model. For example, if four candidate parameters are selected from 16 candidate parameters for one mixture model of the eigenpart and an index indicating each of the four candidate parameters is recorded, 4 × 4 = 16 bits are required to record the four indices.

[0152] FIG. 9 shows an example of a combination designation method, which is an example of a designation method.

[0153] In the combined designation method, candidate parameters are prepared in the same way as in the individual designation method. In the example of Figure 9, as shown by arrow Q61, a set of K(A) distribution functions and weights is prepared as candidate parameters Θ for all J pieces of data to be recorded belonging to a group. j , w j It is prepared as.

[0154] In the individual specification method described above, the distribution function (model) that constitutes the mixture model of the unique part is specified by an index for each piece of data to be recorded.

[0155] In contrast, in the combination specification method, a set of distribution functions (models) constituting the mixture model of the eigenpart, i.e., a plurality of combinations of one or more candidate parameters, are prepared in advance. In other words, a plurality of mixture models each consisting of a combination of one or more candidate parameters are prepared in advance.

[0156] Then, at the time of recording, for each data to be recorded, a combination indicating a set of distribution functions that constitutes a mixture model of the eigenpart is selected (specified) from among these combinations, and flag information (index information) indicating the selected combination is recorded.

[0157] In the example of FIG. 9, five combinations (patterns), from pattern 1 to pattern 5, are prepared in advance, as indicated by arrow Q62.

[0158] For example, in the pattern 1 shown in the upper right corner of the figure, the candidate parameter (distribution function) indicated by index "1" is a combination of the candidate parameter indicated by index "2" and the candidate parameter indicated by index "3".

[0159] At the time of recording, for each piece of data to be recorded, one of patterns 1 to 5 is selected as a combination of candidate parameters for obtaining a mixture model of the unique part of the data to be recorded.

[0160] For example, the portion indicated by arrow Q63 shows "Pattern 1," which is a combination (set) of candidate parameters for obtaining a mixture model of the eigenpart selected for the data to be recorded with index 1. More specifically, flag information indicating "Pattern 1" is recorded during recording.

[0161] Similarly, for example, the portion indicated by arrow Q64 shows "Pattern 3," which is a combination of candidate parameters for obtaining a mixture model of the eigenpart selected for the recording target data with index J.

[0162] In the combination specification method, a finite number of combinations of candidate parameters for obtaining a mixture model of the eigenpart are defined in advance, and flag information is set for each combination.

[0163] The number of bits of the flag information is determined by the number of combinations prepared in advance. For example, if the total number of combinations of candidate parameters is four (four patterns), the flag information can be set to two bits.

[0164] It should be noted that not only the unique parts but also the common parts may be recorded using an individual specification method or a combined specification method.

[0165] In such a case, for example, the model parameters of each model constituting the common mixture model are selected from among a plurality of candidate model parameters, or the combination of model parameters of each model constituting the common mixture model is selected from a plurality of candidate model parameter combinations. For example, when there are a large number of groups at the time of recording, it is possible to use the individual specification method or the combination specification method for the common parts as well.

[0166] The grouping of data to be recorded and the determination of common and unique parts will be explained.

[0167] In grouping, a data series consisting of multiple pieces of recording target data is divided into one or more groups based on a predetermined evaluation index. At this time, recording target data that are considered to have similar shapes are assigned to the same group.

[0168] As an evaluation index for grouping, for example, a correlation value between recording target data, which is calculated from the values ​​of each position (grid) of the original recording target data, can be used.

[0169] Specific examples include the mean absolute difference (SAD) and the mean squared difference (SSD). In this case, when the evaluation index calculated for two pieces of recording target data is equal to or less than a predetermined threshold, the two pieces of recording target data are considered to belong to the same group.

[0170] For example, when the mean absolute error (SAD) is used as an evaluation index, the predetermined recording target data F org_i(x) and other recording target data F org_j The SAD with (x) can be obtained by the following equation (4).

[0171]

[0172] If the SAD obtained by the formula (4) is equal to or less than the threshold value, the data to be recorded F org_i (x) and the data to be recorded F org_j (x) is considered to be in the same group. That is, the data to be recorded F org_i (x) and the data to be recorded F org_j (x) and (x) are considered to be data that have similar shapes (distribution of values ​​on a surface such as a sphere).

[0173] Similarly, the recording target data F org_i (x) or recording target data F org_j If we perform SAD calculation and threshold processing on (x) and other data to be recorded, for example, the data to be recorded F org_i A group is formed that is made up of data to be recorded such that the SAD between the data to be recorded is equal to or less than a threshold, such as (x).

[0174] When the mean square error (SSD) is used as the evaluation index, the predetermined recording target data F org_i (x) and other recording target data F org_j The SSD with (x) can be obtained by the following equation (5).

[0175]

[0176] In the case of the mean square error (SSD), as in the case of the mean absolute error (SAD), when the SSD is equal to or less than a threshold, the data to be recorded is considered to belong to the same group.

[0177] The common and unique parts of each recording target data belonging to a group are determined by the local distribution of the absolute error and squared error of the values ​​at each position (grid) on the spherical surface of the recording target data relative to the recording target data that represents the group (hereinafter also referred to as representative data).

[0178] For example, representative data representing a group is determined based on all of the recording target data belonging to the group. The representative data may be one piece of recording target data selected by any method from all of the recording target data, such as a method of selecting a median, or may be calculated from several pieces of recording target data, such as the average value of all of the recording target data, or may be predetermined.

[0179] Once the representative data is determined, errors such as absolute errors and squared errors at each position (grid) are calculated between the representative data and the recording target data for each of the recording target data belonging to the group. Then, based on the errors such as absolute errors and squared errors at each position obtained for all of the recording target data belonging to the group, it is determined whether those positions are to be common parts or unique parts.

[0180] For example, if the average or sum of errors at the same position in all recording data is equal to or less than a predetermined threshold, the position is considered to be a common portion because it can be said to be a portion with a similar shape among all recording data belonging to the group. On the other hand, if the average or sum of errors at the same position is greater than a predetermined threshold, the shape of each recording data is different at that position, so it is considered to be a unique portion.

[0181] In determining whether a position (grid) on a surface such as a sphere belongs to a common portion or a unique portion, the results of thresholding the error at that position may be taken into consideration, as well as the results of thresholding at positions adjacent to that position. This makes it possible to more appropriately determine the common portion and the unique portion, for example, by preventing a very small area surrounded by the common portion from being determined to be a unique portion.

[0182] In this way, once it is determined for each position (grid) whether it is a common part or a unique part, the data to be recorded is separated into common part data and unique part data, and modeling is performed for each of these data. In this case, for example, when modeling the common part, the values ​​at each position (grid) in the area on a surface such as a sphere that is not determined to be a common part, i.e., the area determined to be a unique part, are set to 0, etc. Similarly, when modeling the unique part, the values ​​at the position determined to be a common part are set to 0, etc.

[0183] Furthermore, for example, for all recording target data belonging to a group, errors such as absolute error or squared error at each position (grid) may be calculated between the recording target data, and the common part and unique part may be determined based on the calculation results.

[0184] In such a case, for example, the common part and the unique part are determined by performing threshold processing on the average value or sum of the errors for each position (grid) calculated for all combinations of two pieces of recording target data.

[0185] Furthermore, for example, all of the data to be recorded that belong to a group may be modeled first, and the common parts and unique parts may be determined based on the modeling results.

[0186] In such a case, for example, each piece of record target data is modeled using a vMF distribution, a Kent distribution, or the like, and a mixture model for each piece of record target data is generated. Then, for each model (distribution function) constituting these mixture models, the distance between a model of a given mixture model and a model of another mixture model is calculated, the similarity between the models is determined according to the distance for each model, and the common part and unique part are determined based on the determination result.

[0187] For example, models with close distribution centers can be said to have similar shapes, so when there are models with close distribution centers among the mixed models of all the data to be recorded, each of those models is said to represent a common part.

[0188] In this example, for each model (distribution function) constituting the mixed model of the recording target data belonging to a group, a determination is made as to whether that model is a common part model (a model that expresses the common part) or a unique part model. In this case, a model determined to be a common part is a model that is close to other models constituting the mixed model of the recording target data, i.e., a model whose distribution center is close, for example.

[0189] The record target data to be grouped may be not only record target data arranged in one dimension, such as the frequency direction or the time direction, but also record target data arranged in multiple dimensions.

[0190] For example, as shown in FIG. 10, assume that there is directivity data for each frequency for each sound source type, such as "guitar," "violin," and "viola."

[0191] In such a case, the directivity data for each frequency prepared for each sound source type may be used as the recording target data for grouping. Here, the recording target data is the directivity data arranged in each direction of two dimensions, the dimension of the sound source type and the dimension of the frequency.

[0192] As a result of grouping such two-dimensionally arranged directional data as data to be recorded, for example, a group consisting of data to be recorded within an area surrounded by a frame W11 is regarded as one group. In this example, one group contains directional data of adjacent frequencies for all sound source types.

[0193] In addition, for example, in addition to the sound source type and frequency shown in Figure 10, elements such as the surrounding environment and performers may be used as dimensions, and directional data arranged in three dimensions may be grouped as data to be recorded.

[0194] <Syntax Example> FIGS. 11 and 12 show examples of data formats, that is, examples of syntax, when a plurality of data to be recorded is grouped and recorded or transmitted.

[0195] In Figures 11 and 12, the directional data (acoustic directional data) for each frequency bin is the data to be recorded, and the common parts and unique parts of each piece of data to be recorded are modeled using at least one of the vMF distribution and the Kent distribution.

[0196] In particular, in FIGS. 11 and 12, one or more consecutive frequency bins are regarded as one band, and one band is regarded as one group.

[0197] FIG. 11 shows an example of syntax when recording or transmitting data to be recorded using the individual recording method.

[0198] 11, the portion of the entire data indicated by W41 on the upper side stores information for identifying which frequency bin belongs to which group (band) and information (data) for obtaining a mixture model (common mixture model) of the common parts of the groups. Hereinafter, the portion indicated by W41 will also be referred to as a common parameter recording portion W41.

[0199] 11, the portion indicated by W42 at the bottom stores information (data) for obtaining a mixture model of the inherent part of each piece of data to be recorded. Hereinafter, the portion indicated by W42 will also be referred to as the inherent parameter recording section W42.

[0200] The common parameter recording section W41 includes a band number "band_count" which is information indicating the number of groups, i.e., the number of bands corresponding to a group, obtained as a result of grouping all data to be recorded.

[0201] The common parameter recording unit W41 includes, for each band, frequency bin information "bin_range_per_band[i_band]" indicating the frequency bins included in the band, and a mixture number "mix_count[i_band]" indicating the number of models (distribution functions) that make up the common mixture model of the common part of the bands (groups).

[0202] For example, the frequency bin information "bin_range_per_band[i_band]" indicates the frequency bin with the highest frequency among the frequency bins included in the band, and it is possible to identify which frequency bins are included in the band from this frequency bin information.

[0203] Furthermore, the common parameter recording unit W41 stores model parameters and the like of each model constituting the common mixture model for each band, the number of which is the number of mixtures "mix_count[i_band]".

[0204] Specifically, for each model in the common mixture model, the model weight “weight[i_band][i_mix]”, the concentration “kappa[i_band][i_mix]”, the azimuth angle of vector γ (mean vector) “gamma1[i_band][i_mix][φ]”, the elevation angle of vector γ “gamma1[i_band][i_mix][θ]”, and the selection flag “dist_flag” are stored.

[0205] Here, the weight "weight[i_band][i_mix]" is the weight w in the above-mentioned formula (3). k The concentration degree "kappa[i_band][i_mix]", the azimuth angle "gamma1[i_band][i_mix][φ]", and the elevation angle "gamma1[i_band][i_mix][θ]" correspond to the concentration degree κ and the azimuth angle and elevation angle of the vector γ in the above-mentioned formula (1).

[0206] When the model is a vMF distribution, the parameter Θ in Equation (3) is a set of parameters consisting of the concentration “kappa[i_band][i_mix]”, the azimuth angle “gamma1[i_band][i_mix][φ]”, and the elevation angle “gamma1[i_band][i_mix][θ]”. k This becomes:

[0207] The selection flag "dist_flag" is flag information indicating whether the distribution function used as a model is a Kent distribution or a vMF distribution.

[0208] The value "1" of the selection flag "dist_flag" indicates that the model is a Kent distribution, and the value "0" of the selection flag "dist_flag" indicates that the model is a vMF distribution.

[0209] As explained with reference to the above equation (1), the vMF distribution can be expressed by the concentration degree κ and the azimuth and elevation angles of the vector γ.

[0210] In contrast, the Kent distribution has the same concentration index κ, azimuth and elevation angles of vector γ as the vMF distribution, and the same ellipticity index β, major axis vector γ 2 , and the minor axis vector γ 3 It can be expressed as follows.

[0211] Therefore, in the common parameter recording unit W41, when the value of the model selection flag "dist_flag" is "1", the ellipticity β and major axis vector γ of the model (Kent distribution) are further recorded. 2 , and the minor axis vector γ 3 The information to obtain the

[0212] That is, the ellipticity "beta[i_band][i_mix]" and the major axis vector γ 2 The azimuth angle "gamma2[i_band][i_mix][φ]" and elevation angle "gamma2[i_band][i_mix][θ]" of the 3 The azimuth angle "gamma3[i_band][i_mix][φ]" and elevation angle "gamma3[i_band][i_mix][θ]" of the major axis vector γ 2 and the minor axis vector γ 3 is expressed in terms of azimuth and elevation angles.

[0213] The inherent parameter recording section W42 stores the number of all frequency bins, that is, the number of frequency points "bin_count" indicating the number of all data to be recorded.

[0214] In addition, the inherent parameter recording unit W42 stores the scale factor "scale_factor[i_bin]", offset value "offset[i_bin]", and sub-mixing number "mix_count_bin[i_bin]", which is the number of models (distribution functions) that make up the mixture model of the inherent part, for the number of frequency points "bin_count", i.e., the number of data to be recorded (frequency bins).

[0215] The scale factor "scale_factor[i_bin]" and the offset value "offset[i_bin]" are the scale factor S in the above equation (3). j and offset value C j The sub-mix number "mix_count_bin[i_bin]" corresponds to K(j) in equation (3).

[0216] Furthermore, the inherent parameter recording unit W42 stores model parameters of each model constituting the mixture model of the inherent part for all data to be recorded, the number of which is the sub-mix number "mix_count_bin[i_bin]".

[0217] Specifically, for each model of the mixture model of the eigenpart, the model weight "weight[i_bin][i_mix]", the concentration "kappa[i_bin][i_mix]", the azimuth angle of vector γ (mean vector) "gamma1[i_bin][i_mix][φ]", the elevation angle of vector γ "gamma1[i_bin][i_mix][θ]", and the selection flag "dist_flag[i_bin][i_mix]" are stored.

[0218] Here, the weight "weight[i_bin][i_mix]" is the weight w in the above-mentioned formula (3). jkThe concentration degree "kappa[i_bin][i_mix]" indicates the concentration degree κ of the vMF distribution or Kent distribution, and the vector consisting of the azimuth angle "gamma1[i_bin][i_mix][φ]" and the elevation angle "gamma1[i_bin][i_mix][θ]" is the vector γ (mean vector) of the vMF distribution or Kent distribution.

[0219] For example, when the model is a vMF distribution, a set of parameters consisting of the concentration “kappa[i_bin][i_mix]”, the azimuth angle “gamma1[i_bin][i_mix][φ]”, and the elevation angle “gamma1[i_bin][i_mix][θ]” corresponds to the parameter Θ in the above equation (3). jk This becomes:

[0220] Note that the index [i_mix] recorded in the specific parameter recording unit W42 is an index indicating the model that constitutes the mixed model of the specific part, and is different from the index [i_mix] of the model that constitutes the common mixed model that is recorded in the common parameter recording unit W41.

[0221] The selection flag "dist_flag[i_bin][i_mix]" is flag information indicating whether the distribution function serving as a model constituting the mixture model of the eigenpart is a Kent distribution or a vMF distribution.

[0222] The value "1" of the selection flag "dist_flag[i_bin][i_mix]" indicates that the model is a Kent distribution, and the value "0" of the selection flag "dist_flag[i_bin][i_mix]" indicates that the model is a vMF distribution.

[0223] If the value of the model selection flag "dist_flag[i_bin][i_mix]" is "1", the ellipticity β and major axis vector γ of the Kent distribution as the model are also 2 , and the minor axis vector γ 3 The information for obtaining this is stored in the inherent parameter recording section W42.

[0224] That is, the ellipticity "beta[i_bin][i_mix]" and the major axis vector γ 2 The azimuth angle "gamma2[i_bin][i_mix][φ]" and elevation angle "gamma2[i_bin][i_mix][θ]" of the 3 The azimuth angle "gamma3[i_bin][i_mix][φ]" and elevation angle "gamma3[i_bin][i_mix][θ]" of the major axis vector γ are stored in the common parameter storage unit W41. 2 and the minor axis vector γ 3 is expressed in terms of azimuth and elevation angles.

[0225] FIG. 12 shows an example of syntax when recording or transmitting data to be recorded using the individual specification method.

[0226] In the example of Figure 12, the portion of the entire data indicated by W51 at the top of the figure is the same portion as the common parameter recording portion W41 shown in Figure 11 (hereinafter also referred to as common parameter recording portion W51).

[0227] The portion of the entire data indicated by W52 in the approximate center of the figure is the portion where candidate parameters are stored (hereinafter also referred to as the candidate parameter recording section W52). The portion of the entire data indicated by W53 in the lower part of the figure is the portion corresponding to the inherent parameter recording section W42 shown in Fig. 11 (hereinafter also referred to as the inherent parameter recording section W53).

[0228] The common parameter recording section W51 is the same as the common parameter recording section W41 in FIG. 11, and stores model parameters of the common parts and the like.

[0229] The candidate parameter recording section W52 stores a sub-mix number "mix_count[i_band]" indicating the number of candidate parameters prepared in advance. The sub-mix number "mix_count[i_band]" here corresponds to the number K(A) of candidate parameters in the example described with reference to FIG. 8.

[0230] The candidate parameter recording section W52 stores model parameters (candidate parameters) of candidate models (hereinafter also referred to as candidate models) that constitute the mixture model of the specific part, for the number of sub-mixtures "mix_count[i_band]".

[0231] That is, the candidate parameters stored include the model weight "weight[i_band][i_mix]", the concentration "kappa[i_band][i_mix]", the azimuth angle of vector γ (mean vector) "gamma1[i_band][i_mix][φ]", the elevation angle of vector γ "gamma1[i_band][i_mix][θ]", and the selection flag "dist_flag[i_band][i_mix]".

[0232] The selection flag "dist_flag[i_band][i_mix]" is flag information indicating whether the distribution function as a candidate model is a Kent distribution or a vMF distribution.

[0233] The value "1" of the selection flag "dist_flag[i_band][i_mix]" indicates that the candidate model is a Kent distribution, and the value "0" of the selection flag "dist_flag[i_band][i_mix]" indicates that the candidate model is a vMF distribution.

[0234] If the value of the selection flag “dist_flag[i_band][i_mix]” of the candidate model is “1”, the ellipticity “beta[i_band][i_mix]” and the major axis vector γ 2 The azimuth angle "gamma2[i_band][i_mix][φ]" and elevation angle "gamma2[i_band][i_mix][θ]" of the 3 The azimuth angle "gamma3[i_band][i_mix][φ]" and elevation angle "gamma3[i_band][i_mix][θ]" are stored.

[0235] For example, a set of parameters such as the weight “weight[i_band][i_mix]” of the candidate model and the concentration “kappa[i_band][i_mix]” is the candidate parameter Θ in the example described with reference to FIG. j , w j Corresponds to.

[0236] The inherent parameter recording section W53 stores the number of frequency points "bin_count" which indicates the total number of data to be recorded.

[0237] 11, the characteristic parameter recording unit W53 stores a scale factor "scale_factor[i_bin]", an offset value "offset[i_bin]", and a sub-mixing number "mix_count_bin[i_bin]", which is the number of models (distribution functions) constituting the mixture model of the characteristic part, for the number of frequency points "bin_count". The sub-mixing number "mix_count_bin[i_bin]" corresponds to K(1) etc. described with reference to FIG. 8.

[0238] The specific parameter recording unit W53 stores indexes "index_mix_sub[i_bin][i_mix]" for each of the data to be recorded (frequency bins) whose number is indicated by the frequency point number "bin_count", as many as the number of sub-mixes of the data to be recorded, "mix_count_bin[i_bin]".

[0239] The index "index_mix_sub[i_bin][i_mix]" is the index i 11 In particular, the index [i_mix] in the index "index_mix_sub[i_bin][i_mix]" is an index indicating a model that constitutes a mixture model of the specific part of the data to be recorded.

[0240] The index "index_mix_sub[i_bin][i_mix]" is index information indicating one of the candidate parameters stored in the candidate parameter recording unit W52. That is, the index "index_mix_sub[i_bin][i_mix]" specifies, from among the candidate parameters stored in the candidate parameter recording unit W52, a parameter to be used as a model parameter of the model that constitutes the mixture model of the inherent part of the data to be recorded.

[0241] <Server Configuration Example> FIG. 13 is a diagram showing a configuration example of an embodiment of a server to which the present technology is applied.

[0242] The server 11 shown in FIG. 13 is an information processing device such as a computer, and functions as an encoding device that distributes content.

[0243] For example, the data for playing content may include audio data for one or more objects (audio objects) and directional data that indicates the directionality, i.e., directional characteristics, of the sound source (object) for each object, more specifically, for each sound source type. Note that the data for playing content may also include video data or haptic data that correspond to the audio data.

[0244] The server 11 includes a parameter generating section 21 , a parameter encoding section 22 , an audio data encoding section 23 , and an output section 24 .

[0245] The parameter generating unit 21 is supplied with directivity data as data to be recorded for each object (sound source type).

[0246] Here, the directivity data is provided as data consisting of directional gains at each position on a sphere centered on the object, prepared for each frequency (frequency band), i.e., for each frequency bin. In this case, the directivity data is discrete data distributed on the sphere surface, having coordinates indicating a position on the sphere surface and a directional gain value at that position. Hereinafter, the positions on the sphere where the directional data has a directional gain value will also be referred to as data points.

[0247] The parameter generation unit 21 groups the supplied multiple pieces of directional data, and generates model parameters for the common parts (common part data) and model parameters for the unique parts (unique part data) based on the grouping results, and supplies them to the parameter encoding unit 22.

[0248] The parameter generating unit 21 includes a group determining unit 31 , a modeling unit 32 , and a determining unit 33 .

[0249] The group determination unit 31 groups all the directional data to be transmitted (recorded) so that each piece of directional data belongs to one or more groups.

[0250] The modeling unit 32 models the directional data and generates model parameters for the common portion and the unique portion. For example, the modeling unit 32 may model the directional data itself, or may model the common portion and the unique portion of the directional data separately. When the common portion and the unique portion are modeled separately, different models may be used for the common portion and the unique portion.

[0251] The determining unit 33 determines whether each part (each area) of the directional data is to be a common part or a unique part based on the directional data belonging to the group.

[0252] The parameter encoding unit 22 encodes the model parameters and the like supplied from the parameter generating unit 21 and supplies the resulting encoded directivity data to the output unit 24 .

[0253] The audio data encoding unit 23 encodes the audio data of each object supplied thereto, and supplies the resulting encoded audio data to the output unit 24 .

[0254] The output unit 24 multiplexes the encoded directivity data supplied from the parameter encoding unit 22 and the encoded audio data supplied from the audio data encoding unit 23 to generate a bit stream and output it.

[0255] For simplicity of explanation, an example will be described in which the encoded directivity data and the encoded audio data are output simultaneously, but the encoded directivity data and the encoded audio data may be generated separately and output at different times. Also, the encoded directivity data and the encoded audio data may be generated by different devices.

[0256] <Description of Encoding Process> Next, a description will be given of the operation of the server 11. That is, the encoding process performed by the server 11 will be described below with reference to the flowchart of FIG.

[0257] In step S11, the group determination unit 31 of the parameter generation unit 21 performs grouping (grouping) based on the supplied directivity data of the plurality of objects and on the degree of similarity of the plurality of directivity data.

[0258] For example, the group determination unit 31 calculates, for each combination of two pieces of omnidirectional data, a correlation value such as SAD or SSD between the two pieces of omnidirectional data based on the directional data, as an evaluation index. The calculation of the evaluation index may be performed for all combinations of two pieces of omnidirectional data, or may be performed only for necessary combinations.

[0259] The group determination unit 31 performs grouping so that each piece of directional data belongs to one or more groups based on the calculation result of the evaluation index for each combination. At this time, directional data of a combination whose evaluation index is equal to or less than a predetermined threshold value are made to belong to the same group.

[0260] In step S12, the parameter generating unit 21 generates model parameters for the common part and unique part of each piece of directivity data for each group.

[0261] For example, the determination unit 33 determines representative data that represents the group based on the omnidirectional data belonging to the group, and calculates errors such as absolute error and squared error at each position (grid) between the representative data and the directional data for each directional data belonging to the group.

[0262] Then, the determination unit 33 determines whether to treat the positions as common parts or unique parts based on the errors at each position obtained for the omnidirectional data belonging to the group. In this case, for example, if the average or sum of the errors at the same position of the omnidirectional data is equal to or less than a predetermined threshold, the position is treated as a common part, and the positions that are not treated as common parts are treated as unique parts.

[0263] It is also possible to calculate errors such as absolute errors or squared errors at each position (grid) between directivity data belonging to a group, and determine the common part and unique part based on the calculation results.

[0264] The modeling unit 32 performs modeling for each piece of directional data belonging to the group based on the determination results of the common part and the unique part and the directional data itself.

[0265] As an example, the modeling unit 32 extracts data from the common area in the representative data of the group, and models the representative data by representing (expressing) the common area of ​​the representative data as a mixed model consisting of one or more models (distribution functions), such as a vMF distribution or a Kent distribution.

[0266] This allows us to obtain the model parameters of a mixture model that represents the common part of the groups, i.e., the common mixture model. For example, the parameters Θ1 to Θ2 in the above-mentioned equation (3) are K and weights w1 to w K are obtained as the model parameters of the common mixture model.

[0267] In addition, the modeling unit 32 extracts data of the unique part of the directional data for each directional data belonging to the group, and models the unique part by representing the unique part of the directional data as a mixed model consisting of one or more models (distribution functions), such as a vMF distribution or a Kent distribution.

[0268] This provides model parameters for the mixture model that represent the intrinsic part of the directional data, e.g., parameter Θ in equation (3) above. j1 ~Θ jK(j) and weight w j1 ~w jK(j) are obtained as model parameters of the mixture model of the eigenpart.

[0269] Furthermore, the modeling unit 32 calculates (generates) a scale factor and an offset value for each piece of directivity data belonging to a group based on the directivity data, the model parameters of the common part, and the model parameters of the unique part. For example, if the directivity data itself is F in the above-mentioned formula (3), j (x) is the scale factor S j and offset value C j is required.

[0270] The model parameters of the common and unique parts may be specified by the creator of the content.

[0271] Furthermore, rather than performing modeling after determining the common portion and the unique portion, the common portion and the unique portion may be determined after modeling.

[0272] In such a case, for example, the modeling unit 32 models the directional data by representing the omnidirectional data belonging to the group as a mixture model consisting of one or more models (distribution functions), such as a vMF distribution, a Kent distribution, etc. Note that the model parameters of the mixture model of the directional data, i.e., the weights and model parameters of each model, may be specified by the content creator or the like.

[0273] The determination unit 33 calculates, for each model, the distance or the like between the models constituting the mixed model of different directional data as the similarity, based on the models (distribution functions) constituting the mixed model obtained for each directional data. That is, the similarity between one model constituting the mixed model of predetermined directional data and one model constituting the mixed model of another directional data is calculated.

[0274] Based on the similarity with other models obtained for each model, the determination unit 33 determines (determines) for each model that constitutes the mixed model for each directional data whether the model should be a model representing the common part or a model representing the unique part.

[0275] As an example, among the models that constitute a mixed model of specified directional data, one model of interest will be referred to as the model of interest, and among the models that constitute the mixed model of other directional data, models whose similarity to the model of interest is greater than or equal to a threshold will be referred to as similar models.

[0276] In this case, for example, for a particular directional data, if there is a similar model similar to the particular directional data in all other directional data that are different from the particular directional data, the particular directional data is considered to be a model representing the common part. Conversely, if there is no similar model similar to the particular directional data in all other directional data, the particular directional data is considered to be a model representing the unique part.

[0277] The modeling unit 32 generates (calculates) model parameters of a common mixed model representing the common part based on a model that is the common part of all the directional data belonging to the group, more specifically, model parameters related to the model. For example, one model parameter that represents the model parameters of the model of the common part of each directional data, or the average value of the model parameters of the model of the common part of each directional data, is used as the model parameter of the common part of the group. The representative model parameter may be, for example, a median value or a parameter specified by a specifying operation, etc.

[0278] Furthermore, the modeling unit 32 sets the model parameters of the model that are set as the inherent part of each piece of directivity data as the model parameters of the inherent part.

[0279] Furthermore, the modeling unit 32 calculates (generates) a scale factor and an offset value for each piece of directivity data belonging to a group, based on the directivity data, the model parameters of the common part, and the model parameters of the unique part.

[0280] As described above, a plurality of pieces of directional data belonging to a group may be further grouped to form subgroups.

[0281] In such a case, the group determination unit 31 performs the same process as in the group determination described above to group the directional data belonging to a group so that the directional data belongs to one or more subgroups as appropriate. Note that there may be directional data that does not belong to any subgroup.

[0282] The parameter generating unit 21 further groups the data of the unique parts of the directional data belonging to a group, and the groups formed as a result are treated as subgroups. When forming subgroups, as a result of modeling the directional data, model parameters of the common parts of the groups, model parameters of the common parts of the subgroups, model parameters of the unique parts of each directional data, and scale factors and offset values ​​of each directional data are obtained (generated).

[0283] When the individual recording method is adopted when transmitting (recording) directional data, the parameter generation unit 21 supplies the parameter encoding unit 22 with the model parameters of the common part of the group, the model parameters of the unique part of each directional data, and the scale factor and offset value of each directional data as the modeling results of the directional data.

[0284] When the individual specification method is adopted, the parameter generating unit 21 prepares model parameters (candidate parameters) of candidate models of the mixture model of the specific parts. For example, each candidate parameter is assigned an index that allows the candidate parameter to be uniquely identified. For example, the candidate parameters may be prepared in advance.

[0285] When the individual specification method is adopted, the parameter generation unit 21 selects, for each model constituting the mixture model of the inherent part of the directional data, one candidate model that is most similar to the model from among the candidate models. In other words, one candidate parameter that is most similar to the model parameter of the model of the inherent part is selected. The degree of similarity to the candidate model is determined by the distance between the models, for example, the proximity of the distribution centers.

[0286] The parameter generation unit 21 uses an index (hereinafter also referred to as a model index) indicating the candidate parameters of the candidate model selected for each model that constitutes the mixture model of the eigenpart as information (eigenpart data) for obtaining a model of the eigenpart by restoration.

[0287] The parameter generating unit 21 supplies the parameter encoding unit 22 with the model parameters of the common parts of the group, the model index of each model of the unique part of each directional data, and the scale factor and offset value of each directional data as the modeling results of the directional data.

[0288] In the individual specification method, candidate parameters are also supplied to the parameter encoding unit 22 as needed, and the candidate parameters are stored in the bit stream (encoded directivity data). Note that in the individual specification method, candidate parameters (candidate models) may be selected, and then the scale factor and offset value may be calculated using the candidate parameters, model parameters of the common part, directivity data, etc.

[0289] When the combination specification method is adopted, the parameter generation unit 21 prepares one or more combinations of candidate parameters (candidate models) for obtaining a mixed model, and the values ​​of flag information indicating each of these combinations are determined in advance. Hereinafter, a combination of candidate parameters (candidate models) will also be referred to as a candidate combination, and a mixed model obtained from each candidate parameter constituting a candidate combination will also be referred to as a candidate mixed model.

[0290] For each directional data, the parameter generator 21 selects, for each mixture model of the intrinsic part of the directional data, one candidate mixture model that is most similar to the mixture model from among the candidate mixture models. In other words, a candidate combination of one candidate mixture model that is most similar to the model parameters of the mixture model of the intrinsic part is selected.

[0291] The parameter generating unit 21 sets flag information having a value indicating the candidate combination selected for the mixture model of the eigenpart as information (eigenpart data) for obtaining the mixture model of the eigenpart by restoration.

[0292] The parameter generating unit 21 supplies the model parameters of the common parts of the group, flag information of the unique parts of each piece of directivity data, and the scale factor and offset value of each piece of directivity data to the parameter encoding unit 22 as the modeling results of the directivity data.

[0293] In the combination designation method, candidate parameters are also supplied to the parameter encoding unit 22 as needed, and the candidate parameters are stored in the bitstream. Note that, also in the combination designation method, after a candidate mixture model (candidate combination) is selected, the scale factor and offset value may be calculated using the candidate combination, model parameters of the common part, directivity data, etc.

[0294] Which method among the individual recording method, the individual designation method, and the combination designation method is adopted when transmitting directional data may be specified, for example, by a content creator, or may be determined in advance.

[0295] In step S13, the parameter encoding unit 22 encodes the model parameters etc. (the modeling results of the directivity data) supplied from the parameter generation unit 21 using a predetermined encoding method, and supplies the resulting encoded directivity data to the output unit 24. As a result, the data shown in Fig. 11 or 12, for example, is obtained as encoded directivity data.

[0296] The encoding method may be any encoding method, such as arithmetic coding, Huffman coding, etc. Similarly, even in the following description of encoding using a predetermined encoding method, the encoding method may be any method.

[0297] In step S 14 , the audio data encoding unit 23 encodes the audio data of each object supplied thereto, and supplies the resulting encoded audio data to the output unit 24 .

[0298] In addition, when metadata is present for the audio data of each object, the audio data encoding unit 23 or the parameter encoding unit 22 also encodes the metadata of each object (audio data) and supplies the resulting encoded metadata to the output unit 24.

[0299] For example, the metadata may include object position information (which may be expressed in xyz format or polar coordinate format, for example) indicating the absolute or relative position of the object in three-dimensional space, object direction information indicating the orientation of the object in three-dimensional space, sound source type information indicating the type of object (sound source), priority information indicating the priority of the object, spread information indicating the extent of the object, etc. For example, the priority indicated by the priority information may be a value from 0 to 7, with 7 being set to the object with the highest priority, or a predetermined number of spread vectors may be used as spread information indicating the extent of the object. Furthermore, data other than those described above may be included in the metadata.

[0300] In step S15, the output unit 24 multiplexes the encoding directivity data supplied from the parameter encoding unit 22 and the encoded audio data supplied from the audio data encoding unit 23 to generate a bit stream, and outputs the bit stream. Note that if there is encoding metadata, the encoding metadata is also stored in the bit stream.

[0301] The output unit 24 transmits the bit stream to the information processing device functioning as a client. The bit stream thus transmitted includes data such as model parameters as common part data for obtaining the common part of each group, and model parameters as specific part data for obtaining the specific part of each directivity data.

[0302] Once the bitstream is transmitted, the encoding process is complete.

[0303] In this way, the server 11 groups a plurality of pieces of directional data as data to be recorded, and outputs a bit stream including model parameters of common parts and model parameters of unique parts.

[0304] By doing so, it is possible to improve the recording efficiency and obtain more accurate data (directional data) at the time of restoration.

[0305] That is, by dividing the directional data belonging to a group into a common part and a unique part, and using one model parameter for the common part for the group, the amount of data required for recording can be reduced and recording efficiency (transmission efficiency) can be improved. Also, by recording (transmitting) model parameters for each directional data for the unique part, data close to the original directional data, i.e., directional data with higher accuracy, can be obtained during restoration.

[0306] <Configuration Example of Information Processing Device> An information processing device that functions as a client that acquires a bitstream output from the server 11 and generates output audio data for playing back the sound of content may be configured, for example, as shown in Fig. 15. The information processing device 61 shown in Fig. 15 may be, for example, a personal computer, a smartphone, a tablet, a head-mounted display, or a game device.

[0307] The information processing device 61 includes an acquisition unit 71 , a directional data decoding unit 72 , an audio data decoding unit 73 , and a rendering processing unit 74 .

[0308] The acquisition unit 71 acquires (receives) the bitstream output from the server 11 and extracts encoded directional data and encoded audio data from the bitstream. The acquisition unit 71 supplies the encoded directional data to a directional data decoding unit 72 and supplies the encoded audio data to an audio data decoding unit 73.

[0309] The directional data decoding unit 72 decodes the encoded directional data supplied from the acquisition unit 71 to restore (calculate) the directional data.

[0310] The directional data decoding unit 72 includes an unpacking unit 81 , a directional data restoration unit 82 , and a frequency interpolation processing unit 83 .

[0311] The unpacking unit 81 unpacks and decodes the encoded directional data supplied from the acquisition unit 71 to extract data such as model parameters contained in the encoded directional data and supplies it to the directional data restoration unit 82.

[0312] The directivity data restoration unit 82 restores (calculates) directivity data based on the data supplied from the unpacking unit 81 and supplies the restored data to a frequency interpolation processing unit 83 .

[0313] For example, the directivity data restoration unit 82 restores the directivity data based on the model parameters of the common parts (common part data), the model parameters of the unique parts (unique part data), the scale factor, the offset value, and the like.

[0314] The frequency interpolation processing unit 83 performs interpolation processing in the frequency direction on the directivity data supplied from the directivity data restoration unit 82, and supplies the resulting directivity data to the rendering processing unit 74. Here, the interpolation processing in the frequency direction may be, for example, linear interpolation using a straight line, or nonlinear interpolation using a spline curve, a cubic function, or the like. Furthermore, the interpolation processing in the frequency direction may be a combination of linear interpolation and nonlinear interpolation. Furthermore, interpolation may be performed using a function other than the above-mentioned straight line or function.

[0315] The audio data decoding unit 73 decodes the encoded audio data supplied from the acquisition unit 71 and supplies the resulting audio data of each object to a rendering processing unit 74 .

[0316] In addition, if the bitstream contains encoded metadata, the audio data decoding unit 73 or the unpacking unit 81 decodes the encoded metadata supplied from the acquisition unit 71 and supplies the resulting metadata to the rendering processing unit 74.

[0317] The rendering processing unit 74 generates output audio data based on the directivity data supplied from the frequency interpolation processing unit 83 and the audio data supplied from the audio data decoding unit 73. That is, the rendering processing unit 74 performs rendering processing using at least the directivity data and the audio data to generate output audio data.

[0318] The rendering processing unit 74 includes a directivity data storage unit 84 , an HRTF data storage unit 85 , a time interpolation processing unit 86 , a directivity convolution unit 87 , and an HRTF convolution unit 88 .

[0319] The directivity data storage unit 84 and the HRTF data storage unit 85 are supplied with viewpoint position information, listener direction information, object position information, and object direction information in accordance with user specifications or measurements by sensors.

[0320] For example, viewpoint position information is information that indicates the viewpoint position (listening position) in three-dimensional space of a user (listener) watching the content, and listener direction information is information that indicates the direction of the face of the user watching the content in three-dimensional space.

[0321] In addition, if the bitstream contains encoded metadata, object position information and object direction information are extracted from the metadata obtained by decoding the encoded metadata and supplied to a directionality data storage unit 84 and an HRTF data storage unit 85.

[0322] In addition, the directivity data storage unit 84 is also supplied with sound source type information obtained by extracting from metadata, and the HRTF data storage unit 85 is appropriately supplied with a user ID indicating the user viewing the content.

[0323] The directivity data storage unit 84 stores the directivity data supplied from the frequency interpolation processing unit 83. The directivity data storage unit 84 also reads out, from the stored directivity data, directivity data corresponding to the supplied viewpoint position information, listener direction information, object position information, object direction information, and sound source type information, and supplies the readout data to the time interpolation processing unit 86.

[0324] The HRTF data storage unit 85 stores, for each user identified by a user ID, HRTFs for each of a plurality of directions as seen from the user (listener).

[0325] The HRTF data storage unit 85 reads out the HRTFs stored therein that correspond to the supplied viewpoint position information, listener direction information, object position information, object direction information, and user ID, and supplies them to the HRTF convolution unit 88.

[0326] The time interpolation processing unit 86 performs interpolation processing in the time direction on the directivity data supplied from the directivity data holding unit 84, and supplies the resulting directivity data to the directivity convolution unit 87. Here, the interpolation processing in the time direction may be, for example, linear interpolation using a straight line, or nonlinear interpolation using a spline curve, a cubic function, or the like. Furthermore, the interpolation processing in the time direction may be a combination of linear interpolation and nonlinear interpolation. Furthermore, interpolation may be performed using a function other than the above-mentioned straight line or function.

[0327] The directivity convolution unit 87 convolves the audio data supplied from the audio data decoding unit 73 with the directivity data supplied from the time interpolation processing unit 86, and supplies the resulting audio data to an HRTF convolution unit 88. By convolving the directivity data, the directional characteristics of the object (sound source) are added to the audio data.

[0328] The HRTF convolution unit 88 convolves the audio data supplied from the directivity convolution unit 87, i.e., the audio data convolved with the directivity data, with the HRTF supplied from the HRTF data storage unit 85, and outputs the resulting audio data as output audio data. HRTF convolution makes it possible to obtain output audio data in which the sound of an object is localized at the position of the object as seen by the user (listener). Note that HRTF convolution may or may not be performed depending on the type and layout of the output audio data destination device. For example, HRTF convolution is performed when the output audio data destination device is a two-channel device such as earphones, headphones, hearing aids, or sound amplifiers. However, HRTF convolution may not be performed when the output audio data destination device is a speaker group with M channels (M>2). It should be noted that the present invention is not limited to the above-mentioned examples. For example, HRTF convolution may be performed when the output device is an HMD (Head Mounted Display) used for AR / VR, or may be performed depending on the positional relationship between a user walking around in a virtual space and a virtual sound source (audio object).

[0329] Here, an example will be described in which HRTF convolution processing is performed as the rendering processing of audio data. More specifically, convolution of directional data can also be considered part of the rendering processing. However, the rendering processing is not limited to this, and processing using VBAP (Vector Based Amplitude Panning), BRIR (Binaural Room Impulse Response), HOA (Higher Order Ambisonics), or the like may also be performed. The VBAP processing may be either two-dimensional VBAP or three-dimensional VBAP, or a combination of both. Note that if the output destination of the audio data is not a two-channel device (left and right channels) such as headphones, earphones, hearing aids, or sound collectors, but is, for example, an M-channel (M>2) speaker, HRTF convolution processing may not be performed.

[0330] In this case, for example, in the rendering process, the audio data of the object is convolved with the directional data, and then VBAP or the like is performed to generate output audio data. Furthermore, the HRTF convolution unit 88 may convert, for example, M-channel (M>2) audio data into two-channel audio data and output it by convolving, for example, HRTF or BRIR. In this way, when listening with a two-channel device such as headphones, earphones, hearing aids, or sound collectors, output audio data can be obtained in which the sound of the object is localized to the position of the object as seen by the user (listener), providing a more realistic stereophonic experience. It can also be said that metadata of the object (audio data) is used in the rendering process.

[0331] The directional data decoding unit 72, the audio data decoding unit 73, and the rendering processing unit 74 may be provided in different devices.

[0332] <Description of Directional Data Generation Process> Next, the operation of the information processing device 61 will be described.

[0333] First, a description will be given of the directivity data generation process performed when the information processing device 61 generates directivity data for each sound source type (object). That is, the directivity data generation process performed by the information processing device 61 will be described below with reference to the flowchart in FIG.

[0334] This directional data generation process begins when the acquisition unit 71 receives (acquires) the bitstream transmitted from the server 11 and supplies the encoded directional data extracted from the bitstream to the unpacking unit 81.

[0335] In step S81 , the unpacking unit 81 unpacks and decodes the encoding directivity data supplied from the acquisition unit 71 .

[0336] For example, the unpacking unit 81 decodes the encoded directional data using a decoding method corresponding to the encoding method used by the parameter encoding unit 22, and supplies the resulting common part model parameters and the like to the directional data restoration unit 82.

[0337] In step S 82 , the directivity data restoration unit 82 restores (generates) directivity data based on the data (information) supplied from the unpacking unit 81 , and supplies the restored directivity data to the frequency interpolation processing unit 83 .

[0338] When the individual recording method is adopted, the directional data restoration unit 82 restores the directional data by calculating the above-mentioned equation (3) based on the model parameters of the common parts of the group, the model parameters of the unique parts of the directional data, and the scale factor and offset value of the directional data.

[0339] Furthermore, when the individual specification method is adopted, the directivity data restoration unit 82 acquires, for each directivity data, candidate parameters indicated by the model indexes of each model in the specific part based on the model indexes. The candidate parameters may be included in the encoded directivity data, or may be acquired in advance from the server 11 or another device and stored in the directivity data restoration unit 82.

[0340] The directional data restoration unit 82 restores the directional data by performing a calculation similar to the above-mentioned equation (3) based on the model parameters of the common parts of the group, the candidate parameters of the unique parts of the directional data, and the scale factor and offset value of the directional data.

[0341] Furthermore, when the combination designation method is adopted, the directional data restoration unit 82 acquires, for each directional data, candidate parameters for obtaining a mixture model of the unique part based on the value of the flag information of the unique part. That is, the directional data restoration unit 82 identifies a candidate combination from the value of the flag information and acquires candidate parameters that constitute the candidate combination.

[0342] The candidate parameters that make up the candidate combinations may be included in the encoding directional data, as in the individual specification method, or may be obtained in advance from the server 11 or other devices and stored in the directional data restoration unit 82.

[0343] The directional data restoration unit 82 restores the directional data by performing a calculation similar to the above-mentioned equation (3) based on the model parameters of the common parts of the group, the candidate parameters of the unique parts of the directional data, and the scale factor and offset value of the directional data.

[0344] In addition, if the encoded directional data includes model parameters of the common parts of the subgroups, the directional data restoration unit 82 restores the directional data based on the model parameters of the common parts of the groups, the model parameters of the common parts of the subgroups, the model parameters of the unique parts, the scale factor, and the offset value.

[0345] Furthermore, when restoring the directional data, some model parameters of the common part and the unique part may not be used, some model parameters may be replaced with other model parameters, or model parameters used for the restoration may be added. Details of such deletion (non-use), addition, replacement, etc. of model parameters will be described later.

[0346] In step S83, the frequency interpolation processing unit 83 performs interpolation processing in the frequency direction on the directivity data supplied from the directivity data restoration unit 82, and supplies the directivity data obtained by the interpolation processing to the directivity data holding unit 84 to be held therein.

[0347] For example, if the audio data of an object is frequency domain data and the audio data has frequency component values ​​for each of a plurality of frequency bins, in the frequency domain interpolation process, directional data for the required frequency bins is generated (calculated) by the interpolation process so that directional data is generated for all frequency bins for which the audio data has frequency component values.

[0348] When the interpolation process in the frequency direction is performed and the directivity data is stored in the directivity data storage unit 84, the directivity data generation process ends.

[0349] In this way, the information processing device 61 generates directional data from the encoded directional data. By doing so, the decoding side can obtain more accurate directional data from a small amount of encoded directional data. In other words, it is possible to improve recording efficiency and obtain more accurate data (directional data) during restoration.

[0350] <Description of Output Audio Data Generation Process> The output audio data generation process performed by the information processing device 61 will be described with reference to the flowchart of Fig. 17. This output audio data generation process is performed at any timing after the directionality data generation process described with reference to Fig. 16 has been performed.

[0351] In step S141, the audio data decoding unit 73 decodes the encoded audio data supplied from the acquisition unit 71, and supplies the resulting audio data to the directivity convolution unit 87. For example, frequency domain audio data is obtained by the decoding.

[0352] When encoded metadata is supplied from the acquisition unit 71, the audio data decoding unit 73 decodes the encoded metadata and supplies the object position information, object direction information, and sound source type information contained in the resulting metadata to the directivity data storage unit 84 and the HRTF data storage unit 85 as appropriate.

[0353] The directivity data holding unit 84 also supplies the time interpolation processing unit 86 with directivity data according to the supplied viewpoint position information, listener direction information, object position information, object direction information, and sound source type information.

[0354] For example, the directivity data storage unit 84 identifies the relationship between the object and the user's viewpoint position (listening position) in three-dimensional space from the viewpoint position information, listener direction information, object position information, and object direction information, and identifies a data point according to the identification result. The data point here is a position on a sphere where the directivity data has a directivity gain value.

[0355] For example, if the direction from the object to the viewpoint is defined as the viewpoint direction, a position on the surface of the sphere in the direction of the viewpoint as viewed from the center of the sphere on which the data points of the directional data are arranged is identified as the target data point position. Note that there may be cases where an actual data point does not exist at the target data point position.

[0356] The directivity data storage unit 84 extracts, for each frequency bin, directivity gains at a plurality of data points in the vicinity of the identified target data point position from the directivity data of the sound source type (object) indicated by the sound source type information.

[0357] The directivity data storage unit 84 then supplies data consisting of the directivity gain for each frequency bin at the extracted multiple data points to the time interpolation processing unit 86 as directivity data corresponding to the relationship between the position and direction of the object and the user (listener).

[0358] Furthermore, the HRTF data storage unit 85 supplies an HRTF corresponding to the supplied viewpoint position information, listener direction information, object position information, object direction information, and user ID to the HRTF convolution unit 88 .

[0359] Specifically, for example, the HRTF data storage unit 85 identifies the relative direction of the object as seen by the listener (user) as the object direction based on the viewpoint position information, listener direction information, object position information, and object direction information.The HRTF data storage unit 85 then supplies the HRTF for the direction corresponding to the object direction, out of the HRTFs for each direction corresponding to the user ID, to the HRTF convolution unit 88.

[0360] In addition, the HRTF data storage unit 85 may supply parameters such as RIR (Room Impulse Response), BRIR, ITD (Interaural Time Difference), and IID (Interaural Intensity Difference) in addition to HRTFs to the HRTF convolution unit 88.

[0361] In step S142 , the time interpolation processing unit 86 performs interpolation processing in the time direction on the directivity data supplied from the directivity data holding unit 84 , and supplies the resulting directivity data to the directivity convolution unit 87 .

[0362] For example, the time interpolation processing unit 86 calculates the directional gain of each frequency bin at the target data point position by interpolation processing based on the directional gain of each frequency bin at multiple data points included in the directional data. That is, the directional gain at a new data point (target data point position) different from the original data point is calculated by interpolation processing.

[0363] The time interpolation processing unit 86 supplies data consisting of the directivity gain of each frequency bin at the target data point position to the directivity convolution unit 87 as directivity data obtained by interpolation processing in the time direction.

[0364] In step S143, the directivity convolution unit 87 convolves the audio data supplied from the audio data decoding unit 73 with the directivity data supplied from the time interpolation processing unit 86, and supplies the resulting audio data to the HRTF convolution unit 88.

[0365] In step S144, the HRTF convolution unit 88 convolves the audio data supplied from the directivity convolution unit 87 with the HRTF supplied from the HRTF data storage unit 85, and outputs the resulting output audio data. The HRTF convolution unit 88 may also convolve RIR, BRIR, ITD, IID, etc. in addition to HRTF.

[0366] In step S145, the information processing device 61 determines whether or not to end the process.

[0367] For example, if encoded audio data of a new frame is supplied from the acquisition unit 71 to the audio data decoding unit 73, it is determined in step S145 not to end the processing. On the other hand, if encoded audio data of a new frame is not supplied from the acquisition unit 71 to the audio data decoding unit 73 and output audio data of all frames of the content is generated, it is determined in step S145 to end the processing.

[0368] If it is determined in step S145 that the process is not yet finished, the process then returns to step S141, and the above-described process is repeated.

[0369] On the other hand, if it is determined in step S145 that the process is to be ended, the information processing device 61 ends the operation of each unit, and the output audio data generation process ends.

[0370] In this way, the information processing device 61 selects appropriate directivity data and HRTFs, and convolves the directivity data and HRTFs with the audio data to generate output audio data. In this way, it is possible to realize high-quality audio reproduction with a more realistic feel, taking into account the directional characteristics of the object (sound source) and the relationship between the object and the position and orientation of the listener.

[0371] <Regarding Non-Use or Addition of Model Parameters> In step S82 of the directional data generation process described with reference to FIG. 16, the directional data is restored (generated) using the model parameters of the common part and unique part.

[0372] In this case, when restoring the directional data, for at least one of the common part and the unique part, some model parameters of the model may not be used, some model parameters may be replaced with other model parameters, or other model parameters to be used for the restoration may be added.

[0373] That is, for example, the directivity data restoration unit 82 does not use a particular model parameter among the model parameters obtained from the encoded directivity data, but restores the directivity data using all model parameters other than the particular model parameter.

[0374] Furthermore, for example, the directivity data restoration unit 82 restores the directivity data using the model parameters obtained from the encoded directivity data and the specified additional model parameters.

[0375] Such deletion (non-use), addition, replacement, etc. of model parameters may be performed based on, for example, the relationship between the position and orientation of the object and the user (listener) in three-dimensional space, the relationship between the speed and acceleration of the object and the user in three-dimensional space, metadata such as object priority information and sound source type information, etc.

[0376] Additionally, model parameters may be deleted (not used), added, replaced, etc. based on, for example, the environment around the object, a user's operation to specify model parameters (models), etc., the computational capabilities (computational resources) of the information processing device 61, the device type, the remaining battery power of the device, etc. The environment around the object here refers to, for example, the type and size of the three-dimensional space in which the object is placed, the ease with which the object reflects and absorbs sound, the placement position of the object in the three-dimensional space, etc.

[0377] That is, the directional data restoration unit 82 determines the model parameters to delete, replace, or add based on at least one of the viewpoint position information, listener direction information, object position information, object direction information, metadata, the environment around the object, a specified operation by a user, etc., the computing power of the information processing device 61, and the device type of the information processing device 61.

[0378] As a specific example, for example, for a given directional data, if there is a model among the models that constitute the mixture model of the unique part that the user has specified as not to be used during restoration by a specified operation, etc., the directional data is restored without using the model parameters of that specified model.

[0379] Conversely, if there is a model to be added to the mixed model of the common or unique parts, for example, by a user specifying operation to be used during restoration, the model parameters of that model are used to restore the directional data.

[0380] Furthermore, for example, when the computational capability of the information processing device 61 is equal to or less than a predetermined threshold, the directional data may be restored without using model parameters of a predetermined specific model of a mixed model of common parts or unique parts. In this case, the restoration accuracy of the directional data can be appropriately adjusted in accordance with the computational capability of the information processing device 61.

[0381] Similarly, for example, when the priority indicated by the priority information of an object is equal to or lower than a threshold, or when the sound source type indicated by the sound source type information of an object is a specific type such as reverb, the directional data may be restored without using the model parameters of a predetermined specific model of the mixed model of the common part or unique part of the object. In this way, for example, it is possible to minimize the restoration accuracy of the directional data for objects with low priority and to increase the restoration accuracy of the directional data for objects with high priority.

[0382] Furthermore, model parameters of additional models constituting a mixed model of common or unique parts may be selected depending on, for example, whether the three-dimensional space in which the object is placed is a space with a ceiling or a space without a ceiling, as the environment surrounding the object.

[0383] In this case, for example, in a space with a ceiling, model parameters for a space with a ceiling are added to the model parameters obtained from the encoded directivity data, and the directivity data is restored. In other words, the directivity data is restored based on a plurality of model parameters consisting of the model parameters obtained from the encoded directivity data and the added model parameters.

[0384] Furthermore, in a use case such as reproduction of sound source directivity, depending on the content, it is conceivable that the user (listener) and the object (sound source) may move relative to each other in a three-dimensional space.

[0385] Furthermore, although the user can move in three-dimensional space, it is possible that the range in which the user can move is limited to a predetermined range, for example, the user can only be positioned in front of an object and cannot be positioned behind the object.

[0386] In such a case, for example, as shown in FIG. 18, the range to be used in the directivity data may be limited depending on the movable range of the user or the object OS11 (sound source).

[0387] For example, suppose there is directional data of a given object OS11 as indicated by an arrow Q101, and this directional data is represented by a mixed model consisting of three models.

[0388] In each model that makes up the mixture model, a directional vector that is the center of the distribution is defined, and in this example, vectors AR11, AR12, and AR13 are defined for each model (distribution function). These vectors AR11 to AR13 are, for example, the vector γ (mean vector) of the vMF distribution or Kent distribution as models.

[0389] With such directivity data, for example, as shown by arrow Q102, the user's movable range is limited to the range indicated by arrow MR11 relative to a predetermined object OS11 that is the sound source. In this example, the user can only move within the range in front of object OS11. In other words, the user cannot move behind object OS11.

[0390] In such a case, an effective range AC11 to be recorded or reproduced is determined based on, for example, the position of the object OS11 and the user's movable range indicated by the arrow MR11. In this example, the effective range AC11 is the surface of a sphere centered on the central position of the object OS11, that is, the area of ​​the sphere surface for which the directivity data is defined, in the direction in which the user may be located as viewed from the center of the sphere.

[0391] As indicated by arrow Q101, of the vectors AR11 to AR13 of each model, vectors AR11 and AR12 are vectors that indicate positions within the effective range AC11. In other words, models that have these vectors AR11 and AR12 represent a highly directional distribution in the direction in which the user may be present, and are therefore highly important models when reproducing the directionality of object OS11.

[0392] On the other hand, vector AR13 indicates a position outside the effective range AC11. In other words, the model having vector AR13 represents a distribution with high directivity in a direction where the user does not exist, and is a model with low importance when reproducing the directivity of object OS11.

[0393] Therefore, the model parameters of a model having a vector γ indicating a position outside the effective range AC11, i.e., a model whose distribution center is located outside the effective range AC11, may not be used to restore the directional data. This allows the directional data to be restored using only models with high importance, thereby reducing the amount of calculation during restoration and the processing time required for restoration while maintaining sufficient restoration accuracy.

[0394] For example, in the example of FIG. 18, the directivity data restoration unit 82 restores the directivity data using the model parameters of the model having the vector AR11 and the model parameters of the model having the vector AR12, without using the model having the vector AR13.

[0395] Note that the deletion (non-use), addition, replacement, etc. of model parameters of common parts and unique parts may be performed not only when restoring directional data, but also when modeling (encoding) the directional data.

[0396] In such a case, when generating model parameters for the common and unique parts, the parameter generation unit 21 of the server 11 deletes model parameters of models that are not used, replaces certain model parameters with other model parameters, or adds new model parameters.

[0397] Even in this case, the deletion (non-use), addition, replacement, etc. of model parameters is performed based on the relationship between the position and orientation of the object and the user (listener) in three-dimensional space, the relationship between the speed and acceleration of the object and the user in three-dimensional space, metadata such as object priority information and sound source type information, the environment around the object, specified operations by the content creator, etc., the computing power and device type of the information processing device 61, remaining battery power, etc.

[0398] That is, the parameter generating unit 21 determines the model parameters to be deleted, replaced, or added based on at least one of the viewpoint position information, listener direction information, object position information, object direction information, metadata, the environment around the object, a designation operation by a content creator or the like, the computing power of the information processing device 61, the remaining battery power of the information processing device 61, and the device type of the information processing device 61. In this case, necessary information such as the viewpoint position information, listener direction information, the computing power and device type of the information processing device 61 is appropriately acquired from the information processing device 61.

[0399] For example, in the example shown in Figure 18, the parameter generation unit 21 determines the effective range AC11 based on the position of the object OS11 and the user's movable range, and deletes the model parameters of the model having the vector AR13 according to the determination result.

[0400] Additionally, when modeling the common portion or unique portion of the directional data, modeling may be performed using a combination of a plurality of different methods.

[0401] In such a case, the parameter generating unit 21 generates, as common part data, data consisting of model parameters of a plurality of mutually different common mixture models for restoring the common part, for example.

[0402] Similarly, for example, the parameter generation unit 21 generates data consisting of model parameters of multiple different mixed models for restoring the unique parts, data consisting of model indices of each of the multiple mixed models, or flag information of each of the multiple mixed models as unique part data.

[0403] For example, suppose there is data to be recorded that has multiple components, such as directional data having directional gain and phase. In this case, the modeling unit 32 may model a predetermined component, such as directional gain, for the common or unique part using a model expressed by a vMF distribution or the like, and model other components, such as phase, using a model expressed by an HOA or the like.

[0404] During restoration, the directional data restoration unit 82 restores data of a specified component based on the model parameters of a model such as the vMF distribution, and restores data of other components based on the model parameters of a model such as the HOA, and combines (adds) these data to obtain the final data to be recorded.

[0405] As another example, the modeling unit 32 may model the recording target data using a predetermined method, i.e., a predetermined model, and further model the difference (residual) between the original recording target data before modeling and the recording target data after modeling (mixed model). Here, the modeled recording target data is obtained from the mixed model of the common part and the mixed model of the unique part.

[0406] In this case, the modeling method (model) for the data to be recorded and the modeling method for the residual may be the same or different. In the directional data restoration unit 82, the residual restored based on the residual model parameters is added to the data to be recorded restored using the model parameters, resulting in the final data to be recorded.

[0407] Alternatively, in the directional data restoration unit 82, data restored using other model parameters supplied from an external source may be added to the data to be recorded obtained from the encoded directional data, thereby forming the final data to be recorded.

[0408] <Example of Computer Configuration> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs constituting the software are installed on a computer. Here, the computer includes a computer built into dedicated hardware, and a general-purpose personal computer, for example, that can execute various functions by installing various programs.

[0409] FIG. 19 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program.

[0410] In the computer, a CPU (Central Processing Unit) 501 , a ROM (Read Only Memory) 502 , and a RAM (Random Access Memory) 503 are interconnected by a bus 504 .

[0411] An input / output interface 505 is further connected to the bus 504. An input unit 506, an output unit 507, a recording unit 508, a communication unit 509, and a drive 510 are connected to the input / output interface 505.

[0412] The input unit 506 includes a keyboard, a mouse, a microphone, an image sensor, etc. The output unit 507 includes a display, a speaker, etc. The recording unit 508 includes a hard disk, a non-volatile memory, etc. The communication unit 509 includes a network interface, etc. The drive 510 drives a removable recording medium 511 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.

[0413] In a computer configured as described above, the CPU 501 loads a program recorded in the recording unit 508, for example, into the RAM 503 via the input / output interface 505 and the bus 504, and executes the program, thereby performing the above-described series of processes.

[0414] The program executed by the computer (CPU 501) can be provided by being recorded on a removable recording medium 511 such as a package medium, for example. The program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.

[0415] In a computer, a program can be installed in the recording unit 508 via the input / output interface 505 by inserting a removable recording medium 511 into the drive 510. The program can also be received by the communication unit 509 via a wired or wireless transmission medium and installed in the recording unit 508. Alternatively, the program can be installed in the ROM 502 or the recording unit 508 in advance.

[0416] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.

[0417] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present technology.

[0418] For example, the present technology can be configured as a cloud computing system in which a single function is shared and processed collaboratively by a plurality of devices via a network.

[0419] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by a plurality of devices.

[0420] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.

[0421] Furthermore, the present technology can also be configured as follows.

[0422] (1) An information processing device comprising: an acquisition unit that acquires common part data for restoring a common part of a plurality of target data and unique part data for each of a plurality of target data for restoring a unique part different from the common part of the target data; and a data restoration unit that restores the target data based on the common part data and the unique part data. (2) The information processing device according to (1), wherein the common part data comprises first model parameters of one or more first models constituting a first mixed model representing the common part, obtained by modeling the target data or the common part. (3) The information processing device according to (2), wherein the unique part data is data for obtaining second model parameters of one or more second models constituting a second mixed model representing the unique part, obtained by modeling the target data or the unique part. (4) The information processing device according to (3), wherein the unique part data comprises the second model parameters of one or more second models. (5) The information processing device according to (3), wherein the unique part data comprises information indicating the second model parameters of one or more second models. (6) The information processing device according to (5), wherein the information is information indicating any one of a plurality of candidate parameters that are prepared in advance and are candidates for the second model parameter. (7) The information processing device according to (3), wherein the unique partial data is information indicating a combination of the second model parameters of one or a plurality of the second models that constitute the second mixture model. (8) The information processing device according to (7), wherein the information is information indicating a combination of one or a plurality of candidate parameters that are prepared in advance and are candidates for the second model parameter. (9) The information processing device according to any one of (3) to (8), wherein the data restoration unit restores the target data by at least one of not using some model parameters, replacing some model parameters, and adding other model parameters for at least one of the one or a plurality of the first model parameters and the one or a plurality of the second model parameters.(10) The information processing device according to (9), wherein the data restoration unit determines some model parameters not to be used in restoring the target data, some model parameters to be replaced when the target data is restored, or the other model parameters to be added, based on at least one of the spatial position of an object corresponding to the target data, the spatial position of a user listening to sound of content including sound of the object, a relationship between the object and at least one of the velocity or acceleration of the user, metadata of the object, an environment surrounding the object, a designation operation by the user, a computing power of the information processing device, a remaining battery level of the information processing device, and a device type of the information processing device. (11) The information processing device according to any one of (3) to (10), wherein at least one of the first model and the second model is a distribution function. (12) The information processing device according to (2), wherein the common part data includes each of the first model parameters of one or more of the first models constituting each of a plurality of first mixture models different from each other for obtaining the common part. (13) The information processing device according to (3), wherein the characteristic part data is data for obtaining each of the second model parameters of one or more of the second models constituting each of the plurality of second mixture models different from each other for obtaining the characteristic parts. (14) The information processing device according to any one of (1) to (13), wherein the acquisition unit further acquires a scale factor related to a dynamic range of the target data and an offset value related to a shift amount of the target data, and the data restoration unit restores the target data based on the common part data and the characteristic part data, the scale factor, and the offset value. (15) The information processing device according to any one of (1) to (14), wherein the target data is data indicating a distribution on a surface.(16) The information processing device according to any one of (1) to (15), wherein the acquisition unit acquires the common part data of a group consisting of a plurality of the target data, the common part data of a subgroup consisting of one or more of the target data belonging to the group, and the unique part data, and the data restoration unit restores the target data based on the common part data of the group, the common part data of the subgroup, and the unique part data. (17) The information processing device according to any one of (1) to (16), wherein the target data is directivity data of an object, and further comprises a rendering processing unit that performs rendering processing based on the directivity data and audio data of the object. (18) The information processing device according to (17), wherein the rendering processing includes convolution of the directivity data and processing using at least one of HRTF, VBAP, BRIR, HOA, RIR, ITD, and IID. (19) An information processing method in which an information processing device acquires common part data for restoring a common part of a plurality of target data and unique part data for each of the plurality of target data for restoring a unique part different from the common part of the target data, and restores the target data based on the common part data and the unique part data. (20) A program that causes a computer to execute a process including the steps of acquiring common part data for restoring the common part of a plurality of target data and unique part data for each of the plurality of target data for restoring a unique part different from the common part of the target data, and restoring the target data based on the common part data and the unique part data. (21) An information processing device comprising: a generation unit that generates, based on the plurality of target data, common part data for restoring the common part of the plurality of target data and unique part data for each of the plurality of target data for restoring the unique part different from the common part of the target data. (22) The information processing device described in (21), in which the generation unit groups the plurality of target data so that each of the plurality of target data belongs to one or more groups, and generates the common part data for each of the groups.(23) The information processing device according to (21) or (22), wherein the generation unit generates the common part data consisting of first model parameters of one or more first models constituting a first mixed model representing the common part by modeling the target data or the common part. (24) The information processing device according to (23), wherein the generation unit generates the unique part data for obtaining second model parameters of one or more second models constituting a second mixed model representing the unique part by modeling the target data or the unique part. (25) The information processing device according to (24), wherein the unique part data consists of the second model parameters of one or more second models. (26) The information processing device according to (24), wherein the unique part data consists of information indicating the second model parameters of one or more second models. (27) The information processing device according to (26), wherein the information is information indicating any one of a plurality of candidate parameters that are prepared in advance and are candidates for the second model parameters. (28) The information processing device according to (24), wherein the specific partial data is information indicating a combination of the second model parameters of the one or more second models constituting the second mixture model. (29) The information processing device according to (28), wherein the information is information indicating a combination of one or more candidate parameters among a plurality of candidate parameters that are prepared in advance and are candidates for the second model parameters. (30) The information processing device according to any one of (24) to (29), wherein the generation unit generates the specific partial data or the specific partial data by at least one of deleting some model parameters, replacing some model parameters, and adding other model parameters for at least one of the one or more first model parameters and one or more second model parameters.(31) The information processing device according to (30), wherein the generation unit determines the unique partial data or model parameters to be deleted, replaced, or added when generating the unique partial data based on at least one of the spatial position of an object corresponding to the target data, the spatial position of a user listening to sound of content including sound of the object, a relationship between the object and at least one of the speed or acceleration of the user, metadata of the object, an environment surrounding the object, a predetermined designation operation, a computing capability of a device that plays back the content, a remaining battery level of the device, and a device type of the device. (32) The information processing device according to any one of (24) to (31), wherein at least one of the first model and the second model is a distribution function. (33) The information processing device according to (23), wherein the common part data consists of each of the first model parameters of one or more of the first models that constitute each of a plurality of first mixture models that are different from each other for obtaining the common part. (34) The information processing device according to (24), wherein the characteristic part data is data for obtaining each of the second model parameters of one or more of the second models constituting each of the plurality of second mixture models different from each other for obtaining the characteristic part. (35) The information processing device according to any one of (21) to (34), wherein the generation unit further generates, for each of the plurality of target data, a scale factor related to a dynamic range of the target data and an offset value related to a shift amount of the target data based on the plurality of target data. (36) The information processing device according to any one of (21) to (35), wherein the target data is data indicating a distribution on a surface. (37) The information processing device according to (22), wherein the generation unit groups the plurality of target data belonging to the group into one or more subgroups and further generates the common part data for each of the subgroups.(38) An information processing method, in which an information processing device generates, based on a plurality of target data, common part data for restoring a common part of the plurality of target data and unique part data of each of the plurality of target data for restoring a unique part different from the common part of the target data. (39) A program that causes a computer to execute processing including the step of generating, based on a plurality of target data, common part data for restoring a common part of the plurality of target data and unique part data of each of the plurality of target data for restoring a unique part different from the common part of the target data.

[0423] REFERENCE SIGNS LIST 11 Server 21 Parameter generation unit 22 Parameter encoding unit 23 Audio data encoding unit 24 Output unit 61 Information processing device 71 Acquisition unit 72 Directional data decoding unit 73 Audio data decoding unit 74 Rendering processing unit 82 Directional data restoration unit

Claims

1. An information processing device comprising: an acquisition unit that acquires common part data for restoring common parts of multiple target data and unique part data of each of the multiple target data for restoring unique parts that differ from the common parts of the target data; and a data restoration unit that restores the target data based on the common part data and the unique part data.

2. The information processing device according to claim 1, wherein the common part data consists of first model parameters of one or more first models constituting a first mixture model representing the common part, obtained by modeling the target data or the common part.

3. The information processing device described in claim 2, wherein the specific part data is data for obtaining second model parameters of one or more second models constituting a second mixed model representing the specific part, obtained by modeling the target data or the specific part.

4. The information processing device according to claim 3, wherein the specific partial data comprises the second model parameters of one or more of the second models.

5. The information processing device according to claim 3, wherein the specific partial data comprises information indicating the second model parameters of one or more of the second models.

6. The information processing device according to claim 5, wherein the information indicates any one of a plurality of candidate parameters prepared in advance as candidates for the second model parameter.

7. The information processing device according to claim 3, wherein the specific partial data comprises information indicating a combination of the second model parameters of one or more of the second models that constitute the second mixture model.

8. The information processing device according to claim 7, wherein the information indicates one or a combination of a plurality of candidate parameters that are prepared in advance and serve as candidates for the second model parameter.

9. The information processing device of claim 3, wherein the data restoration unit restores the target data by at least one of not using some of the one or more first model parameters and one or more of the second model parameters, replacing some of the model parameters, or adding other model parameters.

10. The information processing device according to claim 9, wherein the data restoration unit determines some model parameters not to be used in restoring the target data, some model parameters to be replaced when the target data is restored, or the other model parameters to be added based on at least one of the spatial position of an object corresponding to the target data, the spatial position of a user listening to the sound of content including the sound of the object, the relationship between the object and at least one of the speed or acceleration of the user, metadata of the object, the environment around the object, a specified operation by the user, the computing power of the information processing device, the remaining battery power of the information processing device, and the device type of the information processing device.

11. The information processing device according to claim 3, wherein at least one of the first model and the second model is a distribution function.

12. The information processing device according to claim 2, wherein the common part data comprises each of the first model parameters of one or more of the first models constituting each of the plurality of first mixture models that are different from one another for obtaining the common part.

13. The information processing device described in claim 3, wherein the characteristic part data is data for obtaining each of the second model parameters of one or more of the second models that constitute each of the multiple second mixed models that are different from each other for obtaining the characteristic part.

14. The information processing device of claim 1, wherein the acquisition unit further acquires a scale factor related to the dynamic range of the target data and an offset value related to the shift amount of the target data, and the data restoration unit restores the target data based on the common part data and the unique part data, the scale factor and the offset value.

15. The information processing device according to claim 1, wherein the target data is data indicating a distribution on a surface.

16. The information processing device of claim 1, wherein the acquisition unit acquires the common part data of a group consisting of a plurality of the target data, the common part data of a subgroup consisting of one or more of the target data belonging to the group, and the unique part data, and the data restoration unit restores the target data based on the common part data of the group, the common part data of the subgroup, and the unique part data.

17. The information processing device according to claim 1, wherein the target data is directional data of an object, and further comprising a rendering processing unit that performs rendering processing based on the directional data and audio data of the object.

18. The information processing device according to claim 17, wherein the rendering process includes convolution of the directional data and processing using at least one of HRTF, VBAP, BRIR, HOA, RIR, ITD, and IID.

19. An information processing method in which an information processing device acquires common part data for restoring the common part of multiple target data and unique part data for each of the multiple target data for restoring a unique part different from the common part of the target data, and restores the target data based on the common part data and the unique part data.

20. A program that causes a computer to execute a process including the steps of acquiring common part data for restoring the common part of multiple target data and unique part data for each of the multiple target data for restoring unique parts that differ from the common part of the target data, and restoring the target data based on the common part data and the unique part data.

21. An information processing device comprising a generation unit that generates, based on a plurality of target data, common part data for restoring the common part of the plurality of target data, and unique part data of each of the plurality of target data for restoring the unique part different from the common part of the target data.

22. An information processing method in which an information processing device generates, based on multiple target data, common part data for restoring the common part of the multiple target data, and unique part data for each of the multiple target data for restoring a unique part different from the common part of the target data.

Citation Information

Patent Citations

  • Object Clustering for Rendering Object-Based Audio Content Based on Perceptual Criteria

    JP2016509249A

  • Method for generating filter for audio signal and parameterizing device therefor

    US20160275956A1

  • Information processing device, method, and program

    WO2023074009A1

  • Apparatus, method or computer program for synthesizing a spatially extended sound source using modification data on a potentially modifying object

    WO2023083753A1