Signal processing apparatus, method and storage medium
Patent Information
- Application Number
- CN202311448231.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-10-20
- Filing Date
- 2018-10-05
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2038-10-05
AI Technical Summary
然而,尽管这种方法对于能够以非常长的计算时间产生的运动图像内容(诸如电影)是有效的,但是在实时再现音频对象的情况下使用这种方法是困难的
[0027] According to one aspect of this technology, coding efficiency can be improved.
Smart Images

Figure CN117475983B_ABST
Abstract
Description
[0001] This application is a divisional application of the Chinese national phase application, filed on October 5, 2018, with international application number PCT / JP2018 / 037330, entitled "Signal Processing Apparatus, Method and Procedure". The Chinese national phase application entered the national phase on March 30, 2020, with application number 201880063759.0, entitled "Signal Processing Apparatus, Method and Procedure". Technical Field
[0002] This technology relates to signal processing apparatus, methods, and procedures, and more specifically, to signal processing apparatus, methods, and procedures capable of improving coding efficiency. Background Technology
[0003] Object audio technology has traditionally been used in movies, games, and the like, and encoding methods capable of processing object audio have been developed. Specifically, for example, the MPEG (Moving Picture Experts Group)-H Part 3: 3D Audio standard, which is known as an international standard, is an example (see, for example, Non-Patent Literature 1).
[0004] In this encoding method, similar to traditional two-channel stereo sound methods and multi-channel stereo sound methods such as 5.1 channels, moving sound sources are treated as independent audio objects, and the position information of the objects can be encoded as metadata along with the signal data of the audio objects.
[0005] This arrangement allows for playback in various viewing / listening environments with varying numbers of speakers. Furthermore, it makes it easy to process the sound of specific sound sources during playback, such as adjusting the volume of a particular sound source and adding effects, which is difficult in traditional encoding methods.
[0006] For example, in the standard of Non-Patent Document 1, a method called Amplitude Translation Based on Three-Dimensional Vectors (VBAP) (hereinafter referred to as VBAP) is used for rendering processing.
[0007] This is one of the reproduction methods commonly known as panning, and it is a method of reproduction performed by distributing gain among the three speakers that are closest to the audio object existing on the sphere, which are also located on the sphere with the viewing / listening position as the origin.
[0008] This rendering of audio objects via panning is based on the premise that all audio objects lie on a sphere with the viewing / listening position as the origin. Therefore, distance sensing when an audio object is near or far from the viewing / listening position is controlled solely by the gain of the audio object.
[0009] However, in reality, without considering factors such as the different attenuation rates depending on the frequency components and reflections in the space where the audio object exists, the expression of distance is far removed from actual experience.
[0010] To reflect this effect in the listening experience, one might first think of physically calculating reflections and attenuation in space to obtain the final output audio signal. However, while this method is effective for moving image content (such as movies) that can be generated in very long computation times, it is difficult to use in the case of reproducing audio objects in real time.
[0011] Furthermore, the final output obtained through reflection and attenuation in the physical computation space is difficult to reflect the content creator's intent. This is especially true for musical works such as music clips, which require formats that easily reflect the content creator's intent, such as applying preferred reverb processing to the audio tracks.
[0012] List of citations
[0013] Non-patent literature
[0014] Non-Patent Document 1: International Standard ISO / IEC 23008-3, First Edition, 2015-10-15 Information Technology—Efficient Coding and Media Delivery in Heterogeneous Environments—Part 3: 3D Audio Summary of the Invention
[0015] The problem to be solved by the present invention
[0016] Therefore, in real-time reproduction, it is desirable to store data such as the coefficients required for reverberation processing that take into account the reflection and attenuation of each audio object in space, as well as the location information of the audio objects, in a file or a transmission stream, and obtain the final output audio signal by using them.
[0017] However, storing the reverb processing data required for each audio object in a file or send stream for each frame increases the send rate and requires data transmission with high coding efficiency.
[0018] This technique was developed in light of this situation, and it aims to improve coding efficiency.
[0019] Problem Solving Methods
[0020] A signal processing apparatus according to one aspect of the present technology includes: an acquisition unit that acquires reverberation information and an audio object signal of an audio object, the reverberation information including at least one of spatial reverberation information specific to the space surrounding the audio object or object reverberation information specific to the audio object; and a reverberation processing unit that generates a signal of the reverberation component of the audio object based on the reverberation information and the audio object signal.
[0021] A signal processing apparatus according to one aspect of the present technology includes: an acquisition unit that acquires reverberation information and an audio object signal of an audio object, the reverberation information including at least one of spatial reverberation information specific to the space surrounding the audio object or object reverberation information specific to the audio object; a reverberation processing unit that generates a signal of a reverberation component of the audio object based on the reverberation information and the audio object signal; and a rendering unit that performs rendering processing on the reverberation component of the audio object and the audio object signal to generate an output audio signal.
[0022] A signal processing method or procedure according to one aspect of the present technology includes the following steps: acquiring reverberation information, the reverberation information including at least one of spatial reverberation information specific to the space surrounding an audio object or object reverberation information specific to the audio object and an audio object signal of the audio object; and generating a signal of the reverberation component of the audio object based on the reverberation information and the audio object signal.
[0023] A signal processing method according to one aspect of the present technology includes the following steps: acquiring reverberation information and an audio object signal of an audio object by a signal processing device, the reverberation information including at least one of the following: spatial reverberation information specific to the space surrounding the audio object and object reverberation information specific to the audio object; generating a signal of the reverberation component of the audio object by the signal processing device based on the reverberation information and the audio object signal; and performing rendering processing on the reverberation component of the audio object and the audio object signal by the signal processing device to generate an output audio signal.
[0024] According to one aspect of the present technology, a computer-readable storage medium having instructions stored thereon, when executed by a computer, causes the computer to perform a process comprising the following steps: acquiring reverberation information and an audio object signal of an audio object, the reverberation information including at least one of: spatial reverberation information specific to the space surrounding the audio object and object reverberation information specific to the audio object; generating a signal of a reverberation component of the audio object based on the reverberation information and the audio object signal; and performing rendering processing of the reverberation component of the audio object and the audio object signal to generate an output audio signal.
[0025] In one aspect of this technology, reverberation information is acquired, which includes at least one of spatial reverberation information specific to the space surrounding the audio object or object reverberation information specific to the audio object and an audio object signal of the audio object, and a signal of the reverberation component of the audio object is generated based on the reverberation information and the audio object signal.
[0026] Effects of the present invention
[0027] According to one aspect of this technology, coding efficiency can be improved.
[0028] Note that the effects described herein are not limited and can be any effect described in this disclosure. Attached Figure Description
[0029] [ Figure 1 [Illustration 1] is a diagram showing an example configuration of a signal processing device.
[0030] [ Figure 2 [Illustration 1] is a diagram showing an example of the configuration of a rendering processing unit.
[0031] [ Figure 3 [] is a diagram illustrating a syntax example of audio object information.
[0032] [ Figure 4 [] is a diagram illustrating syntactic examples of object reverb information and spatial reverb information.
[0033] [ Figure 5 [Illustration] is a diagram showing the location of the reverberation component.
[0034] [ Figure 6 [Illustration] is a diagram showing the impulse response.
[0035] [ Figure 7 [Illustration] is a diagram showing the relationship between audio objects and viewing / listening locations.
[0036] [ Figure 8 [Illustration] is a diagram showing the direct sound component, the initial reflected sound component, and the post-reverberation component.
[0037] [ Figure 9 [ ] is a flowchart illustrating the audio output processing.
[0038] [ Figure 10 [Illustration 1] is a diagram showing an example configuration of an encoding device.
[0039] [ Figure 11 [] is a flowchart illustrating the encoding process.
[0040] [ Figure 12 [Illustration 1] is a diagram showing an example of a computer configuration.
[0041] Methods of implementing the present invention
[0042] In the following description, embodiments of the application of this technology will be described with reference to the accompanying drawings.
[0043] <First Embodiment>
[0044] <Signal Processing Device Configuration Example>
[0045] This technology enables the transmission of reverb parameters with high coding efficiency by adaptively selecting the encoding method of reverb parameters based on the relationship between the audio object and the viewing / listening position.
[0046] Figure 1 This is an illustration showing a configuration example of a signal processing device to which the present technology is applied.
[0047] Figure 1 The signal processing device 11 shown includes a core decoding processing unit 21 and a rendering processing unit 22.
[0048] The core decoding processing unit 21 receives and decodes the transmitted input bitstream, and provides the obtained audio object information and audio object signal to the rendering processing unit 22. In other words, the core decoding processing unit 21 is used as an acquisition unit for obtaining audio object information and audio object signal.
[0049] Here, the audio object signal is the audio signal used to reproduce the sound of the audio object.
[0050] Furthermore, audio object information is the metadata of the audio object, i.e., the audio object signal. Audio object information includes information about the audio object, which is necessary for the processing performed by the rendering processing unit 22.
[0051] Specifically, the audio object information includes object location information, direct sound gain, object reverberation information, object reverberation sound gain, spatial reverberation information, and spatial reverberation gain.
[0052] Here, object position information refers to information indicating the position of an audio object in three-dimensional space. For example, object position information includes a horizontal angle indicating the horizontal position of the audio object viewed from a reference viewing / listening position, a vertical angle indicating the vertical position of the audio object viewed from the viewing / listening position, and a radius indicating the distance from the viewing / listening position to the audio object.
[0053] In addition, direct sound gain is a gain value used for gain adjustment when generating the direct sound component of the sound of an audio object.
[0054] For example, when rendering an audio object, i.e., an audio object signal, the rendering processing unit 22 generates a signal of the direct sound component, an object-specific reverberant signal, and a space-specific reverberant signal from the audio object.
[0055] In particular, the signal of object-specific reverberation or space-specific reverberation is a signal of a component of the reverberation, such as a reflected sound or a reverberant sound from an audio object, i.e., a signal of the reverberation component obtained by performing reverberation processing on the audio object signal.
[0056] Object-specific reverberation is the initial reflected sound component of an audio object's sound, and it is the sound that contributes most significantly to the state of the audio object, such as its position in three-dimensional space. In other words, object-specific reverberation is reverberation that depends on the position of the audio object, and it changes dramatically depending on the relative positional relationship between the viewing / listening position and the audio object.
[0057] On the other hand, space-specific reverberant sound is the post-reverberation component of the sound of an audio object, and it is the sound in which the state of the audio object contributes little and the state of the environment surrounding the audio object contributes much, that is, the space surrounding the audio object.
[0058] In other words, space-specific reverberation varies greatly depending on the relative positions of the viewing / listening position and walls in the space surrounding the audio object, as well as the materials of the walls and floors. However, it remains almost unchanged depending on the relative positions of the viewing / listening position and the audio object itself. Therefore, it can be said that space-specific reverberation depends on the sound of the space surrounding the audio object.
[0059] During rendering in the rendering unit 22, reverberation processing of the audio object signal is used to generate a direct sound component from the audio object, an object-specific reverberation component, and a space-specific reverberation component. Direct sound gain is used to generate this direct sound component signal.
[0060] Object reverberation information is information about object-specific reverberant sound. For example, object reverberation information includes object reverberation location information that indicates the location of the sound image of the object-specific reverberant sound, and coefficient information used to generate object-specific reverberant sound components during reverberation processing.
[0061] Since object-specific reverberation is a component specific to the audio object, it can be said that object reverberation information is reverberation information specific to the audio object, which is used to generate object-specific reverberation components during reverberation processing.
[0062] Note that, in the following text, the location of the sound image of object-specific reverberant sound in three-dimensional space, indicated by the object reverberation location information, is also referred to as the object reverberation component location. In other words, the object reverberation component location is the arrangement position of the real or virtual speaker that outputs object-specific reverberant sound in three-dimensional space.
[0063] In addition, the object reverberation gain included in the audio object information is a gain value used for object-specific reverberation gain adjustment.
[0064] Spatial reverberation information is information about space-specific reverberant sound. For example, spatial reverberation information includes spatial reverberation location information that indicates the location of a sound image of space-specific reverberant sound, and coefficient information used to generate space-specific reverberant sound components during reverberation processing.
[0065] Since spatial reverberation is a spatially specific component of an audio object that contributes little to its sound, it can be said that spatial reverberation information is spatially specific reverberation information surrounding an audio object that is used to generate spatially specific reverberation components during reverberation processing.
[0066] Note that, in the following text, the location of the sound image of space-specific reverberant sound in three-dimensional space, indicated by spatial reverberation location information, is also referred to as the spatial reverberation component location. In other words, the spatial reverberation component location is the arrangement position of the actual or virtual loudspeaker that outputs space-specific reverberant sound in three-dimensional space.
[0067] In addition, spatial reverberation gain is a gain value used for gain adjustment of reverberant sound specific to an object.
[0068] The audio object information output from the core decoding processing unit 21 includes at least object location information, direct sound gain, object reverberation information, object reverberation sound gain, spatial reverberation information, and object location information in spatial reverberation gain.
[0069] The rendering processing unit 22 generates an output audio signal based on the audio object information and audio object signal provided by the core decoding processing unit 21, and provides the output audio signal to the speaker, recording unit, etc. in the next part.
[0070] That is, the rendering processing unit 22 performs reverberation processing based on the audio object information, and generates one or more signals of direct sound, object-specific reverberation sound, and space-specific reverberation sound for each audio object.
[0071] Then, the rendering processing unit 22 performs rendering processing on each of the obtained direct sound, object-specific reverberation, and space-specific reverberation signals via VBAP, and generates an output audio signal with a channel configuration corresponding to a reproduction device, such as a speaker system or headphones used as the output destination. Furthermore, the rendering processing unit 22 adds the signals of the same channel included in the output audio signal generated for each signal to obtain a final output audio signal.
[0072] When sound is reproduced based on the output audio signal obtained in this way, the sound image of the direct sound of the audio object is located at the position indicated by the object's position information, the sound image of the object-specific reverberant sound is located at the object's reverberation component position, and the sound image of the space-specific reverberant sound is located at the spatial reverberation component position. As a result, a more realistic audio reproduction is achieved in which the distance sensing of the audio object is appropriately controlled.
[0073] <Configuration example of a rendering unit>
[0074] Next, we will describe Figure 1 A more detailed configuration example of the rendering processing unit 22 of the signal processing device 11 shown.
[0075] Here, we will describe the case where there are two audio objects as a specific example. Note that there can be any number of audio objects, and we can process as many audio objects as our computing resources allow.
[0076] In the following text, when distinguishing between two audio objects, one audio object is also described as audio object OBJ1, and the audio object signal of audio object OBJ1 is also described as audio object signal OA1. Furthermore, the other audio object is also described as audio object OBJ2, and the audio object signal of audio object OBJ2 is also described as audio object signal OA2.
[0077] Furthermore, in the following text, the object position information, direct sound gain, object reverberation information, object reverberation sound gain, and spatial reverberation gain of audio object OBJ1 are also specifically described as object position information OP1, direct sound gain OG1, object reverberation information OR1, object reverberation sound gain RG1, and spatial reverberation gain SG1.
[0078] Similarly, in the following text, the object position information, direct sound gain, object reverberation information, object reverberation sound gain, and spatial reverberation gain of audio object OBJ2 are specifically described as object position information OP2, direct sound gain OG2, object reverberation information OR2, object reverberation sound gain RG2, and spatial reverberation gain SG2.
[0079] In the case of two audio objects as described above, for example... Figure 2 The rendering processing unit 22 is configured as shown.
[0080] exist Figure 2In the example shown, the rendering processing unit 22 includes amplification units 51-1, 51-2, 52-1, 52-2, object-specific reverberation processing unit 53-1, object-specific reverberation processing unit 53-2, amplification unit 54-1, amplification unit 54-2, space-specific reverberation processing unit 55, and rendering unit 56.
[0081] Amplification units 51-1 and 51-2 multiply the direct audio gain OG1 and direct audio gain OG2 provided by the core decoding processing unit 21 by the audio object signal OA1 and audio object signal OA2 provided by the core decoding processing unit 21 to perform gain adjustment. The resulting direct audio signal of the audio object is then provided to the rendering unit 56.
[0082] Note that in the following text, without needing to specifically distinguish between amplification unit 51-1 and amplification unit 51-2, amplification unit 51-1 and amplification unit 51-2 are also referred to as amplification unit 51.
[0083] Amplification units 52-1 and 52-2 multiply the object reverberation gain RG1 and object reverberation gain RG2 provided from the core decoding processing unit 21 with the audio object signal OA1 and audio object signal OA2 provided from the core decoding processing unit 21 to perform gain adjustment. This gain adjustment is used to adjust the loudness of each object-specific reverberation sound.
[0084] Amplification units 52-1 and 52-2 provide the gain-adjusted audio object signals OA1 and OA2 to object-specific reverberation processing units 53-1 and 53-2, respectively.
[0085] Note that in the following text, without specifically distinguishing between amplification unit 52-1 and amplification unit 52-2, amplification unit 52-1 and amplification unit 52-2 are also referred to as amplification unit 52.
[0086] The object-specific reverb processing unit 53-1 performs reverb processing on the gain-adjusted audio object signal OA1 provided by the amplification unit 52-1 based on the object reverb information OR1 provided by the core decoding processing unit 21.
[0087] Through reverb processing, one or more object-specific reverb sounds are generated for the audio object OBJ1.
[0088] Furthermore, based on the object location information OP1 provided from the core decoding processing unit 21 and the object reverberation location information included in the object reverberation information OR1, the object-specific reverberation processing unit 53-1 generates location information indicating the absolute location of the sound image of each object-specific reverberation sound in three-dimensional space.
[0089] As described above, the object position information OP1 includes information such as the horizontal angle, vertical angle, and radius that indicate the absolute position of the audio object OBJ1 based on the viewing / listening position in three-dimensional space.
[0090] On the other hand, object reverberation location information can be information indicating the absolute position (location position) of an object-specific reverberant sound image viewed from a viewing / listening position in three-dimensional space, or information indicating the relative position (location position) of an object-specific reverberant sound image relative to the audio object OBJ1 in three-dimensional space.
[0091] For example, in the case where object reverberation location information is information indicating the absolute position of an object-specific reverberant sound image viewed from a viewing / listening position in three-dimensional space, object reverberation location information includes information indicating the absolute location of an object-specific reverberant sound image based on the horizontal angle, vertical angle, and radius of the viewing / listening position in three-dimensional space.
[0092] In this case, the object-specific reverberation processing unit 53-1 uses the object reverberation location information as location information indicating the absolute location of the sound image of the object-specific reverberation sound.
[0093] On the other hand, when the object reverberation position information is information indicating the relative position of the sound image of the reverberation sound specific to the object with respect to the audio object OBJ1, the object reverberation position information is information including horizontal angle, vertical angle and radius indicating the relative position of the sound image of the reverberation sound specific to the object with respect to the audio object OBJ1 as viewed from a visual / auditory position in three-dimensional space.
[0094] In this case, based on the object location information OP1 and the object reverberation location information, the object-specific reverberation processing unit 53-1 generates information including the horizontal angle, vertical angle and radius of the sound image indicating the absolute location of the object-specific reverberation sound based on the viewing / listening position in three-dimensional space, as location information indicating the absolute position of the sound image of the object-specific reverberation sound.
[0095] The object-specific reverb processing unit 53-1 provides the rendering unit 56 with a pair of signals and position information for each of the one or more object-specific reverb sounds.
[0096] As described above, reverb processing generates object-specific reverb signals and location information, allowing each object-specific reverb signal to be processed as an independent audio object signal.
[0097] Similarly, the object-specific reverb processing unit 53-2 performs reverb processing on the gain-adjusted audio object signal OA2 provided by the amplification unit 52-2 based on the object reverb information OR2 provided by the core decoding processing unit 21.
[0098] Through reverb processing, one or more object-specific reverb sounds are generated for the audio object OBJ2.
[0099] Furthermore, based on the object location information OP2 provided from the core decoding processing unit 21 and the object reverberation location information included in the object reverberation information OR2, the object-specific reverberation processing unit 53-2 generates location information indicating the absolute location of the sound image of each object-specific reverberation sound in three-dimensional space.
[0100] Then, the object-specific reverberation processing unit 53-2 provides the rendering unit 56 with a pair of signals and position information of the object-specific reverberation sound obtained in this way.
[0101] Note that in the following text, without having to specifically distinguish between object-specific reverberation processing unit 53-1 and object-specific reverberation processing unit 53-2, object-specific reverberation processing unit 53-1 and object-specific reverberation processing unit 53-2 are also simply referred to as object-specific reverberation processing unit 53.
[0102] Amplification units 54-1 and 54-2 multiply the spatial reverberation gain SG1 and spatial reverberation gain SG2 provided from the core decoding processing unit 21 by the audio object signals OA1 and OA2 provided from the core decoding processing unit 21 to perform gain adjustment. This gain adjustment is used to adjust the loudness of each space-specific reverberant sound.
[0103] In addition, amplification units 54-1 and 54-2 provide the gain-adjusted audio object signal OA1 and audio object signal OA2 to the space-specific reverberation processing unit 55.
[0104] Note that in the following text, without specifically distinguishing between amplification unit 54-1 and amplification unit 54-2, amplification unit 54-1 and amplification unit 54-2 are also referred to as amplification unit 54.
[0105] The space-specific reverberation processing unit 55 performs reverberation processing on the gain-adjusted audio object signals OA1 and OA2 provided by the amplification units 54-1 and 54-2, based on spatial reverberation information provided by the core decoding processing unit 21. Furthermore, the space-specific reverberation processing unit 55 generates a space-specific reverberant sound signal by adding the signals obtained through the reverberation processing of audio objects OBJ1 and OBJ2. The space-specific reverberation processing unit 55 generates one or more space-specific reverberant sound signals.
[0106] Furthermore, similar to the object-specific reverberation processing unit 53, the space-specific reverberation processing unit 55 generates object position information OP1 and object position information OP2 as position information indicating the absolute location of the sound image of the space-specific reverberation sound, based on the spatial reverberation position information included in the spatial reverberation information provided from the core decoding processing unit 21.
[0107] This location information includes, for example, information such as horizontal angle, vertical angle, and radius, which indicates the absolute location of a spatially specific reverberant sound image based on the viewing / listening location in three-dimensional space.
[0108] The space-specific reverberation processing unit 55 provides the rendering unit 56 with a pair of signals of one or more space-specific reverberations obtained in this way and the location information of the space-specific reverberations. Note that space-specific reverberations can be regarded as independent audio object signals because they have location information similar to that of object-specific reverberations.
[0109] The amplification unit 51 is used as a processing block of the reverberation processing unit provided before the rendering unit 56 by the space-specific reverberation processing unit 55 described above, and performs reverberation processing based on audio object information and audio object signal.
[0110] The rendering unit 56 performs VBAP rendering processing based on each provided sound signal and the position information of each sound signal, and generates and outputs an output audio signal including signals of each channel with a predetermined channel configuration.
[0111] That is, the rendering unit 56 performs rendering processing through VBAP based on the object position information provided by the core decoding processing unit 21 and the direct sound signal provided by the amplification unit 51, and generates output audio signals for each channel for each audio object OBJ1 and audio object OBJ2.
[0112] Furthermore, the rendering unit 56 performs VBAP rendering processing on each pair of signals based on the pair of signals and the location information of the object-specific reverberation sound provided by the object-specific reverberation processing unit 53, and generates an output audio signal for each channel for each object-specific reverberation sound.
[0113] Furthermore, the rendering unit 56 performs VBAP rendering processing for each pair of signals based on the pair of signals and the location information of the space-specific reverberation sound provided by the space-specific reverberation processing unit 55, and generates an output audio signal for each channel for each space-specific reverberation sound.
[0114] Then, the rendering unit 56 adds the signals of the same channel in the output audio signals obtained for each audio object OBJ1, audio object OBJ2, object-specific reverberation and space-specific reverberation to obtain the final output audio signal.
[0115] <Input bitstream format example>
[0116] Here, an example of the format of the input bit stream provided to the signal processing device 11 will be described.
[0117] For example, the format (syntax) of the input bitstream is as follows: Figure 3 As shown. In Figure 3 In the example shown, the part indicated by the characters "object_metadata()" is the metadata of the audio object, which is part of the audio object's information.
[0118] This section of the audio object information includes object position information for the audio objects, indicating the number of audio objects as specified by the characters "num_objects". In this example, the horizontal angle position_azimuth[i], the vertical angle position_elevation[i], and the radius position_radius[i] are stored as the object position information for the i-th audio object.
[0119] In addition, the audio object information includes a reverb information flag, indicated by the characters "flag_obj_reverb", which indicates whether reverb information such as object reverb information and spatial reverb information is included.
[0120] Here, when the reverb information flag flag_obj_reverb has a value of "1", it indicates that the audio object information includes reverb information.
[0121] In other words, when the value of the reverb information flag flag_obj_reverb is "1", it can be said that reverb information, including at least one of spatial reverb information or object reverb information, is stored in the audio object information.
[0122] Note, more specifically, that depending on the value of the reuse flag use_prev described later, there are cases where the audio object information includes identification information used to identify past reverb information (i.e., the reverb ID described later) as reverb information, and does not include object reverb information or spatial reverb information.
[0123] On the other hand, when the reverb information flag flag_obj_reverb has a value of "0", it indicates that the audio object information does not include reverb information.
[0124] When the reverb information flag flag_obj_reverb is set to "1", each of the following is stored as reverb information in the audio object information: the direct sound gain indicated by the character "dry_gain[i]", the object reverb sound gain indicated by the character "wet_gain[i]", and the spatial reverb gain indicated by the character "room_gain[i]".
[0125] The direct sound gain dry_gain[i], the object-specific reverb gain wet_gain[i], and the spatial reverb gain room_gain[i] determine the mixing ratio of the direct sound, object-specific reverb, and spatial reverb in the output audio signal.
[0126] In addition, in the audio object information, the reuse flag indicated by the character "use_prev" is stored as reverb information.
[0127] The reuse flag `use_prev` is a flag indicating whether to reuse past object reverb information specified by the reverb ID as object reverb information for the i-th audio object.
[0128] Here, the reverb ID is used as identification information to identify (specify) the object reverb information sent in the input bitstream for each object reverb information.
[0129] For example, when the reuse flag `use_prev` is set to "1", it indicates that past object reverb information should be reused. In this case, the reverb ID, indicated by the character "reverb_data_id[i]", is stored in the audio object information, specifying the object reverb information to be reused.
[0130] On the other hand, when the reuse flag `use_prev` is set to "0", it indicates that the object reverb information has not been reused. In this case, the object reverb information, indicated by the character "obj_reverb_data(i)", is stored in the audio object information.
[0131] In addition, the spatial reverb information flag indicated by the characters "flag_room_reverb" is stored as reverb information in the audio object information.
[0132] The spatial reverb information flag, flag_room_reverb, is a flag that indicates the presence or absence of spatial reverb information. For example, when the value of the spatial reverb information flag, flag_room_reverb, is "1", it indicates that spatial reverb information is present, and the spatial reverb information indicated by the characters "room_reverb_data(i)" is stored in the audio object information.
[0133] On the other hand, when the spatial reverb information flag `flag_room_reverb` is set to "0", it indicates that no spatial reverb information exists, and in this case, spatial reverb information is not stored in the audio object information. Note that, similar to the case of object reverb information, a reuse flag can be stored for spatial reverb information, and spatial reverb information can be reused appropriately.
[0134] Furthermore, for example, the format (syntax) of the object reverb information obj_reverb_data(i) and the spatial reverb information room_reverb_data(i) in the audio object information of the input bitstream is as follows: Figure 4 As shown.
[0135] exist Figure 4 The example shown includes the reverb ID indicated by the character "reverb_data_id", the number of object-specific reverb sound components to be generated indicated by the character "num_out", and the tap length indicated by the character "len_ir" as object reverb information.
[0136] Note that in this example, it is assumed that the coefficients of the impulse response are stored as coefficient information used to generate object-specific reverberant sound components, and the tap length len_ir indicates the tap length of the impulse response, i.e., the number of coefficients in the impulse response.
[0137] In addition, the object reverberation location information, including num_out of the object-specific reverberation sound component to be generated, is used as object reverberation information.
[0138] That is, the horizontal azimuth angle [i], the vertical elevation angle [i], and the radius [i] are stored as the object reverberation position information of the i-th object-specific reverberation sound component.
[0139] Furthermore, as the coefficient information for the i-th object-specific reverberant sound component, the coefficients of the impulse response_response[i][j] are stored for the number of tap lengths len_ir.
[0140] On the other hand, spatial reverberation information includes the number of space-specific reverberant sound components to be generated, indicated by the character "num_out", and the tap length, indicated by the character "len_ir". The tap length len_ir is the tap length of the impulse response used to generate the coefficient information of the space-specific reverberant sound components.
[0141] In addition, spatial reverberation location information of the space-specific reverberant sound is included as spatial reverberation information, which is used to determine the number of space-specific reverberant sound components to be generated, num_out.
[0142] That is, the horizontal angle position_azimuth[i], the vertical angle position_elevation[i], and the radius position_radius[i] are stored as the spatial reverberation position information of the i-th spatially specific reverberation sound component.
[0143] Furthermore, as coefficient information for the i-th space-specific reverberant sound component, the coefficients of the impulse response_response[i][j] are stored for the number of tap lengths len_ir.
[0144] Note that in Figure 3 and Figure 4 The example illustrated in the diagram has already described an instance where the impulse response is used as coefficient information for generating object-specific and space-specific reverberant sound components. That is, an example has been described where reverberation processing using sampled reverberation is performed. However, this technique is not limited to this, and reverberation processing can be performed using parametric reverberation, etc. Furthermore, the coefficient information can be compressed using lossless coding techniques such as Huffman coding.
[0145] As described above, in the input bitstream, the information required for reverberation processing is divided into information about the direct sound (direct sound gain), information about object-specific reverberation such as object reverberation information, and information about space-specific reverberation such as spatial reverberation information, and the information obtained through the division is sent.
[0146] Therefore, for each piece of information, such as information about direct sound, information about object-specific reverberation, and information about space-specific reverberation, the transmission frequency can be appropriately mixed and the output information can be determined. That is, for example, in each frame of the audio object signal, only the necessary information can be selectively transmitted from multiple pieces of information, such as information about direct sound, based on the relationship between the audio object and the viewing / listening position. As a result, the bit rate of the input bitstream can be reduced, and more efficient information transmission can be achieved. In other words, coding efficiency can be improved.
[0147] <About output audio signal>
[0148] Next, we will describe the direct sound of the audio object reproduced based on the output audio signal, the object-specific reverberant sound, and the space-specific reverberant sound.
[0149] The relationship between the position of an audio object and the position of its reverberation component is as follows: Figure 5 As shown.
[0150] Here, near the location OBJ11 of an audio object, there are four object-specific reverberation sound object reverberation component locations from RVB11 to RVB14.
[0151] Here, the horizontal (azimuth) and vertical (elevation) angles representing the object reverberation component positions RVB11 to RVB14 are shown on the upper side of the diagram. In this example, it can be seen that four object-specific reverberation sound components are arranged around the origin O, which is the viewing / listening position.
[0152] The location and type of object-specific reverberation depend to a large extent on the position of the audio object in three-dimensional space. Therefore, it can be said that object reverberation information is reverberation information that depends on the position of the audio object in space.
[0153] Therefore, in the input bitstream, the object reverb information is not linked to the audio object, but is managed by the reverb ID.
[0154] When object reverberation information is read from the input bitstream, the core decoding processing unit 21 maintains the read object reverberation information for a certain period. That is, the core decoding processing unit 21 always maintains the object reverberation information for a predetermined period.
[0155] For example, suppose the value of the reuse flag use_prev is "1" at a predetermined time, and an instruction is given to reuse the object's reverb information.
[0156] In this case, the core decoding processing unit 21 obtains the reverb ID of the predetermined audio object from the input bitstream. That is, it reads out the reverb ID.
[0157] Then, the core decoding processing unit 21 reads the object reverb information specified by the read reverb ID from the past object reverb information stored in the core decoding processing unit 21, and reuses the object reverb information as object reverb information about the predetermined audio object at a predetermined time.
[0158] By managing object reverb information with reverb IDs in this way, for example, object reverb information sent for audio object OBJ1 can be reused just like object reverb information sent for audio object OBJ2. Therefore, the number of object reverb information entries, i.e., the amount of data, temporarily stored in the core decoding processing unit 21 can be further reduced.
[0159] Incidentally, typically, in the case of pulses being emitted into space, for example, as... Figure 6 As shown, the initial reflected sound is generated by reflections from the floor, walls, etc., which exist in the surrounding space, and in addition to the direct sound, a post-reverberation component is generated by the repetition of the reflections.
[0160] Here, the part indicated by arrow Q11 indicates the direct sound component, and the direct sound component corresponds to the signal of the direct sound obtained by the amplification unit 51.
[0161] Furthermore, the portion indicated by arrow Q12 indicates the initial reflected sound component, and the initial reflected sound component corresponds to the object-specific reverberant sound signal obtained by the object-specific reverberation processing unit 53. Additionally, the portion indicated by arrow Q13 indicates the post-reverberation component, and the post-reverberation component corresponds to the space-specific reverberant sound signal obtained by the space-specific reverberation processing unit 55.
[0162] For example, if described in a two-dimensional plane, this relationship between the direct sound, the initial reflected sound, and the after-reverberation component would be as follows: Figure 7 and Figure 8 As shown. Note that in Figure 7 and Figure 8 In the figures, corresponding parts are indicated by the same reference numerals, and their descriptions will be omitted as appropriate.
[0163] For example, such as Figure 7 As shown, it is assumed that there are two audio objects, OBJ21 and OBJ22, in an indoor space surrounded by walls represented by a rectangular frame. It is also assumed that the viewer / listener U11 is in a reference viewing / listening position.
[0164] Here, we assume that the distance from the viewer / listener U11 to the audio object OBJ21 is R.OBJ21 And the distance from the viewer / listener U11 to the audio object OBJ22 is R. OBJ22 .
[0165] In this case, such as Figure 8 As shown, the sound generated at audio object OBJ21 and directly pointing to the audience / listener U11, drawn by the dashed arrow in the diagram, is the direct sound D of audio object OBJ21. OBJ21 Similarly, the sound generated at audio object OBJ22 and directly pointing to the audience / listener U11, as shown by the dashed arrow in the attached diagram, is the direct sound D of audio object OBJ22. OBJ22 .
[0166] Furthermore, the sound generated at the audio object OBJ21, as shown by the dashed arrow in the attached diagram, and which points towards the audience / listener U11 after being reflected once by indoor walls, is the initial reflected sound E of the audio object OBJ21. OBJ21 Similarly, the sound generated at audio object OBJ22, and directed towards the audience / listener U11 after being reflected once by interior walls, etc., as shown by the dashed arrow in the attached diagram, is the initial reflected sound E of audio object OBJ22. OBJ22 .
[0167] In addition, including sound S OBJ21 And sound S OBJ22 The sound component is the after-reverberation component. Sound S OBJ21 The sound is generated at audio object OBJ21 and repeatedly reflected by interior walls, etc., to reach the audience / listener U11. OBJ22 The reverberation is generated at audio object OBJ22 and repeatedly reflected by interior walls, etc., to reach the audience / listener U11. Here, the post-reverberation component is drawn with a solid arrow.
[0168] Here, at a distance of R OBJ22 It is shorter than ROBJ21, and the audio object OBJ22 is closer to the audience / listener U11 than the audio object OBJ21.
[0169] As a result, for the audio object OBJ22, the direct sound D OBJ22 The sound that U11, as the viewer / listener, can hear is greater than the initial reflected sound E. OBJ22 This gives it an advantage. Therefore, for the reverberation of the audio object OBJ22, the direct sound gain is set to a large value, the object reverberation sound gain and the spatial reverberation gain are set to small values, and these gains are stored in the input bitstream.
[0170] On the other hand, audio object OBJ21 is farther away from the audience / listener U11 than audio object OBJ22.
[0171] As a result, for the audio object OBJ21, the initial reflected sound E of the post-reverberation component is the sound that the viewer / listener U11 can hear. OBJ21 And sound S OBJ21 Compared to direct sound D OBJ21 This gives it an advantage. Therefore, for the reverberation of the audio object OBJ21, the direct sound gain is set to a small value, the object reverberation sound gain and the spatial reverberation gain are set to large values, and these gains are stored in the input bitstream.
[0172] Furthermore, when audio object OBJ21 or audio object OBJ22 moves, the initial reflected sound components change to a large extent depending on the positional relationship between the position of the audio object and the positions of the walls and floor of the room that serves as the surrounding space.
[0173] Therefore, the object reverberation information for audio objects OBJ21 and OBJ22 must be transmitted at the same frequency as the object location information. This object reverberation information is largely dependent on the location of the audio objects.
[0174] On the other hand, since the post-reverberation component largely depends on the materials of the space, such as the walls and the floor, subjective quality can be adequately ensured by sending spatial reverberation information at the minimum required frequency and controlling only the amplitude relationship of the post-reverberation component according to the location of the audio object.
[0175] Therefore, for example, spatial reverberation information is transmitted to signal processing device 11 at a frequency lower than that of object reverberation information. In other words, core decoding processing unit 21 acquires spatial reverberation information at a frequency lower than that at which object reverberation information is acquired.
[0176] In this technique, by dividing the information required for reverberation processing into each sound component such as direct sound, object-specific reverberant sound, and space-specific reverberant sound, the amount of data required for reverberation processing can be reduced.
[0177] Typically, sampling reverberation requires approximately one second of long impulse response data. However, by dividing the necessary information for each sound component as described in this technique, the impulse response can be realized as a combination of fixed-delay and short impulse response data, thus reducing the amount of data. With this arrangement, the number of stages in the dual second-order filters can be similarly reduced, not only in sampling reverberation but also in parametric reverberation.
[0178] Furthermore, in this technology, by dividing the necessary information of each sound component and sending the information obtained through the division, the information required for reverberation processing can be sent at the desired frequency, thereby improving coding efficiency.
[0179] As described above, according to this technology, when transmitting reverberation information for controlling distance sensing, higher transmission efficiency can be achieved even in the presence of a large number of audio objects compared to translation-based rendering methods such as VBAP.
[0180] <Audio Output Processing Instructions>
[0181] Next, the specific operation of the signal processing device 11 will be described. That is, the following will refer to... Figure 9 The flowchart in the diagram describes the audio output processing of the signal processing device 11.
[0182] In step S11, the core decoding processing unit 21 decodes the received input bit stream (data).
[0183] The core decoding processing unit 21 provides the audio object signal obtained through decoding to the amplification units 51, 52 and 54, and provides the direct sound gain, object reverberation sound gain and spatial reverberation gain obtained through decoding to the amplification units 51, 52 and 54 respectively.
[0184] Furthermore, the core decoding processing unit 21 provides the object reverberation information and spatial reverberation information obtained through decoding to the object-specific reverberation processing unit 53 and the space-specific reverberation processing unit 55. Additionally, the core decoding processing unit 21 provides the object position information obtained through decoding to the object-specific reverberation processing unit 53, the space-specific reverberation processing unit 55, and the rendering unit 56.
[0185] Note that at this time, the core decoding processing unit 21 temporarily stores the object reverberation information read from the input bit stream.
[0186] Furthermore, more specifically, when the reuse flag use_prev is "1", the core decoding processing unit 21 provides the object reverb information specified by the reverb ID read from the input bitstream of the object reverb information segment stored in the core decoding processing unit 21 to the object-specific reverb processing unit 53 as the object reverb information of the audio object.
[0187] In step S12, the amplification unit 51 performs gain adjustment by multiplying the direct audio gain provided by the core decoding processing unit 21 by the audio object signal provided by the core decoding processing unit 21. Therefore, the amplification unit 51 generates a direct audio signal and provides it to the rendering unit 56.
[0188] In step S13, the object-specific reverberation processing unit 53 generates an object-specific reverberation sound signal.
[0189] That is, the amplification unit 52 multiplies the object reverberation gain provided by the core decoding processing unit 21 by the audio object signal provided by the core decoding processing unit 21 to perform gain adjustment. Then, the amplification unit 52 provides the gain-adjusted audio object signal to the object-specific reverberation processing unit 53.
[0190] Furthermore, the object-specific reverberation processing unit 53 performs reverberation processing on the audio object signal provided from the amplification unit 52 based on the impulse response coefficients included in the object reverberation information provided from the core decoding processing unit 21. That is, it performs convolution processing on the impulse response coefficients and the audio object signal to generate an object-specific reverberant sound signal.
[0191] Furthermore, the object-specific reverberation processing unit 53 generates object-specific reverberation sound location information based on the object location information provided from the core decoding processing unit 21 and the object reverberation location information included in the object reverberation information. Then, the object-specific reverberation processing unit 53 provides the obtained location information and the object-specific reverberation sound signal to the rendering unit 56.
[0192] In step S14, the space-specific reverberation processing unit 55 generates a space-specific reverberation sound signal.
[0193] That is, the amplification unit 54 performs gain adjustment by multiplying the spatial reverberation gain provided by the core decoding processing unit 21 by the audio object signal provided by the core decoding processing unit 21. Then, the amplification unit 54 provides the gain-adjusted audio object signal to the space-specific reverberation processing unit 55.
[0194] Furthermore, the space-specific reverberation processing unit 55 performs reverberation processing on the audio object signal provided from the amplification unit 54 based on the impulse response coefficients included in the spatial reverberation information provided from the core decoding processing unit 21. That is, it performs convolution processing on the impulse response coefficients and the audio object signal, adds the signals obtained through the convolution processing for each audio object, and generates a space-specific reverberant sound signal.
[0195] Furthermore, the space-specific reverberation processing unit 55 generates space-specific reverberation sound location information based on the object location information provided from the core decoding processing unit 21 and the spatial reverberation location information included in the spatial reverberation information. The space-specific reverberation processing unit 55 provides the obtained location information and the space-specific reverberation sound signal to the rendering unit 56.
[0196] In step S15, the rendering unit 56 performs rendering processing and outputs the obtained output audio signal.
[0197] That is, the rendering unit 56 performs rendering processing based on the object location information provided by the core decoding processing unit 21 and the direct sound signal provided by the amplification unit 51. In addition, the rendering unit 56 performs rendering processing based on the object-specific reverberation sound signal and location information provided by the object-specific reverberation processing unit 53, and performs rendering processing based on the space-specific reverberation sound signal and location information provided by the space-specific reverberation processing unit 55.
[0198] Then, rendering unit 56 adds the signal obtained through the rendering process of each sound component to each channel to generate the final output audio signal. Rendering unit 56 outputs the thus obtained output audio signal to the next part, and the audio output processing ends.
[0199] As described above, the signal processing device 11 performs reverberation processing and rendering processing based on audio object information, including information on the division of each component of direct sound, object-specific reverberation, and space-specific reverberation, and generates an output audio signal. This configuration improves the encoding efficiency of the input bitstream.
[0200] <Encoding device configuration example>
[0201] Next, the encoding device that generates and outputs the above-mentioned input bit stream as an output bit stream will be described.
[0202] For example, such as Figure 10 As shown, this encoding device is configured.
[0203] Figure 10 The encoding device 101 shown includes an object signal encoding unit 111, an audio object information encoding unit 112, and a grouping unit 113.
[0204] The object signal encoding unit 111 encodes the provided audio object signal using a predetermined encoding method and provides the encoded audio object signal to the grouping unit 113.
[0205] The audio object information encoding unit 112 encodes the provided audio object information and provides the encoded audio object information to the grouping unit 113.
[0206] The grouping unit 113 stores the encoded audio object signal provided by the object signal encoding unit 111 and the encoded audio object information provided by the audio object information encoding unit 112 in the bit stream to obtain the output bit stream. The grouping unit 113 sends the obtained output bit stream to the signal processing device 11.
[0207] <Description of encoding processing>
[0208] Next, the operation of the encoding device 101 will be described. That is, the following will refer to... Figure 11 The flowchart in the diagram describes the encoding process performed by the encoding device 101. For example, encoding process is performed on each frame of the audio object signal.
[0209] In step S41, the object signal encoding unit 111 encodes the provided audio object signal using a predetermined encoding method and provides the encoded audio object signal to the grouping unit 113.
[0210] In step S42, the audio object information encoding unit 112 encodes the provided audio object information and provides the encoded audio object information to the grouping unit 113.
[0211] Here, for example, audio object information including object reverberation information and spatial reverberation information is provided and encoded such that the spatial reverberation information is transmitted to the signal processing device 11 at a lower frequency than the object reverberation information.
[0212] In step S43, the grouping unit 113 stores the encoded audio object signal provided by the object signal encoding unit 111 in the bit stream.
[0213] In step S44, the grouping unit 113 stores in the bit stream object location information included in the encoded audio object information provided from the audio object information encoding unit 112.
[0214] In step S45, the grouping unit 113 determines whether the encoded audio object information provided from the audio object information encoding unit 112 includes reverberation information.
[0215] Here, since neither object reverberation information nor spatial reverberation information is included as reverberation information, it is determined that reverberation information is not included.
[0216] If it is determined in step S45 that reverberation information is not included, the process proceeds to step S46.
[0217] In step S46, grouping unit 113 sets the value of the reverberation information flag flag_obj_reverb to "0" and stores the reverberation information flag flag_obj_reverb in the bit stream. As a result, an output bit stream excluding reverberation information is obtained. After obtaining the output bit stream, the process proceeds to step S54.
[0218] On the other hand, if it is determined in step S45 that reverberation information is included, proceed to step S47.
[0219] In step S47, the grouping unit 113 sets the value of the reverberation information flag flag_obj_reverb to "1" and stores the reverberation information flag flag_obj_reverb and gain information included in the encoded audio object information provided from the audio object information encoding unit 112 in the bit stream. Here, the aforementioned direct sound gain dry_gain[i], object reverberation sound gain wet_gain[i], and spatial reverberation gain room_gain[i] are stored in the bit stream as gain information.
[0220] In step S48, grouping unit 113 determines whether to reuse object reverberation information.
[0221] For example, if the encoded audio object information provided from the audio object information encoding unit 112 does not include object reverb information but includes a reverb ID, it is determined that the object reverb information should be reused.
[0222] If it is determined in step S48 that the object reverberation information needs to be reused, the process proceeds to step S49.
[0223] In step S49, grouping unit 113 sets the value of reuse flag use_prev to "1" and stores the reuse flag use_prev and reverb ID included in the encoded audio object information provided from audio object information encoding unit 112 in the bit stream. After storing the reverb ID, the process proceeds to step S51.
[0224] On the other hand, if it is determined in step S48 that the object reverberation information will not be reused, proceed to step S50.
[0225] In step S50, grouping unit 113 sets the value of reuse flag use_prev to "0" and stores reuse flag use_prev and object reverb information included in the encoded audio object information provided from audio object information encoding unit 112 in the bit stream. After storing the object reverb information, the process proceeds to step S51.
[0226] After performing the processing of step S49 or step S50, the processing of step S51 is performed.
[0227] That is, in step S51, the grouping unit 113 determines whether the encoded audio object information provided by the audio object information encoding unit 112 includes spatial reverberation information.
[0228] If spatial reverberation information is determined to be included in step S51, the process proceeds to step S52.
[0229] In step S52, the grouping unit 113 sets the value of the spatial reverberation information flag flag_room_reverb to "1" and stores the spatial reverberation information flag flag_room_reverb and spatial reverberation information included in the encoded audio object information provided from the audio object information encoding unit 112 in the bit stream.
[0230] As a result, an output bitstream including spatial reverberation information is obtained. After obtaining the output bitstream, the processing proceeds to step S54.
[0231] On the other hand, if it is determined in step S51 that there is no spatial reverberation information, proceed to step S53.
[0232] In step S53, grouping unit 113 sets the value of the spatial reverberation information flag flag_room_reverb to "0" and stores the spatial reverberation information flag flag_room_reverb in the bit stream. As a result, an output bit stream excluding spatial reverberation information is obtained. After obtaining the output bit stream, the process proceeds to step S54.
[0233] After performing the processing in step S46, step S52, or step S53 to obtain the output bit stream, the processing in step S54 is performed. Note that the output bit stream obtained through these processes is, for example, having... Figure 3 and 4 The bit stream is in the format shown.
[0234] In step S54, the grouping unit 113 outputs the obtained output bit stream, and the encoding process ends.
[0235] As described above, the encoding device 101 stores audio object information in the bitstream, which appropriately includes information for each component of the direct sound, object-specific reverberation, and space-specific reverberation, and outputs the output bitstream. This configuration improves the encoding efficiency of the output bitstream.
[0236] Note that although examples of gain information such as direct sound gain, object reverberation sound gain, and spatial reverberation gain have been described above as audio object information, gain information can be generated on the decoding side.
[0237] In this case, for example, the signal processing device 11 generates direct sound gain, object reverberation sound gain, and spatial reverberation gain based on object location information, object reverberation location information, spatial reverberation location information, etc., included in the audio object information.
[0238] <Computer Configuration Example>
[0239] Incidentally, the above series of processes can be performed by hardware or software. In the case where the series of processes are performed by software, the program constituting the software is installed in the computer. Here, "computer" includes computers integrated into dedicated hardware, or computers capable of performing various functions by installing various programs, such as general-purpose personal computers.
[0240] Figure 12 This is a block diagram illustrating an example configuration of the computer hardware in which the program performs the series of processes described above.
[0241] In a computer, the central processing unit (CPU) 501, read-only memory (ROM) 502, and random access memory (RAM) 503 are interconnected via a bus 504.
[0242] The input / output interface 505 is also connected to the bus 504. The input unit 506, output unit 507, recording unit 508, communication unit 509, and driver 510 are connected to the input / output interface 505.
[0243] Input unit 506 includes a keyboard, mouse, microphone, and image sensor. Output unit 507 includes a display and speaker. Recording unit 508 includes a hard disk and non-volatile memory. Communication unit 509 includes a network interface. Driver 510 drives a removable recording medium 511 such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory.
[0244] In the computer configured as described above, the CPU 501 loads the program recorded in the recording unit 508 into the RAM 503 via, for example, the input / output interface 505 and the bus 504, and executes the program to perform the series of processes described above.
[0245] Programs executed by the computer (CPU 501) can be provided by recording on a removable recording medium 511, for example, as a packaging medium. Furthermore, programs can be provided via wired or wireless transmission media such as a local area network, the Internet, or digital satellite broadcasting.
[0246] In a computer, a program can be installed in the recording unit 508 via the input / output interface 505 by attaching the removable recording medium 511 to the drive 510. Alternatively, the program can be received by the communication unit 509 via a wired or wireless transmission medium and installed in the recording unit 508. Additionally, the program can be pre-installed in the ROM 502 or the recording unit 508.
[0247] Note that a program executed by a computer can be a program in which processing is performed in a time sequence in the order described in this specification, or a program in which processing is performed in parallel or at necessary time intervals (such as when a call is made).
[0248] Furthermore, the embodiments of this technology are not limited to the above embodiments, and various modifications can be made without departing from the spirit of this technology.
[0249] For example, this technology can have a cloud computing configuration, in which one function is shared and processed jointly by multiple devices via a network.
[0250] Furthermore, each step described in the flowchart above can be performed by a single device or can be performed by multiple devices shared by the device.
[0251] Furthermore, in cases where a step includes multiple types of processing, the multiple types of processing included in a step can be performed by one device or can be performed by multiple devices shared by one device.
[0252] In addition, this technology can have the following configurations.
[0253] (1) The signal processing device includes: The acquisition unit acquires reverberation information, which includes at least one of spatial reverberation information specific to the space surrounding the audio object or object reverberation information specific to the audio object and the audio object signal of the audio object; and A reverberation processing unit that generates a signal of the reverberation component of an audio object based on reverberation information and the audio object signal.
[0254] (2) The signal processing apparatus according to (1), wherein spatial reverberation information is acquired at a frequency lower than that of the object reverberation information.
[0255] (3) The signal processing apparatus according to (1) or (2), wherein when the acquisition unit acquires identification information indicating past reverberation information, the reverberation processing unit generates a signal of the reverberation component based on the reverberation information indicated by the identification information and the audio object signal.
[0256] (4) The signal processing device according to (3), wherein the identification information is information indicating the reverberation information of the object, and
[0257] The reverberation processing unit generates a reverberation component signal based on object reverberation information, spatial reverberation information, and audio object signal indicated by identification information.
[0258] (5) A signal processing apparatus according to any one of (1) to (4), wherein the object reverberation information is information that depends on the position of the audio object.
[0259] (6) A signal processing apparatus according to any one of (1) to (5), wherein the reverberation processing unit
[0260] Signals generating space-specific reverberation components based on spatial reverberation information and audio object signals, and
[0261] A signal with reverberation components specific to the audio object is generated based on the object's reverberation information and the audio object signal.
[0262] (7) Signal processing methods include: Reverberation information is acquired through a signal processing device, the reverberation information including at least one of spatial reverberation information specific to the space surrounding the audio object or spatial reverberation information specific to the audio object and the audio object signal; and The signal processing device generates a signal of the reverberation component of the audio object based on the reverberation information and the audio object signal.
[0263] (8) The procedure that enables the computer to perform the processing includes the following steps: Acquire reverberation information, which includes at least one of spatial reverberation information specific to the space surrounding the audio object or object reverberation information specific to the audio object and the audio object signal of the audio object; and The signal generates the reverberation component of the audio object based on the reverberation information and the audio object signal.
[0264] List of reference numerals
[0265] 11. Signal Processing Device
[0266] 21 Core Decoding Processing Units
[0267] 22 Rendering Processing Units
[0268] 51-1, 51-2, 51 Amplification Unit
[0269] 52-1, 52-2, 52 Amplification Units
[0270] 53-1, 53-2, 53 Object-Specific Reverb Processing Units
[0271] 54-1, 54-2, 54 amplification units
[0272] 55 Space-Specific Reverberation Processing Units
[0273] 56 rendering units
[0274] 101 Encoding Device
[0275] 111 Object signal encoding unit
[0276] 112 Audio Object Information Encoding Unit
[0277] 113 Grouping device.
Claims
1. A signal processing apparatus, comprising: The acquisition unit acquires reverberation information and the audio object signal of the audio object, wherein the reverberation information includes: Spatial reverberation information specific to the space surrounding the audio object and object reverberation information specific to the audio object; A reverb processing unit that generates a signal of the reverb component of the audio object based on the reverb information and the audio object signal; as well as A rendering unit performs rendering processing on the reverberation component of the audio object and the audio object signal to generate an output audio signal. Specifically, the spatial reverberation information is obtained at frequencies lower than those of the object's reverberation information. Wherein, when the acquisition unit acquires identification information indicating past reverberation information, the reverberation processing unit generates the signal of the reverberation component based on the previously acquired reverberation information indicated by the identification information and the audio object signal.
2. The signal processing apparatus according to claim 1, wherein, The identification information is information indicating the reverberation information of the object, and The reverberation processing unit generates a reverberation component signal based on the object reverberation information, the spatial reverberation information, and the audio object signal indicated by the identification information.
3. The signal processing apparatus according to claim 1, wherein, The object reverberation information is information that depends on the location of the audio object.
4. The signal processing apparatus according to claim 1, wherein, The reverberation processing unit: A signal for the space-specific reverberation component is generated based on the spatial reverberation information and the audio object signal. A signal for the reverberation component specific to the audio object is generated based on the object reverberation information and the audio object signal.
5. The signal processing apparatus according to claim 1, wherein, The rendering unit performs rendering processing on the positional information of the audio object signal and object reverberation information to provide the output audio signal.
6. The signal processing apparatus according to claim 1, wherein, The rendering unit performs rendering processing on the location information of the audio object signal and spatial reverberation information to provide the output audio signal.
7. The signal processing apparatus according to claim 1, wherein, The rendering unit performs rendering processing via vector-based amplitude translation (VBAP).
8. A signal processing method, comprising: The reverberation information and the audio object signal of the audio object are obtained by the signal processing device, wherein the reverberation information includes: Spatial reverberation information specific to the space surrounding the audio object and object reverberation information specific to the audio object; The signal processing device generates a signal of the reverberation component of the audio object based on the reverberation information and the audio object signal; as well as The signal processing device performs reverberation component and rendering processing on the audio object signal to generate an output audio signal. Specifically, the spatial reverberation information is obtained at frequencies lower than those of the object's reverberation information. Specifically, when identification information indicating past reverberation information is obtained, the signal of the reverberation component is generated based on the reverberation information previously obtained as indicated by the identification information and the audio object signal.
9. A computer-readable storage medium having instructions stored thereon, which, when executed by a computer, cause the computer to perform a process comprising the following steps: Acquire reverberation information and the audio object signal of the audio object, wherein the reverberation information includes: Spatial reverberation information specific to the space surrounding the audio object and object reverberation information specific to the audio object; A signal for generating the reverberation component of the audio object is generated based on the reverberation information and the audio object signal; as well as Perform rendering processing on the reverb component of the audio object and the audio object signal to generate an output audio signal. Specifically, the spatial reverberation information is obtained at frequencies lower than those of the object's reverberation information. Specifically, when identification information indicating past reverberation information is obtained, the signal of the reverberation component is generated based on the reverberation information previously obtained as indicated by the identification information and the audio object signal.
Citation Information
Patent Citations
Binauralization of rotated higher order ambisonics
CN105325015A
Multimedia Signal Processing Method and Apparatus
JP2016534586A
Speech processing device and method, encoding device, and program
WO2017043309A1