Sound processing device and sound processing method
Patent Information
- Application Number
- JP2024551488
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-02
AI Technical Summary
Current sound processing technologies face challenges in efficiently reducing the computational load for processing reflected sound while maintaining sound localization and spatial understanding, especially in immersive audio environments like VR and AR, where the number of sound rays required is high and battery life is a concern.
A sound processing device that calculates an evaluation value for reflected sound based on sound source, object, and listener position information, allowing selective processing of reflected sound to reduce the calculation load, including options to omit binaural processing and limit the number of selected reflected sounds to prevent excessive computational overhead.
This approach effectively reduces the computational load for reflected sound processing while preserving sound localization and spatial understanding, thereby extending battery life and improving audio processing efficiency in immersive audio environments.
Abstract
Description
Sound processing device and sound processing method
[0001] The present disclosure relates to a sound processing device and the like.
[0002] In recent years, products and services using ER (Extended Reality) (which may also be expressed as XR), including VR (Virtual Reality), AR (Augmented Reality), and MR (Mixed Reality), have become increasingly popular. Accordingly, the importance of acoustic processing technology that provides immersive audio to listeners in a virtual or real space by adding acoustic effects that occur according to the environment of the space to sounds emitted from a virtual sound source is increasing.
[0003] The listener may also be expressed as a listener or a user. Furthermore, Patent Document 1, Patent Document 2, Patent Document 3, and Non-Patent Document 1 disclose techniques related to the sound processing device and sound processing method of the present disclosure.
[0004] Japanese Patent No. 6288100 JP 2019-22049 A International Publication No. 2021 / 180938
[0005] Appl. Sci. 2019, 9, 2854; doi:10.3390 / app9142854 “Psychoacoustic Models for Perceptual Audio Coding-A Tutorial Review”
[0006] For example, Patent Literature 1 discloses a technology for performing signal processing on an object audio signal and presenting the processed signal to a listener. As ER technology becomes more widespread and services using ER technology become more diverse, there is a demand for acoustic processing that corresponds to differences in, for example, the acoustic quality required by each service, the signal processing capabilities of the terminal used, and the sound quality that can be provided by the sound presentation device. In addition, further improvements in acoustic processing technology are required to provide such services.
[0007] Here, an improvement in sound processing technology refers to a change to an existing sound processing. For example, the improvement in sound processing technology may provide a process for adding a new sound effect, a reduction in the amount of sound processing, an improvement in the quality of the sound obtained by the sound processing, a reduction in the amount of data used to perform the sound processing, or an simplification in the acquisition or generation of information used to perform the sound processing. Alternatively, the improvement in sound processing technology may provide a combination of any two or more of these.
[0008] In particular, improvements are required in devices or services that allow listeners to move freely in a virtual space. However, the above-mentioned effects obtained by improvements in sound processing technology are merely examples. One or more aspects grasped based on the present disclosure may be aspects conceived based on a different perspective than the above, aspects that achieve a different object than the above, or aspects that obtain an effect different from the above.
[0009] An acoustic device according to one aspect of the present disclosure includes a circuit and a memory, and the circuit uses the memory to acquire sound space information including information on a sound source in a sound space, information on objects in the sound space, and information on the position of a listener in the sound space, and calculates an evaluation value of reflected sound generated in response to sound emitted from the sound source using the sound space information.
[0010] These comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or may be realized as any combination thereof.
[0011] One aspect of the present disclosure may provide, for example, processing to impart new acoustic effects, reducing the amount of processing required for acoustic processing, improving the sound quality of audio obtained by acoustic processing, reducing the amount of data required for acoustic processing, or facilitating the acquisition or generation of information required for acoustic processing. Alternatively, one aspect of the present disclosure may provide any combination of these. As a result, one aspect of the present disclosure may provide acoustic processing suited to the listener's usage environment, thereby contributing to an improved acoustic experience for the listener.
[0012] In particular, the above-described effects can be achieved in devices or services that allow listeners to move freely within a virtual space. However, the above-described effects are merely examples of the effects of various aspects grasped based on the present disclosure. Each of one or more aspects grasped based on the present disclosure may be an aspect conceived based on a different perspective than the above, an aspect that achieves a different purpose than the above, or an aspect that obtains a different effect than the above.
[0013] FIG. 1 is a diagram showing an example of direct sound and reflected sound generated in a sound space. FIG. 2 is a diagram showing an example of a stereophonic sound reproduction system according to an embodiment. FIG. 3A is a block diagram showing an example of the configuration of an encoding device according to an embodiment. FIG. 3B is a block diagram showing an example of the configuration of a decoding device according to an embodiment. FIG. 3C is a block diagram showing another example of the configuration of an encoding device according to an embodiment. FIG. 3D is a block diagram showing another example of the configuration of a decoding device according to an embodiment. FIG. 4A is a block diagram showing an example of the configuration of a decoder according to an embodiment. FIG. 4B is a block diagram showing another example of the configuration of a decoder according to an embodiment. FIG. 5 is a diagram showing an example of the physical configuration of an audio signal processing device according to an embodiment. FIG. 6 is a diagram showing an example of the physical configuration of an encoding device according to an embodiment. FIG. 7 is a block diagram showing an example of the configuration of a rendering unit according to an embodiment. FIG. 8 is a flowchart showing an example of the operation of the audio signal processing device according to an embodiment. FIG. 9 is a diagram showing a positional relationship between a listener and an obstacle object when they are relatively far apart. FIG. 10 is a diagram showing a positional relationship between a listener and an obstacle object when they are relatively close together. FIG. 11 is a flowchart showing an example of a selection process according to an embodiment. FIG. 12 is a flowchart showing an example of an evaluation process according to an embodiment. FIG. 13 is a diagram showing an example of the angles of arrival of direct sound and reflected sound. Fig. 14 is a diagram showing an example of a method for setting threshold data based on the temporal masking phenomenon. Fig. 15 is a diagram showing an example of threshold data. Fig. 16 is a diagram showing the relationship between the time difference between direct sound and reflected sound and the threshold. Fig. 17 is a diagram showing an example of a configuration for a rendering unit to perform pipeline processing.
[0014] (Findings that form the basis of the present disclosure) Figure 1 is a diagram showing an example of direct sound and reflected sound generated in a sound space. In acoustic processing that expresses the characteristics of a virtual space with sound, it is effective to reproduce not only direct sound but also reflected sound in order to express the size of the space, the material of the walls, etc., and to accurately grasp the position of the sound source (localization of the sound image).
[0015] For example, when listening to sound in a rectangular room as shown in Figure 1, six primary reflections are generated for a single sound source, corresponding to the six walls. Reproducing these reflections provides clues for a proper understanding of the space and sound image. Furthermore, for each reflection, secondary reflections are generated from surfaces other than the surface that generated the reflection. These reflections also provide useful perceptual clues.
[0016] However, even if only secondary reflections are taken into account, one sound source will produce one direct sound and 36 (6 + 6 x 5) reflected sounds, resulting in 37 sound rays, and a considerable amount of calculation is required to process these sound rays.
[0017] Furthermore, in recent applications envisioned for the Metaverse, such as virtual meetings, virtual shopping, or virtual concerts, multiple sound sources will inevitably be present, requiring even greater amounts of computation.
[0018] In addition, listeners who listen to sounds in a virtual space use headphones or VR goggles. To provide such listeners with stereophonic sound, binaural processing is performed on each sound ray, which provides a sound pressure ratio and phase difference between the two ears to reproduce the direction of sound arrival and the sense of perspective. Therefore, if all reflected sounds are to be reproduced, the amount of calculation required becomes enormous.
[0019] On the other hand, for convenience, small storage batteries are sometimes used as the batteries for VR goggles worn by listeners who experience virtual space. In order to extend the battery life, it is desirable to reduce the computational load required for the above-mentioned processing. To achieve this, it is desirable to reduce the number of sound rays, which may number on the order of several hundred, to a degree that does not impair sound localization and spatial understanding.
[0020] Furthermore, in some sound reproduction systems, degrees of freedom such as 6 DoF (6 Degrees of Freedom) are allowed for the position and orientation of the listener. In this case, the positional relationship between the listener, the sound source, and the object that reflects the sound is not determined until playback (rendering). Therefore, the reflected sound is also not determined until playback. Therefore, it is difficult to determine the reflected sound to be processed in advance.
[0021] That is, the number of sound rays for expressing the characteristics of a virtual space with sound and the transition of the sound volume of each sound ray are calculated during rendering, so it is not easy to reduce the amount of calculation during rendering.
[0022] As a method for reducing the number of sound rays in a space, for example, Patent Document 1 discloses a method for detecting the importance of audio objects and not reproducing sounds caused by audio objects with low importance.
[0023] However, when degrees of freedom such as 6 DoF are allowed for the position and orientation of the listener, the listener perceives the positional relationship between the listener, the sound source, and the object that reflects the sound based on the direct sound generated from the sound source and the reflected sound generated when the direct sound is reflected by the object. Therefore, reducing the direct sound and reflected sound originating from a specific sound source may make it difficult to accurately perceive the sound localization and space.
[0024] Therefore, an object of the present disclosure is to provide a sound processing device and the like that can reduce the computational load for processing reflected sounds while enabling sound localization and spatial understanding.
[0025] (Summary of the Disclosure) A sound processing device according to a first aspect grasped based on the present disclosure includes a circuit and a memory, and the circuit uses the memory to acquire sound space information including information on a sound source in the sound space, information on objects in the sound space, and information on the position of a listener in the sound space, and calculates an evaluation value of a reflected sound generated in response to a sound generated from the sound source using the sound space information.
[0026] The device of the above aspect can use sound space information to appropriately calculate the evaluation value of reflected sounds that depends on information about the sound source, the object, and the listener's position. Therefore, it is possible to appropriately select reflected sounds to be processed based on the evaluation value of the reflected sounds. This makes it possible to grasp the sound localization and spatial information while reducing the computational load for reflected sounds.
[0027] An audio processing device according to a second aspect that can be understood based on the present disclosure may be an audio processing device according to the first aspect, in which the circuit controls whether or not to select reflected sound based on the evaluation value.
[0028] The device according to the above aspect can appropriately select reflected sounds to be processed based on the evaluation values of the reflected sounds.
[0029] An audio processing device according to a third aspect of the present disclosure may be an audio processing device according to the second aspect, in which the circuit does not perform binaural processing on the reflected sound if the reflected sound is not selected.
[0030] The device of the above aspect can reduce the computational load for reflected sounds by omitting binaural processing.
[0031] A sound processing device according to a fourth aspect as understood based on the present disclosure may be a sound processing device according to any one of the first to third aspects, in which the circuit calculates the volume of the reflected sound and, when the volume exceeds a predetermined threshold, calculates an evaluation value of the reflected sound.
[0032] The device of the above aspect can omit calculation of the evaluation value of the reflected sound when the volume of the reflected sound is equal to or less than a predetermined threshold, thereby reducing the computational load for the reflected sound.
[0033] A sound processing device according to a fifth aspect grasped based on the present disclosure may be the sound processing device of the second aspect, in which the circuit calculates the total computational load of one or more selected reflected sounds including the reflected sound when the reflected sound is selected based on the evaluation value, and cancels the selection of the reflected sound when the total computational load exceeds a predetermined upper limit.
[0034] The device of the above aspect can prevent the total calculation load from exceeding a predetermined upper limit, thereby reducing the calculation load for reflected sounds.
[0035] A sound processing device according to a sixth aspect as understood based on the present disclosure may be a sound processing device according to the fifth aspect, in which the total computational load is defined by the number of one or more selectively reflected sounds or the processing amount of one or more selectively reflected sounds.
[0036] The device of the above aspect can prevent the number of one or more selectively reflected sounds or the processing amount of one or more selectively reflected sounds from exceeding a predetermined upper limit, thereby reducing the calculation load for the reflected sounds.
[0037] A sound processing device according to a seventh aspect grasped based on the present disclosure may be a sound processing device according to any one of the first to sixth aspects, wherein the circuit calculates the volume of each of a plurality of reflected sounds generated as reflected sounds in a sound space, and calculates an evaluation value of the reflected sound for each of one or more reflected sounds among the plurality of reflected sounds that have a volume equal to or greater than a predetermined threshold.
[0038] The device of the above aspect can omit calculation of the evaluation value of each of the plurality of reflected sounds when the volume of the reflected sound is below a predetermined threshold, thereby reducing the computational load for the reflected sounds.
[0039] An audio processing device according to an eighth aspect that can be understood based on the present disclosure may be the audio processing device of the seventh aspect, wherein the circuit calculates a total computational load for one or more reflected sounds, and if the total computational load exceeds a predetermined upper limit, calculates an evaluation value of the reflected sounds for each of the one or more reflected sounds.
[0040] The device of the above aspect can omit calculation of the evaluation value of the reflected sound when the total calculation load is equal to or less than a predetermined upper limit, thereby reducing the calculation load for the reflected sound.
[0041] A sound processing device according to a ninth aspect grasped based on the present disclosure may be the sound processing device of any one of the first to eighth aspects, in which the circuit calculates an evaluation value of each of a plurality of reflected sounds generated as reflected sounds in a sound space, adds the computational load of the reflected sound to a total computational load for each of the plurality of reflected sounds in descending order of evaluation value, compares the total computational load with a predetermined upper limit each time the computational load of the reflected sound is added to the total computational load, selects the reflected sound if the total computational load obtained by adding the computational loads of the reflected sounds does not exceed the predetermined upper limit, and does not select one or more remaining reflected sounds after the reflected sound from among the plurality of reflected sounds if the total computational load obtained by adding the computational loads of the reflected sounds exceeds the predetermined upper limit.
[0042] The device of the above aspect can exclude the remaining reflected sounds from selection when the total calculation load, obtained by adding up the calculation loads in order, exceeds a predetermined upper limit. Therefore, the device of the above aspect can limit the reflected sounds to be processed, thereby suppressing the calculation load.
[0043] A sound processing device according to a tenth aspect as understood based on the present disclosure may be a sound processing device according to any one of the first to ninth aspects, in which the evaluation value is a sum of at least one index value among an index value relating to volume, a visual index value, an index value relating to an object, and an index value indicating the relationship between a direct sound corresponding to a reflected sound and the reflected sound.
[0044] The device of the above aspect can calculate, as an evaluation value, the sum of at least one index value from among an index value related to volume, a visual index value, an index value related to an object, and an index value indicating the relationship between direct sound and reflected sound. Therefore, it becomes possible to appropriately select reflected sound to be processed based on the index value related to volume, a visual index value, an index value related to an object, or an index value indicating the relationship between direct sound and reflected sound.
[0045] An audio processing device according to an eleventh aspect that can be understood based on the present disclosure may be the audio processing device of the tenth aspect, wherein the circuit is an audio processing device that increases the index value related to the volume the greater the volume of the sound generated from the sound source.
[0046] The device of the above aspect can calculate a higher evaluation value for a louder sound generated from a sound source, thereby making it possible to appropriately select a reflected sound to be processed based on a higher evaluation value for a louder sound generated from a sound source.
[0047] A sound processing device according to a twelfth aspect as understood based on the present disclosure may be the sound processing device of the tenth or eleventh aspect, wherein the circuitry is configured to increase the visual index value when the sound source is within the field of view of the listener compared to when the sound source is not within the field of view of the listener.
[0048] The device of the above aspect can calculate a higher evaluation value when the sound source is in the field of view than when the sound source is not in the field of view, and therefore it becomes possible to appropriately select reflected sounds to be processed based on the higher evaluation value when the sound source is in the field of view than when the sound source is not in the field of view.
[0049] A sound processing device according to a thirteenth aspect of the present disclosure may be a sound processing device according to any one of the tenth to twelfth aspects, in which the circuit increases the visual index value as the moving speed of the sound source decreases.
[0050] The device of the above aspect can calculate a higher evaluation value the slower the moving speed of the sound source, and therefore can appropriately select reflected sounds to be processed based on the higher evaluation value the slower the moving speed of the sound source.
[0051] A sound processing device according to a fourteenth aspect understood based on the present disclosure may be a sound processing device according to any one of the tenth to thirteenth aspects, in which an index value relating to an object is assigned to each object in a sound space and is included in the sound space information.
[0052] The device of the above aspect can calculate an evaluation value based on the index value assigned to each object, thereby making it possible to appropriately select reflected sounds to be processed based on the index value assigned to each object.
[0053] A sound processing device according to a fifteenth aspect as understood based on the present disclosure may be a sound processing device that is any one of the tenth to fourteenth aspects, and the circuit is a sound processing device that increases an index value indicating the relationship between direct sound and reflected sound as the angle between the direction from which direct sound arrives and the direction from which reflected sound arrives increases.
[0054] The device of the above aspect can calculate a higher evaluation value the larger the angle between the arrival direction of the direct sound and the arrival direction of the reflected sound, and therefore can appropriately select the reflected sound to be processed based on the higher evaluation value the larger the angle between the arrival direction of the direct sound and the arrival direction of the reflected sound.
[0055] A sound processing device according to a sixteenth aspect that can be understood based on the present disclosure may be the sound processing device according to any one of the tenth to fifteenth aspects, wherein the circuit is a sound processing device that increases an index value indicating the relationship between direct sound and reflected sound as the difference between the distance from the sound source that direct sound takes to reach the listener and the distance from the sound source that reflected sound takes to reach the listener after reflection increases.
[0056] The device of the above aspect can calculate a higher evaluation value the greater the difference between the distance of the direct sound and the distance of the reflected sound, and therefore can appropriately select the reflected sound to be processed based on the higher evaluation value the greater the difference between the distance of the direct sound and the distance of the reflected sound.
[0057] A sound processing device according to a seventeenth aspect grasped based on the present disclosure may be a sound processing device according to any one of the tenth to sixteenth aspects, wherein the circuit is a sound processing device that increases an index value indicating the relationship between direct sound and reflected sound as the amplitude value of the reflected sound exceeds a temporal masking threshold, which is a threshold for a temporal masking phenomenon in which the reflected sound is masked by the direct sound when the amplitude value of the reflected sound is equal to or less than a threshold.
[0058] The device of the above aspect can calculate a higher evaluation value the more the amplitude value of the reflected sound exceeds the temporal masking threshold, thereby making it possible to appropriately select the reflected sound to be processed based on the higher evaluation value the more the amplitude value of the reflected sound exceeds the temporal masking threshold.
[0059] An 18th aspect of the sound processing device grasped based on the present disclosure may be the sound processing device of any one of the 10th to 17th aspects, wherein the circuit repeatedly performs a process of reducing an index value for an object related to a selected reflected sound among a plurality of reflected sounds generated as reflected sounds in a sound space, calculating an evaluation value for reflected sounds that have not yet been selected, and selecting reflected sounds in descending order of evaluation value, and terminating the repeatedly performed process when the total calculation load of one or more reflected sounds selected from the plurality of reflected sounds exceeds a predetermined upper limit.
[0060] The device of the above aspect can terminate the process of selecting a new reflected sound when the total computational load of one or more selected reflected sounds exceeds a predetermined upper limit. Therefore, the device of the above aspect can limit the reflected sounds to be processed and reduce the computational load.
[0061] A sound processing device according to a 19th aspect of the present disclosure includes a circuit and a memory, and the circuit uses the memory to acquire information about the volume of the sound output from the sound source, corrects an evaluation value of the reflected sound corresponding to the sound using the volume information, and controls whether or not to select the reflected sound based on the corrected evaluation value.
[0062] The device of the above aspect can appropriately correct the evaluation value of the reflected sound corresponding to the sound by using the information on the volume of the sound, and can appropriately control the selection of the reflected sound.
[0063] A sound processing device according to a twentieth aspect grasped based on the present disclosure may be the sound processing device of the nineteenth aspect, in which the volume has a transition.
[0064] The device of the above aspect can appropriately correct the evaluation value of the reflected sound corresponding to the sound using information on the transition of volume, and can appropriately control the selection of the reflected sound.
[0065] An acoustic processing method according to a 21st aspect grasped based on the present disclosure includes a step of acquiring sound space information including information on a sound source in the sound space, information on an object in the sound space, and information on the position of a listener in the sound space, and a step of calculating an evaluation value of a reflected sound generated in response to a sound generated from the sound source using the sound space information.
[0066] The method of the above aspect can achieve the same effects as the sound processing device of the first aspect.
[0067] A program according to a twenty-second aspect grasped based on the present disclosure is a program for causing a computer to execute the acoustic processing method of the twenty-first aspect.
[0068] The program of the above aspect can achieve the same effect as the acoustic processing method of the 21st aspect when used with a computer.
[0069] These comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, or a recording medium.
[0070] The sound processing device, encoding device, decoding device, and stereophonic reproduction system according to the present disclosure will be described in detail below with reference to the drawings. The stereophonic reproduction system may also be expressed as an audio signal reproduction system.
[0071] Note that the embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, component placement and connection configurations, steps, and step sequences shown in the following embodiments are merely examples and are not intended to limit the aspects understood based on the present disclosure. Furthermore, among the components in the following embodiments, for example, components not included in the basic aspects described in the present disclosure or components not described in the independent claims showing the highest concepts will be described as optional components.
[0072] (Embodiment) (Example of a stereophonic sound reproduction system) Fig. 2 is a diagram showing an example of a stereophonic sound reproduction system. Specifically, Fig. 2 shows a stereophonic sound reproduction system 1000, which is an example of a system to which the acoustic processing or decoding processing of the present disclosure can be applied. Stereophonic sound is also expressed as immersive audio. The stereophonic sound reproduction system 1000 includes an audio signal processing device 1001 and an audio presentation device 1002.
[0073] The audio signal processing device 1001, also referred to as an audio processing device, performs audio processing on an audio signal emitted by a virtual sound source to generate an audio signal after the audio processing to be presented to a listener. The audio signal is not limited to a voice, and may be any audible sound. The audio processing is, for example, signal processing performed on the audio signal to reproduce one or more effects that the sound undergoes from the time it is generated by the sound source until it reaches the listener.
[0074] The audio signal processing device 1001 performs acoustic processing based on spatial information that describes factors that cause the above-mentioned effects. The spatial information includes, for example, information indicating the positions of a sound source, a listener, and surrounding objects, information indicating the shape of a space, and parameters related to sound propagation. The audio signal processing device 1001 is, for example, a PC (Personal Computer), a smartphone, a tablet, a game console, or the like.
[0075] The signal after acoustic processing is presented to the listener from the audio presentation device 1002. The audio presentation device 1002 is connected to the audio signal processing device 1001 via wireless or wired communication. The audio signal after acoustic processing generated by the audio signal processing device 1001 is transmitted to the audio presentation device 1002 via wireless or wired communication.
[0076] When the audio presentation device 1002 is configured with a plurality of devices, such as a device for the right ear and a device for the left ear, the plurality of devices present sounds in synchronization through communication between the plurality of devices or communication between each of the plurality of devices and the audio signal processing device 1001. The audio presentation device 1002 is, for example, headphones, earphones, or a head-mounted display worn on the head of a listener, or a surround speaker configured with a plurality of fixed speakers.
[0077] The stereophonic sound reproduction system 1000 may be used in combination with an image presentation device or a stereoscopic video presentation device that provides a visual ER experience, including AR / VR. For example, the space handled by the spatial information is a virtual space, and the positions of a sound source, a listener, and an object in the space are the virtual positions of a virtual sound source, a virtual listener, and a virtual object in the virtual space. The space may also be expressed as a sound space. The spatial information may also be expressed as sound space information.
[0078] 2 shows an example of a system configuration in which the audio signal processing device 1001 and the audio presentation device 1002 are separate devices, the stereophonic sound reproduction system 1000 to which the audio processing method or decoding method of the present disclosure can be applied is not limited to the configuration shown in Fig. 2. For example, the audio signal processing device 1001 may be included in the audio presentation device 1002, which may perform both audio processing and sound presentation.
[0079] The acoustic processing described in the present disclosure may be shared between the audio signal processing device 1001 and the audio presentation device 1002. A server connected to the audio signal processing device 1001 or the audio presentation device 1002 via a network may perform part or all of the acoustic processing described in the present disclosure.
[0080] Furthermore, the audio signal processing device 1001 may perform audio processing by decoding a bit stream generated by encoding at least a portion of the data of the audio signal and spatial information used for the audio processing. Therefore, the audio signal processing device 1001 may be referred to as a decoding device.
[0081] (Example of Encoding Device) Fig. 3A is a block diagram showing an example configuration of an encoding device. Specifically, Fig. 3A shows the configuration of an encoding device 1100, which is an example of an encoding device of the present disclosure.
[0082] Input data 1101 is data to be coded, including spatial information and / or an audio signal, that is input to an encoder 1102. Details of the spatial information will be described later.
[0083] The encoder 1102 encodes the input data 1101 to generate encoded data 1103. The encoded data 1103 is, for example, a bit stream generated by the encoding process.
[0084] The memory 1104 stores the encoded data 1103. The memory 1104 may be, for example, a hard disk or a solid-state drive (SSD), or may be other memory.
[0085] In the above description, a bitstream generated by an encoding process is given as an example of the encoded data 1103 stored in memory 1104, but the encoded data 1103 may be data other than a bitstream. For example, the encoding device 1100 may store converted data generated by converting a bitstream into a predetermined data format in memory 1104. The converted data may be, for example, a file or a multiplexed stream corresponding to one or more bitstreams.
[0086] Here, the file is a file having a file format such as ISO Base Media File Format (ISOBMFF), etc. The encoded data 1103 may be in the form of a plurality of packets generated by dividing the bit stream or file.
[0087] For example, the bitstream generated by the encoder 1102 may be converted into data different from the bitstream. In this case, the encoding device 1100 may include a conversion unit (not shown) and perform the conversion process in the conversion unit, or may perform the conversion process in a CPU (Central Processing Unit), which is an example of a processor described later.
[0088] (Example of Decoding Device) Fig. 3B is a block diagram showing an example configuration of a decoding device. Specifically, Fig. 3B shows the configuration of a decoding device 1110, which is an example of a decoding device according to the present disclosure.
[0089] The memory 1114 stores, for example, the same data as the coded data 1103 generated by the coding device 1100. The stored data is read from the memory 1114 and input to the decoder 1112 as input data 1113. The input data 1113 is, for example, a bitstream to be decoded. The memory 1114 may be, for example, a hard disk or an SSD, or may be some other memory.
[0090] Note that the decoding device 1110 may convert the data read from the memory 1114 and input the converted data to the decoder 1112 as input data 1113, rather than inputting the data directly to the decoder 1112 as input data 1113. The data before conversion may be, for example, multiplexed data including one or more bitstreams. Here, the multiplexed data may be a file having a file format such as ISOBMFF.
[0091] The data before conversion may also be a plurality of packets generated by dividing the bitstream or file. Data different from the bitstream may be read from memory 1114 and converted into a bitstream. In this case, decoding device 1110 may include a conversion unit (not shown) and perform the conversion process, or a CPU (an example of a processor, described later) may perform the conversion process.
[0092] Decoder 1112 decodes input data 1113 to produce an audio signal 1111 representing the audio to be presented to the listener.
[0093] (Another Example of Encoding Device) Fig. 3C is a block diagram showing another example of the configuration of an encoding device. Specifically, Fig. 3C shows the configuration of encoding device 1120, which is another example of an encoding device of the present disclosure. In Fig. 3C, the same components as those in Fig. 3A are assigned the same reference numerals as those in Fig. 3A, and descriptions of these components will be omitted.
[0094] Coding device 1100 stores coded data 1103 in memory 1104. On the other hand, coding device 1120 differs from coding device 1100 in that coding device 1120 includes a transmitting unit 1121 that transmits coded data 1103 to the outside.
[0095] The transmitter 1121 transmits to another device or a server a transmission signal 1122 generated based on the encoded data 1103 or data converted into another data format from the encoded data 1103. The data used to generate the transmission signal 1122 is, for example, the bit stream, multiplexed data, file, or packet described in the encoding device 1100.
[0096] (Another Example of Decoding Device) Fig. 3D is a block diagram showing another example of the configuration of a decoding device. Specifically, Fig. 3D shows the configuration of a decoding device 1130, which is another example of a decoding device of the present disclosure. In Fig. 3D, the same components as those in Fig. 3B are assigned the same reference numerals as those in Fig. 3B, and descriptions of these components will be omitted.
[0097] The decoding device 1110 reads input data 1113 from a memory 1114. On the other hand, the decoding device 1130 differs from the decoding device 1110 in that it includes a receiving unit 1131 that receives the input data 1113 from an external source.
[0098] The receiving unit 1131 receives a received signal 1132, acquires received data, and outputs input data 1113 to be input to the decoder 1112. The received data may be the same as the input data 1113 to be input to the decoder 1112, or may be data in a data format different from that of the input data 1113.
[0099] If the data format of the received data is different from the data format of the input data 1113, the receiving unit 1131 may convert the received data into the input data 1113. Alternatively, a conversion unit or a CPU (not shown) of the decoding device 1130 may convert the received data into the input data 1113. The received data is, for example, a bit stream, multiplexed data, a file, or a packet, as described in the encoding device 1120.
[0100] (Example of Decoder) Fig. 4A is a block diagram showing an example of the configuration of a decoder. Specifically, Fig. 4A shows the configuration of a decoder 1200, which is an example of the decoder 1112 in Fig. 3B or 3D.
[0101] The input data 1113 is an encoded bitstream, and includes encoded audio data, which is an encoded audio signal, and metadata used in acoustic processing.
[0102] The spatial information management unit 1201 acquires and analyzes metadata included in the input data 1113. The metadata includes information describing elements that act on sounds arranged in a sound space. The spatial information management unit 1201 manages spatial information used for acoustic processing obtained by analyzing the metadata, and provides the spatial information to the rendering unit 1203.
[0103] In the present disclosure, the information used for acoustic processing is expressed as spatial information, but other expressions may be used. For example, the information used for acoustic processing may be expressed as sound space information or scene information. Furthermore, when the information used for acoustic processing changes over time, the spatial information input to the rendering unit 1203 may be information expressed as a spatial state, a sound space state, a scene state, or the like.
[0104] Furthermore, the spatial information may be managed for each sound space or each scene. For example, when a plurality of different rooms are represented as virtual spaces, the rooms may be managed as a plurality of different scenes. Furthermore, the spatial information may be managed as different scenes depending on the situation represented in the same space.
[0105] Therefore, a plurality of pieces of spatial information may be managed for a plurality of sound spaces or a plurality of scenes. In managing the plurality of pieces of spatial information, an identifier for identifying each piece of spatial information may be assigned to the spatial information.
[0106] The spatial information data may be included in a bitstream, which is an example of input data 1113. Alternatively, the bitstream may include an identifier of the spatial information, and the spatial information data may be acquired from an information source other than the bitstream. Specifically, when the bitstream includes only the identifier of the spatial information, the identifier of the spatial information may be used in rendering to acquire the spatial information data stored in a memory within the device or an external server as input data 1113.
[0107] It should be noted that the information managed by the spatial information management unit 1201 is not limited to information included in the bitstream. For example, the input data 1113 may include data that is not included in the bitstream and indicates the characteristics and structure of a space acquired from software or a server that provides VR or AR.
[0108] The input data 1113 may also include data indicating the characteristics and positions of listeners or objects, etc. The input data 1113 may also include information about the positions of listeners acquired by sensors provided in the terminal including the decoding device (1110, 1130), or may include information indicating the position of the terminal estimated based on the information acquired by the sensors.
[0109] That is, the spatial information management unit 1201 may communicate with an external system or server to acquire spatial information and listener positions. The spatial information management unit 1201 may also acquire clock synchronization information from the external system and execute processing to synchronize with the clock of the rendering unit 1203.
[0110] Note that the space in the above description may be a virtually formed space, i.e., a VR space, or may be a real space or a virtual space corresponding to a real space, i.e., an AR space or an MR space. The virtual space may also be expressed as a sound field or a sound space. Furthermore, the information indicating a position in the above description may be information such as coordinate values indicating a position within a space, information indicating a relative position with respect to a predetermined reference position, or information indicating the movement or acceleration of a position within a space.
[0111] The audio data decoder 1202 decodes the encoded audio data included in the input data 1113 to obtain an audio signal.
[0112] The encoded audio data acquired by the stereophonic sound reproduction system 1000 is a bitstream encoded in a predetermined format such as MPEG-H 3D Audio (ISO / IEC 23008-3). Note that MPEG-H 3D Audio is merely one example of an encoding method that can be used to generate the encoded audio data contained in the bitstream. The encoded audio data may also be a bitstream encoded using another encoding method.
[0113] For example, the encoding method may be a lossy codec such as MP3 (MPEG-1 Audio Layer-3), AAC (Advanced Audio Coding), WMA (Windows Media Audio), AC3 (Audio Codec-3), or Vorbis. Alternatively, the encoding method may be a lossless codec such as ALAC (Apple Lossless Audio Codec) or FLAC (Free Lossless Audio Codec).
[0114] Alternatively, any other encoding method may be used. For example, PCM data may be a type of encoded audio data. In this case, the decoding process may be, for example, a process of converting an N-bit binary number into a number format (e.g., floating-point format) that can be processed by the rendering unit 1203, where the number of quantization bits of the PCM data is N.
[0115] The rendering unit 1203 acquires the audio signal and spatial information, performs acoustic processing on the audio signal using the spatial information, and outputs the audio signal after the acoustic processing (audio signal 1111).
[0116] Before starting rendering, the spatial information management unit 1201 reads metadata of the input signal, detects rendering items such as objects and sounds defined in the spatial information, and transmits them to the rendering unit 1203. After starting rendering, the spatial information management unit 1201 grasps changes over time in the spatial information and the listener's position, updates and manages the spatial information, and transmits the updated spatial information to the rendering unit 1203.
[0117] The rendering unit 1203 generates and outputs an audio signal to which acoustic processing has been applied, based on the audio signal included in the input data 1113 and the spatial information received from the spatial information management unit 1201 .
[0118] The spatial information update process and the audio signal output process with added acoustic processing may be executed in the same thread. Furthermore, the spatial information management unit 1201 and the rendering unit 1203 may each allocate their processes to independent threads. When the spatial information management unit 1201 and the rendering unit 1203 execute the spatial information update process and the audio signal output process with added acoustic processing in different threads, they may set the thread startup frequency individually, or may execute the processes in parallel.
[0119] When the spatial information management unit 1201 and the rendering unit 1203 execute processes in different independent threads, it is possible to allocate computing resources preferentially to the rendering unit 1203. This makes it possible to safely execute sound output processing in which even the slightest delay is unacceptable, for example, in which a delay of one sample (0.02 msec) would cause a popping noise.
[0120] In this case, the allocation of computational resources to the spatial information management unit 1201 is limited. However, because updating of spatial information is a process that occurs less frequently than output processing of audio signals (for example, a process such as updating the direction of the listener's face), it does not necessarily have to be performed instantaneously like output processing of audio signals. Therefore, even if the allocation of computational resources is limited, it does not have a significant impact on acoustic quality.
[0121] The spatial information may be updated periodically at preset times or intervals, or when preset conditions are met. The spatial information may also be updated manually by a listener or a sound space manager, or may be updated in response to a change in an external system.
[0122] For example, the spatial information may be updated when a listener operates a controller to instantly warp the position of his / her avatar or instantly advance or reverse the time. Alternatively, the spatial information may be updated when an administrator of the virtual space suddenly changes the environment of the space. In these cases, the thread for updating the spatial information managed by the spatial information management unit 1201 may be started as a one-off interrupt process in addition to being started periodically.
[0123] For example, the update process of the spatial information managed by the spatial information management unit 1201 is performed in an information update thread.
[0124] The role of the information update thread is, for example, to update the position and orientation of the listener's avatar placed in the virtual space based on the position and orientation of the VR goggles worn by the listener, or to update the position of an object moving in the virtual space, etc. Such processing is handled within a processing thread that runs at a relatively low frequency of about several tens of Hz.
[0125] The process of updating information indicating the characteristics of the direct sound may be performed in such a processing thread that occurs less frequently. This is because the characteristics of the direct sound change less frequently than the frequency with which audio processing frames for audio output occur. This makes it possible to relatively reduce the computational load of this process. Furthermore, updating information at an unnecessarily high frequency poses a risk of generating pulsive noise. Updating information at a low frequency makes it possible to avoid such a risk.
[0126] Fig. 4B is a block diagram showing another example of the configuration of a decoder. Specifically, Fig. 4B shows the configuration of a decoder 1210, which is another example of the decoder 1112 in Fig. 3B or 3D.
[0127] Figure 4B differs from Figure 4A in that the input data 1113 includes an unencoded audio signal rather than encoded audio data. The input data 1113 includes a bitstream including metadata and an audio signal.
[0128] The spatial information management unit 1211 is the same as the spatial information management unit 1201 in FIG. 4A, and therefore a description thereof will be omitted.
[0129] The rendering unit 1213 is the same as the rendering unit 1203 in FIG. 4A, and therefore a description thereof will be omitted.
[0130] The decoders 1112, 1200, and 1210 may be expressed as audio processing units that perform audio processing. The decoding devices 1110 and 1130 may be the audio signal processing devices 1001, and may be expressed as audio processing devices.
[0131] (Physical configuration of audio signal processing device) Fig. 5 is a diagram showing an example of the physical configuration of the audio signal processing device 1001. Note that the audio signal processing device 1001 in Fig. 5 may be the decoding device 1110 in Fig. 3B or the decoding device 1130 in Fig. 3D. The multiple components shown in Fig. 3B or Fig. 3D may be implemented by the multiple components shown in Fig. 5. Furthermore, part of the configuration described here may be provided in the audio presentation device 1002.
[0132] The audio signal processing device 1001 in FIG. 5 includes a processor 1402 , a memory 1404 , a communication IF (Interface) 1403 , a sensor 1405 , and a speaker 1401 .
[0133] The processor 1402 is, for example, a CPU, a DSP (Digital Signal Processor), or a GPU (Graphics Processing Unit). The CPU, DSP, or GPU may perform the acoustic processing or decoding processing of the present disclosure by executing a program stored in the memory 1404. The processor 1402 is, for example, a circuit that performs information processing. The processor 1402 may also be a dedicated circuit that performs signal processing on audio signals, including the acoustic processing of the present disclosure.
[0134] The memory 1404 is configured, for example, with a RAM (Random Access Memory) or a ROM (Read Only Memory). The memory 1404 may include a magnetic recording medium such as a hard disk or a semiconductor memory such as an SSD. The memory 1404 may also be an internal memory incorporated in the CPU or GPU. The memory 1404 may also store spatial information managed by the spatial information management units (1201, 1211). Threshold data, which will be described later, may also be stored.
[0135] The communication IF 1403 is a communication module compatible with a communication method such as Bluetooth (registered trademark) or WIGIG (registered trademark). The audio signal processing device 1001 communicates with another communication device via the communication IF 1403, for example, to acquire a bitstream to be decoded. The acquired bitstream is stored in the memory 1404, for example.
[0136] The communication IF 1403 is configured with, for example, a signal processing circuit and an antenna corresponding to a communication method. The communication method is not limited to Bluetooth (registered trademark) and WIGIG (registered trademark), but may also be LTE (Long Term Evolution), NR (New Radio), Wi-Fi (registered trademark), or the like.
[0137] Furthermore, the communication method is not limited to the wireless communication method described above, but may be a wired communication method such as Ethernet (registered trademark), USB (Universal Serial Bus), or HDMI (registered trademark) (High-Definition Multimedia Interface).
[0138] The sensor 1405 performs sensing to estimate the position and orientation of the listener. Specifically, the sensor 1405 estimates the position and / or orientation of the listener based on one or more detection results of the position, orientation, movement, velocity, angular velocity, acceleration, etc. of a part or the whole of the body, and generates position / or orientation information indicating the position and / or orientation of the listener.
[0139] Note that a device external to the audio signal processing device 1001 may be equipped with the sensor 1405. The part of the body may be the listener's head, etc. The position / orientation information may be information indicating the position and / or orientation of the listener in real space, or information indicating a displacement of the position and / or orientation of the listener based on the position and / or orientation of the listener at a predetermined time. Furthermore, the position / or orientation information may be information indicating a position and / or orientation relative to the stereophonic sound reproduction system 1000 or an external device equipped with the sensor 1405.
[0140] The sensor 1405 is, for example, an imaging device such as a camera or a ranging device such as a LiDAR (Laser Imaging Detection and Ranging). The sensor 1405 may capture an image of the listener's head movement and detect the head movement by processing the captured image. Alternatively, the sensor 1405 may be a device that performs position estimation using a wireless signal of any frequency band, such as a millimeter wave.
[0141] Furthermore, the audio signal processing device 1001 may acquire position information from an external device equipped with a sensor 1405 via the communication IF 1403. In this case, the audio signal processing device 1001 may not include the sensor 1405. Here, the external device is, for example, the audio presentation device 1002 described in Fig. 2 or a 3D video playback device worn on the head of a listener. In this case, the sensor 1405 is configured by combining various sensors such as a gyro sensor and an acceleration sensor.
[0142] For example, the sensor 1405 may detect the angular velocity of rotation around at least one of three mutually orthogonal axes in the sound space as the axis of rotation as the speed of movement of the listener's head, or may detect the acceleration of displacement with at least one of the three axes as the direction of displacement.
[0143] For example, the sensor 1405 may detect the amount of rotation about at least one of three mutually orthogonal axes in the sound space as the rotation axis, or the amount of displacement about at least one of the three axes as the displacement direction, as the amount of movement of the listener's head. Specifically, the sensor 1405 detects the 6 DoF positions (x, y, z) and angles (yaw, pitch, roll) as the position of the listener. The sensor 1405 is configured by combining various sensors used for detecting movement, such as a gyro sensor and an acceleration sensor.
[0144] The sensor 1405 may be realized by a camera for detecting the position of the listener, a GPS (Global Positioning System) receiver, or the like. Position information obtained by performing self-position estimation using a LiDAR or the like as the sensor 1405 may also be used. For example, when the stereophonic sound reproduction system 1000 is realized by a smartphone, the sensor 1405 is built into the smartphone.
[0145] The sensor 1405 may also include a temperature sensor such as a thermocouple that detects the temperature of the audio signal processing device 1001. The sensor 1405 may also include a sensor that detects the remaining charge of a battery provided in the audio signal processing device 1001 or a battery connected to the audio signal processing device 1001.
[0146] The speaker 1401 has, for example, a diaphragm, a drive mechanism such as a magnet or a voice coil, and an amplifier, and presents an audio signal after acoustic processing as sound to a listener. The speaker 1401 operates the drive mechanism in response to an audio signal (more specifically, a waveform signal indicating the waveform of the sound) amplified via the amplifier, and the drive mechanism vibrates the diaphragm. In this way, the diaphragm vibrating in response to the audio signal generates sound waves, which propagate through the air to the listener's ears, causing the listener to perceive the sound.
[0147] Here, an example has been given in which the audio signal processing device 1001 is equipped with a speaker 1401 and presents an audio signal after acoustic processing via the speaker 1401, but the means for presenting the audio signal is not limited to the above configuration.
[0148] For example, the audio signal after acoustic processing may be output to an external audio presentation device 1002 connected via a communication module. Communication via the communication module may be wired or wireless. As another example, the audio signal processing device 1001 may have a terminal for outputting an analog audio signal, and a cable for earphones or the like may be connected to the terminal to present the audio signal from the earphones or the like.
[0149] In the above case, the audio presentation device 1002 may be headphones, earphones, a head-mounted display, a neck speaker, a wearable speaker, or the like that are worn on the head or part of the body of the listener. Alternatively, the audio presentation device 1002 may be a surround speaker or the like that is composed of multiple fixed speakers. The audio presentation device 1002 may then reproduce an audio signal.
[0150] (Physical Configuration of Encoding Apparatus) Fig. 6 is a diagram showing an example of the physical configuration of an encoding apparatus. Encoding apparatus 1500 in Fig. 6 may be encoding apparatus 1100 in Fig. 3A or encoding apparatus 1120 in Fig. 3C, and multiple components shown in Fig. 3A or 3C may be implemented by multiple components shown in Fig. 6.
[0151] The encoding device 1500 in FIG. 6 includes a processor 1501 , a memory 1503 , and a communication IF 1502 .
[0152] The processor 1501 is, for example, a CPU, a DSP, or a GPU. The CPU, DSP, or GPU may perform the encoding process of the present disclosure by executing a program stored in the memory 1503. The processor 1501 is, for example, a circuit that performs information processing. The processor 1501 may be a dedicated circuit that performs signal processing on an audio signal, including the encoding process of the present disclosure.
[0153] The memory 1503 is configured with, for example, a RAM or a ROM. The memory 1503 may include a magnetic recording medium such as a hard disk or a semiconductor memory such as an SSD. The memory 1503 may also be an internal memory incorporated in the CPU or GPU.
[0154] The communication IF 1502 is a communication module compatible with a communication method such as Bluetooth (registered trademark) or WIGIG (registered trademark). The encoding device 1500 communicates with another communication device via the communication IF 1502, for example, and transmits an encoded bitstream.
[0155] The communication IF 1502 is configured with, for example, a signal processing circuit and an antenna corresponding to the communication method. The communication method is not limited to Bluetooth (registered trademark) and WIGIG (registered trademark), but may be LTE, NR, Wi-Fi (registered trademark), or the like. Furthermore, the communication method is not limited to a wireless communication method. The communication method may be a wired communication method such as Ethernet (registered trademark), USB, or HDMI (registered trademark).
[0156] (Configuration of Rendering Unit) Fig. 7 is a block diagram showing an example of the configuration of the rendering unit. Specifically, Fig. 7 shows an example of the detailed configuration of a rendering unit 1300 corresponding to the rendering units 1203 and 1213 in Figs. 4A and 4B.
[0157] The rendering unit 1300 is composed of an analysis unit 1301, a selection unit 1302, and a synthesis unit 1303, and applies acoustic processing to the sound data contained in the input signal and outputs the result.
[0158] The input signal may be composed of, for example, spatial information, sensor information, and sound data. The input signal may also include a bitstream composed of sound data and metadata (control information), in which case the metadata may include spatial information.
[0159] The spatial information is information about the sound space (three-dimensional sound field) created by the stereophonic sound reproduction system 1000, and is composed of information about objects included in the sound space and information about the listener. Objects include sound source objects that emit sound and act as sound sources, and non-sound-emitting objects that do not emit sound. Sound source objects can also be simply referred to as sound sources.
[0160] A non-sound-emitting object acts as an obstacle object that reflects the sound emitted by a sound source object, but a sound source object may also act as an obstacle object that reflects the sound emitted by another sound source object. Obstacle objects may also be referred to as reflecting objects.
[0161] Information commonly assigned to sound source objects and non-sound generating objects includes position information, shape information, and the rate of attenuation of the volume when the object reflects sound.
[0162] The position information is expressed as coordinate values on three axes, for example, the X-axis, Y-axis, and Z-axis, in Euclidean space, but does not necessarily have to be three-dimensional information. For example, the position information may be two-dimensional information expressed as coordinate values on two axes, the X-axis and the Y-axis. The position information of an object is determined by a representative position of a shape expressed by a mesh or voxels.
[0163] The shape information may include information about the surface material.
[0164] The attenuation rate may be expressed as a real number between 0 and 1, or may be expressed as a negative decibel value. In real space, the volume is not amplified by reflection, so a negative decibel value is set as the attenuation rate, but for example, to create an eerie feeling in an unreal space, an attenuation rate of 1 or more, i.e., a positive decibel value, may be set.
[0165] The attenuation rate may be set to a different value for each of the frequency bands constituting the plurality of frequency bands, or may be set independently for each frequency band. Furthermore, if the attenuation rate is set for each type of material on the object surface, a corresponding attenuation rate value may be used based on information about the surface material.
[0166] The spatial information may also include information indicating whether the object belongs to a living thing, information indicating whether the object is a moving object, etc. If the object is a moving object, the position indicated by the position information may move over time. In this case, information on the changed position or the amount of change is transmitted to the rendering unit 1300.
[0167] The information about the sound source object includes information commonly assigned to the sound source object and the non-sound-producing object, as well as sound data and information necessary for radiating the sound data into the sound space. The sound data is data indicating information about the frequency and intensity of the sound, and is data that expresses the sound perceived by a listener.
[0168] The sound data is typically a PCM signal, but may also be data compressed using an encoding method such as MP3. In this case, the signal must be decoded at least before it reaches the synthesis unit 1303, so the rendering unit 1300 may include a decoding unit (not shown). Alternatively, the signal may be decoded by the audio data decoder 1202.
[0169] One piece of sound data may be set for one sound source object, or multiple pieces of sound data may be set for one sound source object. Furthermore, identification information for identifying each piece of sound data may be assigned to the sound data, and the information about the sound source object may include the identification information of the sound data.
[0170] The information necessary to radiate sound data into a sound space may include, for example, information on the reference volume used as a reference for playing the sound data, information on the position of the sound source object, and information on the orientation of the sound source object (i.e., information on the directionality of the sound emitted by the sound source object).
[0171] The reference volume information may be, for example, the effective value of the amplitude value of the sound data at the sound source position when the sound data is emitted into the sound space, and may be expressed as a floating-point decibel (db) value.
[0172] For example, a reference volume of 0 db may indicate that sound is emitted into the sound space from the position indicated by the information regarding the position of the sound source object at the same volume as the signal level indicated by the sound data, without increasing or decreasing the volume.Alternatively, a reference volume of -6 db may indicate that sound is emitted into the sound space from the position indicated by the information regarding the position of the sound source object, with the volume of the signal level indicated by the sound data reduced to approximately half.
[0173] The reference volume information may be attached to each piece of sound data, or may be attached to a plurality of pieces of sound data collectively.
[0174] The information required to radiate sound data into a sound space may include, as volume information, information indicating time-series fluctuations in the volume of a sound source, for example.
[0175] For example, if the sound space is a virtual conference room and the sound source is a speaker, the volume transitions intermittently over a short period of time. That is, sound and silence alternate. If the sound space is a concert hall and the sound source is a performer, the volume is maintained for a certain period of time. If the sound space is a battlefield and the sound source is an explosive, the volume of the explosion will increase for a moment and then remain silent or low.
[0176] In this way, the information on the volume of the sound source may include not only information on the loudness of the sound but also information on the transition of the loudness of the sound. Such information may be used as information indicating the properties of the sound data.
[0177] The transition information may be represented by data indicating frequency characteristics in a time series. The transition information may be represented by data indicating the duration of a sound section. The transition information may be represented by data indicating a time series of the duration of a sound section and the duration of a silent section. The transition information may be represented by data listing, in a time series, multiple pairs of durations during which the amplitude of a sound signal can be considered steady (considered to be roughly constant) and the amplitude values of the signal during those durations.
[0178] The transition information may be expressed as data on the duration for which the frequency characteristics of the sound signal can be considered stationary. The transition information may be expressed as data listing, in time series, multiple pairs of durations for which the frequency characteristics of the sound signal can be considered stationary and the frequency characteristics during those periods. The transition information may be expressed, for example, in the form of data indicating the outline of a spectrogram.
[0179] Furthermore, the volume used as the reference for the frequency characteristics may be the reference volume. Information on the reference volume and information indicating the properties of the sound data may be used in a process of calculating the volume of direct sound or reflected sound to be perceived by the listener, or may be used in a process of selecting whether or not to perceive the direct sound or reflected sound. Other examples of volume information and methods of using it will be described later.
[0180] Information about the direction of the sound source object (orientation information) is typically expressed using yaw, pitch, and roll. Alternatively, the roll rotation may be omitted, and the direction information of the sound source object may be expressed using azimuth (yaw) and elevation (pitch). The direction information of the sound source object may change over time, and if it changes, it is transmitted to the rendering unit 1300.
[0181] Information about the listener is information about the listener's position and orientation in sound space. The information about the position (position information) is expressed as a position on the XYZ axes in Euclidean space, but it does not necessarily have to be three-dimensional information and may be two-dimensional information. Information about the listener's orientation (orientation information) is typically expressed using yaw, pitch, and roll. Alternatively, the roll rotation may be omitted, and the listener's orientation information may be expressed using azimuth (yaw) and elevation (pitch).
[0182] The position information and orientation information of the listener may change over time, and if so, is transmitted to the rendering unit 1300 .
[0183] The sensor information includes the amount of rotation or displacement detected by a sensor 1405 worn by the listener, as well as the listener's position and orientation. The sensor information is transmitted to the rendering unit 1300, which updates the listener's position and orientation information based on the sensor information. The sensor information may include, for example, position information obtained by a mobile terminal performing self-position estimation using a GPS, a camera, LiDAR, or the like.
[0184] Furthermore, information acquired from outside via a communication module may be detected as sensor information instead of the sensor 1405. Information indicating the temperature of the audio signal processing device 1001 and information indicating the remaining battery capacity may be acquired from the sensor 1405. Furthermore, the computational resources (CPU capacity, memory resources, PC performance, etc.) of the audio signal processing device 1001 or the audio presentation device 1002 may be acquired in real time.
[0185] The analysis unit 1301 analyzes the audio signal contained in the input signal and the spatial information received from the spatial information management unit (1201, 1211), and detects the information necessary to generate direct sound and reflected sound, as well as the information necessary to select whether or not to generate reflected sound.
[0186] The information required to generate the direct sound and the reflected sound is, for example, information on the characteristics of the direct sound and the reflected sound that can occur in the sound space. The reflected sound detected here is a candidate for the reflected sound that is selected by the selection unit 1302 as the reflected sound that will ultimately be generated by the synthesis unit 1303. The characteristics of the direct sound and the reflected sound are, for example, the arrival time (arrival time) and the volume at the time of arrival of each of the direct sound and the reflected sound to the listener. If multiple objects exist in the sound space as reflecting objects, the characteristics of the reflected sound are calculated for each of the multiple objects.
[0187] The information required for selecting the reflected sound to be output may be, for example, information indicating the evaluation value of the reflected sound and an upper limit of the computational resources, or information for calculating the evaluation value of the reflected sound and the information indicating the upper limit of the computational resources. In other words, the analysis unit 1301 may acquire the evaluation value of the reflected sound from an external device, a storage unit, or an input signal. Alternatively, the analysis unit 1301 or the selection unit 1302 may calculate the evaluation value of the reflected sound and the information indicating the upper limit of the computational resources using information acquired by the analysis unit 1301 from an external device, a storage unit, or an input signal.
[0188] The selection unit 1302 determines whether to select a reflected sound based on the evaluation value of the reflected sound. That is, the selection unit 1302 preferentially selects a reflected sound with a high evaluation value over a reflected sound with a low evaluation value. The evaluation value of a reflected sound is the value of the reflected sound and corresponds to, for example, the perceptual importance of the reflected sound. The higher the perceptual importance of the reflected sound, the higher the evaluation value. The perceptual importance of the reflected sound is the degree of necessity of the reflected sound used for the listener to correctly perceive the position of the sound source object in the sound space and the size of the space.
[0189] By preferentially selecting and processing reflected sounds with high evaluation values, i.e., reflected sounds with high perceptual importance, the listener is able to grasp the positioning of the sound image, such as the direction from which the sound is coming and the sense of distance to the sound source object, as well as the size and material of the space.
[0190] Furthermore, by determining which reflected sounds will not be selected before the reflected sound generation process begins, it is possible to determine not to execute processes subsequent to the process of applying acoustic effects to the reflected sounds. Therefore, it is possible to reduce the computational load compared to when determining whether to execute binaural processing after applying acoustic effects to all detected reflected sounds, or when executing binaural processing on all detected reflected sounds.
[0191] In other words, by determining which reflected sounds will not be selected based on the perceptual importance of the reflected sounds, it is possible to reduce the computational load used to generate the reflected sounds while preventing the listener's perception of sound localization and spatial perception from being impaired.
[0192] The selection unit 1302 evaluates the perceptual importance of the reflected sound based on, for example, the volume of the sound source, the visibility of the sound source, the localization of the sound source, the visibility of the reflecting object (obstacle object), information about the material of the reflecting object, and the geometric relationship between the direct sound and the reflected sound, and calculates an evaluation value. Other indices may be used to evaluate the perceptual importance of the reflected sound. The evaluation value of the reflected sound may be calculated based on any one of multiple indices related to the perceptual importance of the reflected sound, or the evaluation value of the reflected sound may be calculated comprehensively using multiple indices.
[0193] The selection unit 1302 may also obtain the evaluation value of the reflected sound from an external device or a storage unit, or may obtain the evaluation value from the input signal.
[0194] Specifically, the louder the volume of the sound source, the higher the evaluation value may be. Furthermore, in order to match the visual localization with the acoustic localization, the evaluation value may be high when the sound source object or a reflective object (obstacle object) is visible to the listener, or when the localization of the sound source object is high.
[0195] Furthermore, the difference in the arrival angle between the direct sound and the reflected sound and the difference in the arrival time between the direct sound and the reflected sound have a significant impact on the perception of the space, so if the difference in the arrival angle between the direct sound and the reflected sound is large or if the difference in the arrival time between the direct sound and the reflected sound is large, the evaluation value may be high.
[0196] Alternatively, an evaluation value of a reflected sound may be calculated using information on the difference in arrival time between the direct sound and the reflected sound. In this case, for example, a masking threshold in a well-known temporal masking phenomenon (post-masking phenomenon) may be used.
[0197] A specific method for evaluating and selecting reflected sounds by the selection unit 1302 will be described later.
[0198] The synthesis unit 1303 synthesizes the audio signal of the direct sound with the audio signal of the reflected sound that the selection unit 1302 has selected to generate.
[0199] Specifically, the synthesis unit 1303 processes the input audio signal to generate a direct sound based on information about the direct sound arrival time and volume at the time of direct sound arrival calculated by the analysis unit 1301. The synthesis unit 1303 also processes the input audio signal to generate a reflected sound based on information about the reflected sound arrival time and volume at the time of reflected sound arrival for the reflected sound selected by the selection unit 1302. The synthesis unit 1303 then synthesizes and outputs the generated direct sound and reflected sound.
[0200] (Operation of Rendering Unit) Fig. 8 is a flowchart showing an example of operation of the audio signal processing device 1001. Fig. 8 mainly shows processing executed by the rendering unit 1300 of the audio signal processing device 1001.
[0201] (Detection of direct sound and reflected sound) In the input signal analysis process (S101 in FIG. 8), the analysis unit 1301 analyzes the input signal input to the audio signal processing device 1001 to detect direct sound and reflected sound that may be generated in the sound space. The reflected sound detected here is a candidate for reflected sound that is selected by the selection unit 1302 as the reflected sound that will ultimately be generated by the synthesis unit 1303. The analysis unit 1301 also analyzes the input signal to calculate information necessary for generating direct sound and reflected sound, and information necessary for selecting the reflected sound to be generated.
[0202] First, the characteristics of each of the direct sound and the reflected sound are calculated. Specifically, the arrival time and volume of each of the direct sound and the reflected sound when they reach the listener are calculated. If multiple objects exist in the sound space as reflecting objects, the characteristics of the reflected sound are calculated for each of the multiple objects.
[0203] The direct sound arrival time (td) is calculated based on the direct sound arrival path (pd). The direct sound arrival path (pd) is a path connecting the position information S (xs, ys, zs) of the sound source object and the position information A (xa, ya, za) of the listener. The direct sound arrival time (td) is a value obtained by dividing the length of the path connecting the position information S (xs, ys, zs) and the position information A (xa, ya, za) by the speed of sound (approximately 340 m / s).
[0204] For example, the path length (X) can be calculated as (xs-xa)^2 + (ys-ya)^2 + (zs-za)^2)^0.5. The volume attenuates in inverse proportion to the distance. Therefore, if the volume of the sound source object at the position information S(xs, ys, zs) is N and the unit distance is U, the volume of the direct sound (ld) when it arrives can be calculated as ld=N*U / X.
[0205] The volume N at the sound source position may be the reference volume described above.
[0206] The reflected sound arrival time (tr) is calculated based on the reflected sound arrival path (pr), which is a path connecting the position of the sound image of the reflected sound and the position information A (xa, ya, za).
[0207] The position of the sound image of the reflected sound may be derived using, for example, the "mirror image method" or "ray tracing method," or any other method for deriving the sound image position. The mirror image method is a method for simulating a sound image by assuming that a mirror image of a wave reflected from a wall in a room exists at a position symmetrical to the sound source with respect to the wall, and that a sound wave is emitted from the position of the mirror image. The ray tracing method is a method for simulating an image (sound image) observed at a certain point by tracing waves that propagate in a straight line, such as light rays or sound rays.
[0208] Fig. 9 is a diagram showing a positional relationship between a listener and an obstacle object that is relatively far away. Fig. 10 is a diagram showing a positional relationship between a listener and an obstacle object that is relatively close. That is, Fig. 9 and Fig. 10 each show an example in which a sound image of a reflected sound is formed at a position symmetrical with respect to the sound source position across a wall. By determining the position of the sound image of the reflected sound on the x, y, and z axes based on this relationship, the arrival time of the reflected sound can be determined in the same way as the method for calculating the arrival time of a direct sound.
[0209] The arrival time of a reflected sound (tr) is a value obtained by dividing the length (Y) of the path connecting the position of the sound image of the reflected sound and the position information A (xa, ya, za) by the speed of sound (approximately 340 m / sec). The volume attenuates inversely proportional to the distance. Therefore, if the volume at the sound source position is N, the unit distance is U, and the rate of attenuation of the volume upon reflection is G, the volume at the time of arrival of the reflected sound (lr) can be calculated as lr = N * G * U / Y.
[0210] As explained above, the attenuation factor G may be expressed as a real number between 0 and 1, or may be expressed as a negative decibel value. In this case, the volume of the entire signal is attenuated by G. The attenuation factor may also be set for each frequency band constituting multiple frequency bands. In this case, the analysis unit 1301 multiplies each frequency component of the signal by a specified attenuation factor. In order to reduce the amount of calculation, the analysis unit 1301 may use a representative value or average value of multiple attenuation factors for multiple frequency bands as the overall attenuation factor, and attenuate the volume of the entire signal by that amount.
[0211] (Reflected Sound Selection Process) Next, in the reflected sound selection process (S102 in FIG. 8), the selection unit 1302 selects whether or not to generate the reflected sound calculated by the analysis unit 1301. In other words, the selection unit 1302 determines whether or not to select the reflected sound as a reflected sound to be generated. When there are multiple reflected sounds, the selection unit 1302 selects whether or not to generate each of the reflected sounds. As a result of selecting whether or not to generate each reflected sound, the selection unit 1302 may select one or more reflected sounds to be generated from among the multiple reflected sounds, or may not select any reflected sounds to be generated.
[0212] The selection unit 1302 may select reflected sounds to which other processing is to be applied, not limited to the generation processing. For example, the selection unit 1302 may select reflected sounds to which binaural processing is to be applied. Furthermore, the selection unit 1302 basically selects only one or more reflected sounds to be processed. However, the selection unit 1302 may also select only one or more reflected sounds that are not to be processed. Then, processing may be applied to one or more reflected sounds that are not selected.
[0213] For example, the selection of the reflected sounds is performed based on the allowable computational load and the perceptual importance of the reflected sounds. The flow of the selection process of the reflected sounds will be described with reference to the flowchart of FIG.
[0214] 11 is a flowchart showing an example of a selection process for reflected sounds. In this example, the selection process is performed based on the computational load and the perceptual importance of the reflected sounds, but the selection process may be performed based on only one of them.
[0215] (Acquisition of information indicating upper limit of computational load) First, the selection unit 1302 acquires information indicating an upper limit of the computational load in the audio signal processing device 1001 (S201). The information indicating the upper limit of the computational load may be determined in advance by a listener or may be acquired from an input signal.
[0216] Here, the information indicating the upper limit of the computational load may indicate the number of (one or more) reflected sounds as the upper limit, or may indicate the processing amount of (one or more) reflected sounds. When the information indicating the upper limit of the computational load indicates the number of reflected sounds as the upper limit, the predicted value of the number of reflected sounds is also used in the predicted value of the computational load of reflected sound candidates, which will be described later, so it is possible to reduce the processing amount of the selection unit 1302 compared to calculating a predicted value of the processing amount of reflected sounds.
[0217] When the information indicating the upper limit of the computational load indicates the processing amount of the reflected sound as the upper limit, the predicted value of the processing amount of the reflected sound is also used in the predicted value of the computational load of the reflected sound candidate, which will be described later, making it possible to predict the computational load more accurately. Note that the processing amount of (one or more) reflected sounds is, for example, the processing amount required to generate (one or more) reflected sounds, or the total computation amount required for processing to generate (one or more) reflected sounds.
[0218] The reflected sound processing is, for example, processing for generating reflected sounds, and is included in pipeline processing, which includes, for example, reverberation processing, early reflection processing, distance attenuation processing, binaural processing, diffraction processing, and occlusion processing.
[0219] However, these processes are merely examples, and the pipeline processing may include other processes or may not include some of the processes. For example, the rendering unit 1300 may perform diffraction processing and occlusion processing as part of the pipeline processing. Also, for example, reverberation processing may be omitted if it is not required.
[0220] The information indicating the upper limit of the computational load may be determined according to the computational resources (CPU capacity, memory resources, PC performance, remaining battery capacity, etc.) of the audio signal processing device 1001 or the audio presentation device 1002. For example, since CPU processing capacity generally increases in the order of head-mounted display, VR / AR goggles, smartphone, notebook PC, desktop PC, and supercomputer, the upper limit of the computational load may also be set to increase in the same order.
[0221] The selection unit 1302 may also acquire information indicating the temperature of the device or information indicating the remaining battery capacity from a sensor 1405 provided in the audio signal processing device 1001 or the audio presentation device 1002. The selection unit 1302 may also acquire information on the computing resources (CPU capacity, memory resources, PC performance, etc.) of the audio signal processing device 1001 or the audio presentation device 1002 in real time.
[0222] In the above case, the selection unit 1302 may acquire information indicating the upper limit of the computational load in real time, or may acquire the information periodically every time the spatial information is updated by the spatial information management unit (1201, 1211).
[0223] Furthermore, the information indicating the upper limit of the calculation load may be set according to the battery life of the audio signal processing device 1001 or the audio presentation device 1002 .
[0224] Alternatively, an upper limit on the computational load may be set for each mode, such as an "energy saving mode" that requires less computation and allows the device to be used for a longer period of time, or a "high performance mode" that requires more computation but allows more reflected sounds to be heard. In this case, the listener, an administrator who manages the stereophonic sound reproduction system 1000, or a creator of the stereophonic content may specify a desired battery life or a desired mode. Alternatively, the upper limit on the computational load may be input directly without selecting a mode.
[0225] Furthermore, information indicating an upper limit of the computational load may be set for each piece of content reproduced by the stereophonic sound reproduction system 1000. For example, for content for which immersion is more important, the upper limit of the computational load may be set high, and more reflected sounds may be selected. For content for which real-time performance is important, the upper limit of the computational load may be set low so as to prevent delays caused by an increase in the amount of processing. This prevents too many reflected sounds from being selected.
[0226] The input signal containing the content may include information indicating an upper limit of the computational load. Furthermore, the selection unit 1302 may determine the upper limit of the computational load based on information indicating the type of content or the type of mode included in the input signal. Alternatively, the selection unit 1302 may determine the upper limit of the computational load based on other flags or parameters included in the input signal, not limited to the information indicating the type of content or the type of mode.
[0227] (Extraction of reflected sounds whose volume is equal to or greater than a threshold) Next, the selection unit 1302 extracts, as selection candidates, one or more reflected sounds whose volume upon arrival is equal to or greater than a threshold, from among one or more reflected sounds detected by the analysis unit 1301 (S202). In other words, the selection unit 1302 determines not to perform subsequent processing on one or more reflected sounds whose volume upon arrival is smaller than the threshold.
[0228] Furthermore, if the volume of the direct sound when it arrives is lower than the threshold, the selection unit 1302 does not need to extract the reflected sound caused by the direct sound. The volume of the reflected sound when it arrives is lower than the volume of the direct sound when it arrives. Therefore, if the volume of the direct sound when it arrives is lower than the threshold, the volume of the reflected sound when it arrives caused by the direct sound is also lower than the threshold.
[0229] Therefore, the selection unit 1302 may extract reflected sounds whose volume upon arrival is equal to or greater than a threshold from reflected sounds resulting from direct sounds whose volume upon arrival is equal to or greater than a threshold.
[0230] That is, the selection unit 1302 may first compare the volume of the direct sound when it arrives with a threshold value. As a result, if the volume of the direct sound when it arrives is smaller than the threshold value, it is possible to determine not to extract multiple reflected sounds caused by the direct sound. Therefore, it is possible to reduce the amount of calculation compared to when the volume of the reflected sound when it arrives is calculated for each of the multiple reflected sounds caused by the direct sound and then it is determined whether to extract the reflected sound.
[0231] The threshold value to be compared with the volume of the direct sound or the reflected sound at the time of arrival may be the minimum volume reproduced in the sound space. That is, the threshold value may be the minimum audible limit indicating the volume at which the listener can perceive the sound. For example, a sound lower than this threshold may not be reproduced in the virtual space as a sound that cannot be perceived by the listener.
[0232] Furthermore, the threshold may be determined in advance by the listener or may be acquired from the input signal. The threshold of the volume at the time of arrival may be determined according to the computational resources (CPU performance, memory resources, PC performance, remaining battery capacity, etc.) of the audio signal processing device 1001 or the audio presentation device 1002. For example, since the processing power of the CPU generally increases in the order of head-mounted display, VR / AR goggles, smartphone, notebook PC, desktop PC, and supercomputer, the threshold of the volume at the time of arrival may also be set to increase in the same order.
[0233] The selection unit 1302 may also acquire information indicating the temperature of the device or information indicating the remaining battery capacity from a sensor 1405 provided in the audio signal processing device 1001 or the audio presentation device 1002. The selection unit 1302 may also acquire information on the computing resources (CPU capacity, memory resources, PC performance, etc.) of the audio signal processing device 1001 or the audio presentation device 1002 in real time.
[0234] In the above case, the selection unit 1302 may also acquire the threshold value of the sound volume at the time of arrival in real time, or may acquire it periodically every time the spatial information management units (1201, 1211) update the spatial information.
[0235] The threshold value of the sound volume at the time of arrival may be set according to the battery life of the audio signal processing device 1001 or the audio presentation device 1002 .
[0236] Alternatively, the threshold value of the sound volume at the time of arrival may be set for each mode, such as an "energy saving mode" that requires less calculation and allows the device to be used for a longer period of time, or a "high performance mode" that requires more calculation but allows more reflected sounds to be heard. In this case, the listener, an administrator who manages the stereophonic sound reproduction system 1000, or a creator of the stereophonic content may specify a desired battery life or a desired mode. Alternatively, the threshold value of the sound volume at the time of arrival may be input directly without selecting a mode.
[0237] Furthermore, a threshold for the volume of sound at the time of arrival may be set for each piece of content reproduced by the stereophonic sound reproduction system 1000. For example, for content for which immersion is more important, the threshold for the volume of sound at the time of arrival may be set high, and more reflected sounds may be selected. For content for which real-time performance is important, the threshold for the volume of sound at the time of arrival may be set low, so as to prevent delays due to increased processing volume. This prevents too many reflected sounds from being selected.
[0238] The input signal including the content may include a threshold for the volume at the time of arrival. Furthermore, the selection unit 1302 may determine the threshold for the volume at the time of arrival based on information indicating the type of content or the type of mode included in the input signal. Alternatively, the selection unit 1302 may determine the threshold for the volume at the time of arrival based on other flags or parameters included in the input signal, not limited to the information indicating the type of content or the type of mode.
[0239] (Calculation of predicted value of total computational load) Next, the selection unit 1302 calculates a predicted value of the total computational load of all reflected sounds extracted as selection candidates whose arrival volume is equal to or greater than a threshold (S203). Here, the predicted value of the computational load may be the number of (one or more) reflected sounds or a predicted value of the processing amount of (one or more) reflected sounds.
[0240] Whether to use the predicted value of the number of reflected sounds or the predicted value of the processing volume of reflected sounds as the predicted value of the computational load may be determined depending on whether the information indicating the upper limit of the computational load indicates the upper limit as the number of reflected sounds or the processing volume of reflected sounds.
[0241] When the predicted value of the calculation load is the number of reflected sounds, it is possible to reduce the processing amount of the selection unit 1302 more than when the predicted value of the calculation load is the predicted value of the processing amount of reflected sounds. When the predicted value of the calculation load is the predicted value of the processing amount of reflected sounds, it is possible to predict the calculation load more accurately by calculating the total amount of calculation required to generate (one or more) reflected sounds.
[0242] As described above, the processing of reflected sounds is, for example, processing for generating reflected sounds, and is processing included in pipeline processing.
[0243] Whether each process included in the pipeline processing is necessary or not depends on the properties of the reflected sound. Therefore, the predicted value of the amount of calculation of the pipeline processing (i.e., the amount of processing for one reflected sound) may be different for each reflected sound.
[0244] Furthermore, in order to reduce the processing load for predicting the amount of calculation in pipeline processing, a predicted value of the processing amount for all reflected sounds may be calculated by assuming that the same processing is performed on each reflected sound. In other words, the same predicted value may be applied to the predicted value of the processing amount for each reflected sound to calculate a predicted value of the processing amount for all reflected sounds.
[0245] In calculating the predicted value of the total calculation load, the predicted value of the total calculation load for a plurality of reflected sounds may be calculated, or the predicted value of the total calculation load for one reflected sound may be calculated.
[0246] Furthermore, the reflected sounds used to calculate the predicted value of the total calculation load may be all of the reflected sounds extracted as selection candidates, or only some of the reflected sounds extracted as selection candidates. When only some of the reflected sounds extracted as selection candidates are used to calculate the predicted value of the total calculation load, the predicted value of the number or processing amount of some of the reflected sounds may be used as the predicted value of the total calculation load.
[0247] (Comparison of predicted total computational load with upper limit of computational load) Next, the selection unit 1302 compares the calculated predicted total computational load with the upper limit of the computational load, and determines whether the predicted total computational load exceeds the upper limit of the computational load (S204). If the predicted total computational load exceeds the upper limit of the computational load (Yes in S204), the selection unit 1302 performs selection processing (S205 to S211) based on the evaluation value. If the predicted total computational load does not exceed the upper limit of the computational load (No in S204), the selection unit 1302 selects all of the reflected sounds extracted as selection candidates, and ends the processing.
[0248] (Selection of Reflected Sound Based on Evaluation Value) In the selection process based on the evaluation value, the selection unit 1302 calculates an evaluation value of each of the reflected sounds that are candidates for selection based on the perceptual importance of the reflected sound, and controls whether or not to select the reflected sound based on the evaluation value. For example, the selection unit 1302 selects the reflected sounds in descending order of evaluation value. A specific method for calculating the evaluation value of the reflected sound will be described later. Here, an example of the selection process in which the reflected sound is selected based on the evaluation value will be described.
[0249] For example, the selection unit 1302 executes a loop process in which the computational loads of the selected reflected sounds are sequentially added up, and ends the selection process when the cumulative total exceeds the upper limit of the computational load (S205 to S211). That is, when the cumulative total of the computational loads of one or more selected reflected sounds exceeds the upper limit of the computational load (Yes in S209), the selection unit 1302 determines that the remaining undetermined reflected sounds are not selected reflected sounds, and ends the selection process.
[0250] Specifically, the selection unit 1302 first sets the count of the total calculation load to zero (S205). The selection unit 1302 then calculates an evaluation value for each extracted reflected sound (S206). The selection unit 1302 then determines to select the reflected sound with the highest evaluation value (S207). The selection unit 1302 then adds the calculation load of the reflected sound determined to be selected to the total calculation load (S208).
[0251] If the total computational load exceeds the upper limit of the computational load (Yes in S209), the selection unit 1302 determines that the remaining undetermined reflected sounds are not to be selected, and ends the selection process. In this case, the selection unit 1302 may re-determine that the reflected sounds that were last determined to be selected are not to be selected. This makes it possible to keep the total computational load below the upper limit of the computational load.
[0252] The processes after the selection process are not applied to the reflected sounds that are determined not to be selected, that is, it is determined that the remaining reflected sounds will not be generated.
[0253] If there is an undetermined reflected sound (selection or non-selection) among the reflected sounds extracted as selection candidates (Yes in S210), the selection unit 1302 repeats the processing (S207 to S209), and if there is no undetermined reflected sound (No in S210), it ends the selection processing.
[0254] Furthermore, for the selected reflected sound, a process may be performed to lower the importance of the sound source object and the reflecting object that generate the reflected sound by a predetermined amount (S211).
[0255] As a result, when a reflected sound caused by a sound source object and a reflection object is selected, another undetermined reflected sound caused by the same sound source object or the same reflection object is less likely to be selected in the next selection process. In other words, when no reflected sound caused by a sound source object or a reflection object is selected, any reflected sound caused by that sound source object or that reflection object is more likely to be selected in the next selection process.
[0256] As a result, the selection of only the reflected sound caused by a specific sound source object or reflective object is suppressed, and the presence of only the specific sound source object or reflective object in the sound space is increased, while the presence of other sound source objects or reflective objects is suppressed from being lost.
[0257] In other words, when a certain reflected sound is selected, the value of the "sound source" of that reflected sound may be lowered. This makes it more likely that a reflected sound related to a different sound source will be selected in the next turn. Also, when a certain reflected sound is selected, the value of the "wall" (reflective object) that generated that reflected sound may be lowered. This makes it more likely that a reflected sound generated by a different wall will be selected in the next turn.
[0258] For example, if three sound sources (direct sounds) exist in a cubic room, theoretically, 18 reflected sounds (3 x 6 sides) will be generated. However, generating all of the reflected sounds that theoretically occur in a sound space is difficult due to the computational load. When it is difficult to select all 18 reflected sounds, the reflected sounds are selected so that the influence of the three "sound sources" and six "walls" is reflected evenly and evenly. This makes it possible to reduce the amount of computation required to generate reflected sounds while maintaining the presence of the three "sound sources" and six "walls."
[0259] In the above example, for example, the sound source objects are represented as X, Y, and Z, the walls are represented as R1 to R6, and the reflected sounds are represented as x1 to x6, y1 to y6, and z1 to z6. If the reflected sounds x1 to x6 are selected even though there is only the amount of calculation required to generate six reflected sounds, the sound source objects Y and Z will have a weak presence in the sound space.
[0260] Furthermore, if six reflected sounds x1, y1, z1, x2, y2, and z2 are selected, the sense of reality of the walls R3 to R6 in the sound space will not be reproduced. On the other hand, for example, if the volume of sound source Y is almost zero, expressing the sense of reality of sound source Y is not that important. Therefore, the evaluation values of the reflected sounds y1 to y6 may be low. In this way, when determining which reflected sounds to select given limited computational resources, the reflected sounds may not be selected randomly, but may be selected evenly based on their importance from acoustic, auditory, and visual perspectives.
[0261] The method for determining the reflected sounds to be selected is not limited to determining the reflected sounds in descending order of evaluation value. For example, reflected sounds with evaluation values equal to or greater than a threshold may be selected, and reflected sounds with evaluation values below the threshold may not be selected. Also, reflected sounds from layers with high evaluation values may be selected at a predetermined rate. Alternatively, reflected sounds from layers with low evaluation values may not be selected at a predetermined rate. In these cases, a loop process for sequentially adding up the computational load of reflected sounds may not be performed.
[0262] (Evaluation Process) Fig. 12 is a flowchart showing an example of the evaluation process. A specific method for determining an evaluation value will be described with reference to the flowchart shown in Fig. 12.
[0263] The selection unit 1302 may calculate an evaluation value of the reflected sound using a pre-set evaluation method based on, for example, the volume of the sound source, the visibility of the sound source, the positioning of the sound source, the visibility of the reflecting object (obstacle object), or the geometric relationship between the direct sound and the reflected sound.
[0264] Specifically, the selection unit 1302 acquires a plurality of reflected sounds, each of which has been extracted as a selection candidate, and calculates an evaluation value for each of the plurality of reflected sounds based on the perceptual importance of the reflected sound.
[0265] For example, an evaluation score may be assigned to the reflected sound for each of the following multiple indices, and an evaluation value may be assigned to the reflected sound based on the evaluation score. Of course, the multiple indices for evaluation are not limited to the following multiple indices. Furthermore, any one of the multiple indices may be used, any two or more of the multiple indices may be used, or all of the multiple indices may be used. Furthermore, the order of evaluation of the multiple indices may be determined based on a predetermined priority order of the indices.
[0266] (Calculation of Evaluation Points) Specifically, an index related to a sound source object may be used as an evaluation index for a reflected sound. Furthermore, when a reflected sound caused by a sound source object is selected as described above, the value of the sound source object may be reduced. This makes it possible to evenly reproduce reflected sounds caused by many sound source objects without favoring reflected sounds caused by a specific sound source object. Therefore, it is possible to secure clues for the listener to correctly perceive the localization of each sound source.
[0267] For example, when 30 reflected sounds are selected from 300 reflected sounds caused by 10 sound sources, selecting 30 reflected sounds caused by a specific sound source makes it difficult to grasp the localization of other sound sources. Also, allocating three reflected sounds to each of the 10 sound sources is not necessarily optimal. Therefore, evaluation points may be assigned to reflected sounds caused by a sound source object based on the importance of the sound source object or the importance of the direct sound emitted by the sound source object.
[0268] For example, as will be described later, the importance of a sound source object, i.e., the importance of a direct sound, may be evaluated based on the audibility of the direct sound or the visibility of the sound source object. This evaluation may be used to evaluate a reflected sound caused by the direct sound generated from the sound source object. In other words, an evaluation score for an index related to the sound source object may be assigned to a reflected sound based on the audibility of the direct sound or the visibility of the sound source object.
[0269] It goes without saying that the evaluation of the sound source object and the direct sound may be used not only to select the reflected sound but also to select the direct sound.
[0270] The selection unit 1302 may evaluate the sound source object based on the audibility, i.e., ease of hearing, of the direct sound, and use the evaluation as an evaluation index for the reflected sound (S301). For example, an evaluation score A obtained by evaluating the audibility using information about the loudness of the direct sound may be assigned to the sound source object (direct sound) and the reflected sound.
[0271] Specifically, a sound source object with a high volume may be assigned a higher evaluation score A than a sound source object with a low volume. Similarly, a reflected sound caused by a sound source object with a high volume may be assigned a higher evaluation score A than a reflected sound caused by a sound source object with a low volume.
[0272] It goes without saying that since the loudness of a sound is generally determined by a volume or amplitude value, an amplitude value may be used instead of the volume. In other words, the information regarding the loudness of a sound may be a volume (decibel value) or an amplitude value. Since the volume or amplitude value of a sound usually changes from moment to moment, it goes without saying that the information regarding the loudness of a sound used for evaluation may be a reference volume assigned to a sound source object or information indicating the loudness of a sound that changes over time.
[0273] Note that both information on the reference volume and information on the volume that transitions over time may be used as information indicating the loudness of the direct sound. For example, the evaluation score of the sound source object may be calculated based on the information on the reference volume, and then the evaluation score of the direct sound may be calculated by correcting the evaluation score using information indicating the loudness of the transitioning sound. Of course, the evaluation score of the direct sound may first be calculated using information indicating the loudness of the transitioning sound, and then the evaluation score of the direct sound may be corrected using the reference volume assigned to the sound source object.
[0274] Furthermore, the evaluation score of the sound source object (direct sound) may be calculated using only either the reference volume information or the volume information that transitions over time.
[0275] For example, if the virtual space is a virtual conference room and the direct sound is conversation, the volume transitions intermittently over a short period of time. That is, sound and silence alternate. If the virtual space is a concert hall and the direct sound is a musical performance, the volume is maintained for a certain period of time. If the virtual space is a battlefield and the direct sound is an explosion, the volume increases for a moment and then remains silent or low.
[0276] In this way, the volume information of the sound source may include not only information about the loudness of the sound but also information about the transition of the sound volume. For example, the information may be information that lists in chronological order multiple pairs of durations during which the volume is roughly constant and the volume values for those durations.
[0277] Furthermore, efforts to use temporal transitions in the frequency characteristics of signals in acoustic processing of virtual spaces have been widely undertaken in the past (see, for example, Patent Document 1). In light of such prior art, it goes without saying that the above pair may be a pair of a time length during which the frequency characteristics are constant and the frequency characteristics themselves.
[0278] The selection unit 1302 may evaluate the sound source object based on its visibility and use the evaluation as an evaluation index for the reflected sound (S302).
[0279] Specifically, the selection unit 1302 may detect a sound source object that is visible to the listener in a video provided from a video providing device in synchronization with the sound provided from the audio presentation device 1002 .
[0280] That is, a sound source object included in a video provided from a video providing device in synchronization with a sound provided from the audio presentation device 1002 may be detected as a visible sound source object. The determination of whether or not the sound source object is visible may be performed according to an update process of the spatial information managed by the spatial information management units (1201, 1211), that is, according to a process in an information update thread.
[0281] The selection unit 1302 may then assign a higher evaluation score V to the sound source object detected as a visible object compared to a sound source object that is not visible to the listener.
[0282] Similarly, the selection unit 1302 may assign a higher evaluation score V to direct sounds and reflected sounds resulting from sound source objects that are visible to the listener, compared to direct sounds and reflected sounds resulting from sound source objects that are not visible to the listener.
[0283] Furthermore, the method for detecting an object visible to the listener is not limited to the method based on the video provided in synchronization with the sound as described above. For example, the visible object may be determined based on the relationship between the position of the listener and the position of the object in the sound space.
[0284] That is, if there is no obstacle object that acts as an obstruction between the position of the listener and the position of the sound source object based on the spatial information managed by the spatial information management units (1201, 1211), the sound source object may be identified as being visible to the listener. More specifically, if there is no obstacle object on the propagation paths of direct sound and reflected sound that may be generated in the sound space calculated by the analysis unit 1301, the sound source object or the reflecting object may be identified as being visible to the listener.
[0285] Alternatively, sound source objects located within a predetermined distance range from the listener's position may be identified as being visible to the listener.
[0286] The selection unit 1302 may then evaluate the importance of reflected sounds resulting from sounds generated from sound source objects identified as being visible to the listener, and assign a high evaluation score V to such reflected sounds.
[0287] By using an index of the visibility of the sound source object, it becomes possible to appropriately select reflected sounds that match the visual localization in the video and the auditory localization (acoustic localization) in the sound. If the visual localization of the sound source object visible to the listener does not match the acoustic localization based on the direct sound, reflected sound, and their relationship provided by the audio presentation device 1002, the sense of localization becomes unnatural, causing the listener to feel uncomfortable and reducing the sense of immersion.
[0288] On the other hand, for non-visible sound source objects, even if the acoustic localization is slightly different from the original localization, it does not feel strange, so the evaluation score V of the reflected sound caused by the non-visible sound source object may be low.
[0289] The audio presentation device 1002 and the video provision device may be the same device, such as VR goggles and a head-mounted display, or may be separate devices, such as earphones and a smartphone.
[0290] The selection unit 1302 may evaluate the sound source object based on its localization, and use the evaluation result as an evaluation index for the reflected sound (S303).
[0291] Specifically, the selection unit 1302 may detect the moving speed of a sound source object visible to the listener in a video provided from the video providing device in synchronization with the sound provided from the audio presentation device 1002. Then, the selection unit 1302 may assign a higher evaluation score S to a sound source object moving slowly compared to a sound source object moving quickly. Similarly, the selection unit 1302 may assign a higher evaluation score S to a direct sound and a reflected sound caused by a sound source object moving slowly compared to a direct sound and a reflected sound caused by a sound source object moving quickly.
[0292] Furthermore, when a sound source object is stationary, the selection unit 1302 may assign the highest evaluation score S to the direct sound and reflected sound originating from the sound source object in the index of the sound source object's localization. For example, the selection unit 1302 may assign a higher evaluation score to the reflected sound originating from the sound emitted by a stationary sound source object than to the reflected sound originating from the sound emitted by a moving sound source object.
[0293] Furthermore, for example, the selector 1302 may assign higher evaluation points to the direct sound and the reflected sound caused by the sound source object as the moving speed of the sound source object is slower.
[0294] The localization of a fast-moving sound source is dominated by visual localization and the direction of arrival of the direct sound. Therefore, the selection unit 1302 may assign a low evaluation score to a reflected sound from a fast-moving sound source so that the reflected sound is not selected.
[0295] By using an index of the localization of the sound source object, it becomes possible to appropriately select reflected sounds that match the visual localization and the acoustic localization, thereby making it possible to prevent the sense of localization from becoming unnatural due to a mismatch between the visual localization and the acoustic localization.
[0296] The selection unit 1302 may use the importance of the reflecting object as an evaluation index for the reflected sound, that is, the selection unit 1302 may evaluate the importance of the reflecting object (S304).
[0297] For example, the spatial information may include information about a reflecting object. The selection unit 1302 may then evaluate the importance of the reflecting object based on the information about the reflecting object. The selection unit 1302 may then assign an evaluation score to the reflected sound caused by the object based on the importance of the reflecting object.
[0298] For example, the selection unit 1302 may determine the importance of a reflection object based on information included in the input signal or metadata included in the bitstream. Alternatively, the selection unit 1302 may determine the importance of a reflection object based on other flags, parameters, or the like included in the input signal.
[0299] For example, the importance of a reflective object may be determined based on the visibility of the reflective object (obstacle object), information about the material of the reflective object, etc. For example, the importance may be determined to be high according to the visibility of the reflective object (obstacle object), that is, for a sound source object that is visible to a listener.
[0300] Specifically, a reflective object (obstacle object) visible to the listener may be detected in the video provided from the video providing device in synchronization with the sound provided from the audio presentation device 1002. Then, the selection unit 1302 may assign a higher importance to the reflective object detected as a visible object compared to a reflective object not visible to the listener.
[0301] Similarly, the selection unit 1302 may assign a higher evaluation score V to a reflected sound caused by a reflecting object visible to the listener, compared to a reflected sound caused by a reflecting object not visible to the listener. In other words, the selection unit 1302 may evaluate the importance of a reflected sound caused by a reflecting object within the listener's field of view as being high, and assign a higher evaluation score V to such a reflected sound.
[0302] As a method for detecting a reflecting object visible to the listener, a method similar to the above-described method for detecting a visible sound source object can be used.
[0303] Here, a method of using information about the material of a reflective object as an index for evaluating the perceptual importance of a reflected sound will be described. For example, multiple parameters such as a reflection coefficient (reflectance), a diffusion coefficient, a transmittance, and a sound absorption coefficient may be acquired from metadata as information about the material of the reflective object. Then, the perceptual importance of a reflected sound may be evaluated according to the ratio of each parameter.
[0304] Specifically, for example, among multiple parameters related to each material that can be set for the reflective surface of a reflective object, if the ratio of reflectivity or diffusion rate is high, the volume of the reflected sound that is reflected from that reflective surface and reaches the listener will be higher than if the ratio of transmittance or sound absorption rate is high. In this case, the perceptual importance is likely to be high. Therefore, if the ratio of reflectivity or diffusion rate is high among multiple parameters related to the material that can be set for the reflective surface of a reflective object, the evaluation value of the reflected sound reflected by that reflective object may be high.
[0305] Furthermore, the information about the material of the reflective object is not limited to the reflection coefficient (reflectance), diffusion rate, transmittance, and sound absorption rate, but may be information that can identify the importance of the material. For example, a set of multiple parameters, such as the reflection coefficient (reflectance), diffusion rate, transmittance, and sound absorption rate, may be acquired from metadata as information that identifies the material. Furthermore, the importance may be defined in advance for each material identifier. Then, an evaluation value of the reflected sound may be calculated according to the importance associated with the material identifier.
[0306] Furthermore, it is not necessary to take all of the multiple parameters into consideration, and the importance of the material of the reflective object may be determined using only some of the parameters.
[0307] Furthermore, the information for specifying (identifying) the material is not limited to information for uniquely identifying the material (material identification information), but may be, for example, information classifying the material (material classification information). The information classifying the material may be, for example, information classified according to a classification method preset by the content creator.
[0308] By using an index of the importance of a reflecting object, it becomes possible to appropriately select reflected sounds caused by reflecting objects with high importance based on the importance of the reflecting object.
[0309] Furthermore, when a reflected sound caused by a reflective object is selected, the importance of the reflective object may be updated to be lower. This makes it possible to reproduce reflected sounds caused by many reflective objects evenly without favoring reflected sounds caused by a specific reflective object (e.g., a specific wall or a specific ceiling). This therefore makes it possible to secure clues for the listener to correctly perceive the size of the sound space.
[0310] For example, when selecting 30 reflected sounds from 300 reflected sounds generated in a sound space, selecting 30 reflected sounds reflected from a specific wall surface makes it difficult to grasp the overall space. Furthermore, allocating five reflected sounds to each of the six walls is not necessarily optimal. Therefore, as described above, whether or not to select reflected sounds caused by each reflective object (e.g., a wall) may be controlled based on the value (importance) of the reflective object.
[0311] The selection unit 1302 may use the relationship between the direct sound and the reflected sound (e.g., geometric relationship) as an evaluation index for the reflected sound. Specifically, the selection unit 1302 may evaluate the geometric relationship between the direct sound and the reflected sound based on the arrival angle between the direct sound and the reflected sound, and use the evaluation as an evaluation index for the reflected sound (S305). Here, the arrival angle between the direct sound and the reflected sound corresponds to the angle between the arrival direction of the direct sound and the arrival direction of the reflected sound, and corresponds to the angular difference between the angle of the arrival direction of the direct sound relative to a reference direction and the angle of the arrival direction of the reflected sound relative to the reference direction.
[0312] The angle formed between the direction from which the direct sound arrives and the direction from which the reflected sound arrives may be detected, and the larger the angle, the higher the evaluation score given to the reflected sound.
[0313] For example, the analysis unit 1301 calculates a direct sound arrival path (pd) and a reflected sound arrival direction path (pr). The analysis unit 1301 or the selection unit 1302 calculates the arrival direction of the direct sound and the arrival direction of the reflected sound based on the direct sound arrival path (pd), the reflected sound arrival direction path (pr), and orientation information (D) of the avatar (listener) included in the input signal. The arrival direction of the direct sound and the arrival direction of the reflected sound are expressed using the orientation of the listener as a reference.
[0314] The selection unit 1302 then calculates an evaluation score for the reflected sound based on the angle formed between the direction from which the direct sound comes and the direction from which the reflected sound comes.
[0315] Fig. 13 is a diagram showing an example of the arrival angles of direct sound and reflected sound. For example, an avatar, a sound source object, and an obstacle object are arranged as shown in Fig. 13. Position information of the avatar, sound source object, and obstacle object, as well as orientation information (D) of the avatar, are obtained from the input signal. Then, from this information, the direction of the direct sound (θ) and the direction of the sound image of the reflected sound (γ) are calculated, assuming that the orientation of the avatar is 0 degrees.
[0316] 13, the direction of the direct sound (θ) is about 20 degrees, and the direction of the sound image of the reflected sound (γ) is about 265 degrees (−95 degrees). In this case, the angle between the direction of arrival of the direct sound and the direction of arrival of the reflected sound is about 115 degrees.
[0317] If the angle between the direction of arrival of the direct sound and the direction of arrival of the reflected sound is large, the reflected sound is assigned a high score. As a result, for example, the reflected sound of a sound emitted from a sound source visible in front of the listener but heard from behind the listener is assigned a high score. As a result, it is possible to preferentially select reflected sounds that help the listener predict the presence of a large object behind the listener, thereby creating a sense of claustrophobia and tension.
[0318] The selection unit 1302 may evaluate the relationship between the direct sound and the reflected sound based on the time difference between the direct sound and the reflected sound, and use the evaluation as an evaluation index for the reflected sound (S306). For example, the selection unit 1302 may assign a higher evaluation score to a reflected sound with a large difference in arrival time between the direct sound and the reflected sound than to a reflected sound with a small difference in arrival time. For example, the echo that is returned when shouting "Yahoo!" from the top of a mountain has a decisive impact on spatial perception. Therefore, such a reflected sound may be assigned a higher evaluation score.
[0319] The selection unit 1302 may evaluate the relationship between the direct sound and the reflected sound using the time difference between the direct sound and the reflected sound and a threshold value corresponding to the time difference. For example, a reflected sound that arrives at the listener's position immediately after the direct sound is likely to be masked by the direct sound and is therefore difficult to perceive. On the other hand, a reflected sound that arrives at the listener's position with a time lag from the direct sound is likely to be masked by the direct sound and is therefore easy to perceive. An evaluation score may be assigned to the reflected sound based on such a perception model.
[0320] The time difference (T) between the direct sound and the reflected sound may be, for example, the time difference between the time it takes for the direct sound and the reflected sound to reach the listening position. For example, the time difference (T) between the time it takes for the direct sound and the reflected sound to reach the listening position can be calculated as T = tr - td.
[0321] For example, when a reflected sound is evaluated using the relationship between a direct sound and a reflected sound, a comparison process is performed using a threshold determined in accordance with the time difference between the direct sound and the reflected sound. The threshold indicates a volume that is preset in accordance with the time difference between the direct sound and the reflected sound, and is determined by referring to threshold data. The threshold data may use an index that indicates the boundary between whether a reflected sound relative to the direct sound is perceptible by a listener.
[0322] For example, the threshold value refers to a value expressed by a numerical value or the like determined corresponding to the time difference (T), and the threshold value data refers to table data or a relational expression used to identify or calculate the threshold value for the time difference (T). However, the format and type of the threshold value data are not limited to table data or a relational expression.
[0323] 14 is a diagram showing an example of a method for setting threshold data based on the temporal masking phenomenon. The threshold data may be set, for example, by referring to a masking threshold, which is a known threshold. The temporal masking phenomenon is widely known, as described in Non-Patent Document 1 and the like. The shaded area in the diagram indicates the time period during which a masker (an interfering signal that interferes with the perception of the signal S to be heard) occurs and its amplitude.
[0324] In Fig. 14, the masking threshold indicates the audible level (SPL: Sound Pressure Level) of the signal S. Naturally, the masking threshold is high while the masker is occurring. On the other hand, even after the masker stops, the masking threshold does not immediately become zero, but gradually decays. In other words, the masking threshold remains high for a while (the period during which post-masking exists) immediately after the masker stops.
[0325] For example, the post-masking tendency shown in the area surrounded by the dotted line in Fig. 14 may be used as threshold data for evaluating reflected sound based on the relationship between direct sound and reflected sound. In other words, the threshold data may be determined based on the post-masking tendency, assuming that the direct sound corresponds to the Masker and the reflected sound corresponds to the signal S to be heard.
[0326] Fig. 15 is a diagram showing an example of threshold data. In the above case, the threshold data may be determined as shown in the curve in Fig. 15. In Fig. 15, in a graph having the time difference between direct sound and reflected sound on the horizontal axis and the volume of the reflected sound on the vertical axis, the boundary (threshold) at which the reflected sound is perceived or not is shown by a curve. The curve corresponds to the threshold data.
[0327] The threshold data according to this embodiment is stored in the memory 1404 of the audio signal processing device 1001. The stored threshold data may be in any format and type. For example, the threshold data may be expressed by an approximation formula having the time difference between a direct sound and a reflected sound as a variable. Alternatively, the threshold data may be expressed as an array of the time difference between a direct sound and a reflected sound and a threshold.
[0328] 16 is a diagram showing the relationship between the time difference between a direct sound and a reflected sound and a threshold value. As shown in FIG. 16, the threshold value data may be stored in an area of memory 1404 as an array of indexes of the time difference between a direct sound and a reflected sound and threshold values corresponding to the indexes.
[0329] Of course, the graphs and numerical values shown in FIGS. 15 and 16 are merely examples, and the threshold data is not limited to these.
[0330] The memory 1404 may store information regarding a relational expression showing the relationship between the time difference (T) and the threshold value. That is, an expression having the time difference (T) as a variable may be stored. The threshold value of each time difference (T) may be approximated by a straight line or a curve, and parameters indicating the geometric shape of the line or curve may be stored. For example, if the geometric shape is a straight line, a starting point and a slope for expressing the straight line may be stored.
[0331] When multiple types and types of thresholds are stored, it may be determined which type and type of threshold to use in the process of selecting reflected sounds.
[0332] Furthermore, the threshold for evaluating reflected sound is not limited to a known masking threshold. Other thresholds may be determined based on the time difference between the direct sound and the reflected sound and a value indicating the amplitude or volume. For example, the threshold may be determined based on the minimum time difference at which a listener perceptually detects a discrepancy between the two sounds. Specific numerical values may be derived from known research results or determined through listening experiments conducted with the assumption that the method will be applied to the virtual space.
[0333] For example, the threshold value is set by referring to threshold data based on the time difference between the arrival time of the direct sound and the arrival time of the reflected sound. If the volume of the reflected sound at the time of arrival is greater than the set threshold value, the selection unit 1302 may increase the evaluation score.
[0334] The time difference between the arrival time of the direct sound and the arrival time of the reflected sound is, in other words, the difference in the time it takes for the direct sound and the reflected sound to arrive at the listening position. Therefore, the difference in the distance of the arrival path of the direct sound and the reflected sound may be used as a value related to the time difference between the arrival times of the direct sound and the reflected sound.
[0335] Alternatively, the time difference between the end of the direct sound and the arrival of the reflected sound at the listening position may be used as the time difference between the direct sound and the reflected sound. Here, the end time of the direct sound may be calculated by adding the duration of the direct sound to the arrival time of the direct sound, for example.
[0336] A method for evaluating reflected sounds using threshold data will be described with reference to FIGS. 9 and 10 showing the positional relationship between the listener and an obstacle object, and FIG. 15 showing an example of threshold data.
[0337] The graph in Figure 15 has the time difference between direct sound and reflected sound on the horizontal axis, and the volume ratio between direct sound and reflected sound on the vertical axis. The curve represents the threshold at which reflected sound is perceived or not. A, B, and C in the graph each represent reflected sound. Note that here, the vertical axis uses the volume ratio, i.e., the volume of reflected sound determined relatively to the volume of direct sound, but it is also possible to use the volume of reflected sound determined absolutely, regardless of the volume of direct sound.
[0338] It goes without saying that when the volume is expressed in decibel units on a logarithmic axis (when the volume is expressed in the decibel domain), the volume ratio of two signals is expressed as the difference in decibel values. Specifically, the volume ratio of two signals may be the difference between the amplitude values of each signal when expressed in the decibel domain. This value may be calculated based on an energy value, a power value, or the like. Furthermore, in the decibel domain, this difference may be referred to as a gain difference or simply a gain difference.
[0339] That is, the volume ratio in the present disclosure is essentially a ratio of signal amplitudes, and may be expressed as a sound volume ratio, a volume ratio, an amplitude ratio, a sound level ratio, a sound intensity ratio, a gain ratio, etc. Furthermore, when the unit of volume is decibels, the volume ratio in the present disclosure can of course be rephrased as a volume difference.
[0340] In the present disclosure, the term "volume ratio" typically refers to the gain difference when the volume of two sounds is expressed in decibel units, and in the example embodiments, the threshold data is also typically defined as a gain difference expressed in the decibel domain. However, the volume ratio is not limited to a gain difference in the decibel domain. When a volume ratio expressed in a domain other than the decibel domain is used, the threshold data defined in the decibel domain may be converted into the unit of the calculated volume ratio and used. Alternatively, threshold data defined in each unit may be stored in advance in memory.
[0341] In other words, it is clear that the algorithm in the present disclosure can be applied to solving the problem of the present disclosure even if a ratio of energy values or power values, for example, is used instead of the volume ratio.
[0342] Fig. 9 shows the positional relationship between a listener, a sound source object, and an obstacle object (wall). In Fig. 9, the sound source object and the obstacle object are relatively far away, and the listener hears the reflected sound C in Fig. 15. Fig. 10 shows another positional relationship between the listener, the sound source object, and the obstacle object (wall). In Fig. 10, the sound source object and the obstacle object are relatively close, and the listener hears the reflected sound A or B in Fig. 15.
[0343] For example, as shown in Figure 9, when the listener is relatively far from the obstacle object, the arrival time of the reflected sound is delayed, and the time difference between the direct sound and the reflected sound of reflected sound C is larger than that of reflected sounds A and B.
[0344] In other words, as shown in Figure 15, reflected sound C is located to the right of reflected sounds A and B on the graph. As shown by the curve in the graph, the greater the time difference between the direct sound and the reflected sound, the smaller the threshold value. As a result, reflected sound B, which has the same volume as reflected sound C, is smaller than the threshold value, and reflected sound C is larger than the threshold value. Therefore, the evaluation score of reflected sound C is higher than the evaluation score of reflected sound B.
[0345] Furthermore, for reflected sounds A and B, the arrival time is the same, but the volume of reflected sound A is louder than the volume of reflected sound B, which is in turn quieter than the volume of reflected sound A. Furthermore, the volume of reflected sound A is louder than the threshold indicated by the curve, and the volume of reflected sound B is quieter than the threshold indicated by the curve. In this case, reflected sound A is assigned a higher evaluation score than reflected sound B.
[0346] The reflected sound is evaluated based on a threshold value indicating the volume determined in accordance with the time difference between the direct sound and the reflected sound. This allows the evaluation of the reflected sound to reflect the nature of human perception, namely, that reflected sound that arrives at the listener's position with a time lag from the direct sound is not masked by the direct sound and is therefore easily perceived.
[0347] By using a threshold value indicating the volume that is determined in accordance with the time difference between the direct sound and the reflected sound, it becomes possible to more appropriately select the reflected sound that has a greater impact on the listener's perception than using only the time difference between the direct sound and the reflected sound or only the volume of the reflected sound.
[0348] Alternatively, calculation of the arrival times and arrival volumes of the direct sound and the reflected sound may be omitted, and the reflected sound may be evaluated based on the path lengths. When the reflected sound is evaluated based on the path lengths of the direct sound and the reflected sound when they reach the listener, a threshold value for the path length of the reflected sound may be set corresponding to the value of the path length difference. In this case, the reflected sound may be evaluated based on whether the path length of the reflected sound is greater than the threshold value set corresponding to the value of the path length difference.
[0349] When the selection process is performed based on the path length, it is possible to perform the selection process based on information that affects the time difference while reducing the amount of calculation compared to when the selection process is performed based on the time difference. In addition to the difference in path length, a parameter that indicates the sound propagation speed or a parameter that affects the sound propagation speed parameter may be used.
[0350] The geometric relationship may be the relationship between the positions of the sound source, the listener, and the reflecting object in the virtual space. These relationships allow the geometric calculation of the path lengths of the direct sound and the reflected sound. Therefore, by utilizing the relationship in which the volume is inversely proportional to the distance, it is possible to calculate the reference volume of the reflected sound relative to the reference volume of the direct sound.
[0351] The reference volume of the reflected sound may be calculated using the reflection coefficient of the reflecting object. A commonly used typical value may also be used as the reflection coefficient. On the other hand, if a special condition exists, such as the reflecting object being covered with a sound-absorbing material, a specially assigned reflection coefficient may be used as the reflection coefficient of the reflecting object.
[0352] The reflected sound may be evaluated based on its volume, which may be calculated from the geometric relationship between the direct sound and the reflected sound and the index assigned to the reflective object, as described above, and may be evaluated by comparing the volume with a predetermined threshold.
[0353] Furthermore, information indicating the temporal transition of the volume of the sound source may be reflected in the evaluation. For example, if the information indicating the temporal transition of the volume of the sound source indicates the duration of a sound section, and the time is within the sound section, the evaluation value of the reflected sound may be maintained as is. On the other hand, if the time is outside the sound section, processing may be performed to reduce or set the evaluation value of the reflected sound to zero even if the reference volume of the reflected sound exceeds the threshold.
[0354] Alternatively, the information indicating the temporal transition of the volume of the sound source may be data listing, in time series, multiple pairs of durations during which the amplitude of a sound signal is considered to be roughly constant and the amplitude values of the signal during those durations. In this case, the reference volume of the reflected sound may be changed in conjunction with changes in the amplitude values in the data to evaluate the reflected sound.
[0355] Note that evaluation points for all of the above-mentioned indices may be assigned to one reflected sound, or evaluation points for only some of the indices may be assigned. Furthermore, the number of indices used for evaluation may vary for each reflected sound, or the same indices may be used for all reflected sounds. Which indices are used to assign evaluation points to reflected sounds may be set based on predetermined information, and may be determined based on information included in the input signal, or may be determined based on information set by a listener or an administrator, for example.
[0356] Furthermore, a high evaluation score corresponds to a large evaluation score, and a low evaluation score corresponds to a small evaluation score. Similarly, a high evaluation value corresponds to a large evaluation value, and a low evaluation value corresponds to a small evaluation value. These expressions may be interchangeable.
[0357] (Calculation of Evaluation Value) Next, for the reflected sounds to which evaluation points have been assigned using each index, the selection unit 1302 calculates an evaluation value indicating the importance of the reflected sound based on the evaluation points. For example, the selection unit 1302 determines the sum of multiple evaluation points as the evaluation value of the reflected sound (S307). The sum of multiple evaluation points may be a weighted sum. If there are unevaluated reflected sounds (Yes in S308), the selection unit 1302 repeats the above-described processing (S301 to S307), and if there are no unevaluated reflected sounds (No in S308), the selection unit 1302 ends the evaluation process.
[0358] The evaluation value of a reflected sound is not limited to the sum of multiple evaluation points obtained using multiple indices. For example, a predetermined reference evaluation value and an already calculated evaluation value may be corrected using multiple evaluation points. Also, only the evaluation points of some of the indices may be used for the evaluation value of the reflected sound, or may be used to correct the evaluation value of the reflected sound. Furthermore, when multiple evaluation points are assigned to one reflected sound using multiple indices, the highest evaluation point may be determined as the evaluation value of the reflected sound.
[0359] The evaluation score of each index to be used for calculating or correcting the evaluation value may be determined based on predetermined information, may be determined based on information contained in the input signal, or may be determined based on information set by the listener or administrator.
[0360] In the above description, the evaluation points and evaluation values are divided into evaluation points obtained by each index and evaluation values obtained using multiple evaluation points obtained by multiple indexes for convenience. However, since both indicate the evaluation results of reflected sounds, the evaluation points and evaluation values may be treated in the same way. Furthermore, the audio signal processing device 1001 may use an evaluation point obtained by one index as an evaluation value as it is in the selection process of reflected sounds, or may use multiple evaluation points obtained by multiple indexes in the selection process of reflected sounds.
[0361] For example, when multiple evaluation points are used in the reflected sound selection process, the audio signal processing device 1001 determines whether to select a reflected sound based on each of the multiple evaluation points.The audio signal processing device 1001 may then ultimately determine that a reflected sound should be selected when all of the multiple determination results based on the multiple evaluation points indicate that a reflected sound should be selected.Alternatively, the audio signal processing device 1001 may ultimately determine that a reflected sound should be selected when any one of the multiple determination results based on the multiple evaluation points indicates that a reflected sound should be selected.
[0362] Furthermore, priorities may be assigned to multiple evaluation points based on multiple indices. For example, the audio signal processing device 1001 determines whether to select a reflected sound based on each of the first to third evaluation points based on the first to third indices.
[0363] In the above case, when the judgment result based on the first evaluation point indicates that the reflected sound will not be selected, the audio signal processing device 1001 may ultimately judge that the reflected sound will not be selected without relying on the judgment results based on the second and third evaluation points.
[0364] Furthermore, when the judgment results based on the first and second evaluation points indicate that reflected sound should be selected, the audio signal processing device 1001 may ultimately judge that reflected sound should be selected without relying on the judgment results based on the third evaluation point.
[0365] For example, after the evaluation value of the reflected sound is determined, the processing is carried out as described above according to the flowchart shown in FIG.
[0366] (Processing Order and Omission) Some of the processes included in the flowcharts shown in FIGS. 11 and 12 may be omitted, or the processing order may be changed.
[0367] For example, in the flowchart shown in Fig. 11, the selection process for the reflected sounds is performed based on both the calculation load and the evaluation value (importance) of the reflected sounds. However, the selection process for the reflected sounds may be performed based on only one of them.
[0368] Specifically, the selection unit 1302 may omit calculation of the evaluation value of each reflected sound, and may determine not to select the reflected sound if the computational load of the reflected sound is greater than a threshold. Alternatively, the selection unit 1302 may omit obtaining information indicating the upper limit of the computational load, calculating the total computational load of the extracted reflected sounds, and comparing the total computational load with the upper limit of the computational load, and may perform the selection process of the reflected sounds based only on the evaluation value of the reflected sound.
[0369] Furthermore, extraction of reflected sounds whose volume is equal to or greater than a threshold may be performed after the evaluation value is determined, or after it is determined that a reflected sound is to be selected. For example, even if it is determined that a reflected sound is to be selected based on the evaluation value or the calculation load, if the volume of the reflected sound is below the threshold, the reflected sound may be redetermined not to be selected.
[0370] (Generation of direct sound and reflected sound) Next, in the process of generating direct sound and reflected sound (S103 in Figure 8), the synthesis unit 1303 generates and synthesizes an audio signal of the direct sound and an audio signal of the reflected sound selected by the selection unit 1302 as the reflected sound to be generated.
[0371] The audio signal of the direct sound is generated by applying the arrival time (td) and arrival volume (ld) calculated by the analysis unit 1301 to the sound data of the sound source object included in the input information. Specifically, the sound data is delayed by the arrival time (td) and multiplied by the arrival volume (ld). The process of delaying the sound data is a process of moving the position of the sound data forward or backward on the time axis. For example, a process of delaying sound data without degrading sound quality, as disclosed in Patent Document 2, may be applied.
[0372] The audio signal of the reflected sound is generated by applying the arrival time (tr) and arrival volume (ld) calculated by the analysis unit 1301 to the sound data of the sound source object, just like the direct sound.
[0373] However, unlike the volume of direct sound arriving at the time of arrival, the volume of arrival (lr) when generating reflected sound is a value to which an attenuation rate G of the volume of reflection is applied. G may be an attenuation rate applied to all frequency bands at once. Alternatively, a reflectance rate may be specified for each predetermined frequency band to reflect the bias in frequency components caused by reflection. In this case, the process of applying the volume of arrival (lr) may be performed as a frequency equalizer process, which multiplies each band by an attenuation rate.
[0374] (Pipeline Processing) The processing performed by the above-described analysis unit 1301, selection unit 1302, and synthesis unit 1303 may be performed as pipeline processing as described in, for example, Patent Document 3.
[0375] FIG. 17 is a block diagram showing an example of the configuration for the rendering unit 1300 to perform pipeline processing.
[0376] 17 includes a reverberation processing unit 1311, an early reflection processing unit 1312, a distance attenuation processing unit 1313, a selection unit 1314, a generation unit 1315, and a binaural processing unit 1316. The reverberation processing unit 1311, the early reflection processing unit 1312, and the distance attenuation processing unit 1313 perform reverberation processing, early reflection processing, and distance attenuation processing, respectively. The selection unit 1314 selects reflected sounds, the generation unit 1315 generates direct sounds and reflected sounds, and the binaural processing unit 1316 applies binaural processing to the direct sounds and reflected sounds.
[0377] These multiple components may be composed of multiple components of the rendering unit 1300 shown in Figure 7, or may be composed of at least some of the multiple components of the audio signal processing device 1001 shown in Figure 5.
[0378] Pipeline processing refers to dividing the process for applying sound effects into multiple processes and executing the multiple processes one by one in sequence. Each of the multiple processes performs, for example, signal processing on an audio signal or generation of parameters used in the signal processing.
[0379] The rendering unit 1300 may perform reverberation processing, early reflection processing, distance attenuation processing, binaural processing, and the like as pipeline processing. However, these processes are merely examples, and the pipeline processing may include other processes or may not include some of the processes. For example, the pipeline processing may include diffraction processing and occlusion processing. Furthermore, for example, reverberation processing may be omitted if it is not necessary.
[0380] Each process may be expressed as a stage. An audio signal such as a reflected sound generated as a result of each process may be expressed as a rendering item. The multiple stages in the pipeline process and their order are not limited to the example shown in FIG. 17 .
[0381] Here, the parameters used in the selection process (arrival paths and arrival times for direct sound and reflected sound) are calculated in one of the multiple stages for generating a rendering item. In other words, the parameters used to select reflected sound are calculated as part of the pipeline processing for generating a rendering item. Note that not all stages need to be performed by the rendering unit 1300. For example, some stages may be omitted or may be performed by a unit other than the rendering unit 1300.
[0382] The following describes reverberation processing, early reflection processing, distance attenuation processing, selection processing, generation processing, and binaural processing that may be included as stages in the pipeline processing. At each stage, metadata included in the input signal may be analyzed to calculate parameters used to generate reflected sounds.
[0383] In the reverberation processing, the reverberation processor 1311 generates an audio signal indicating a reverberant sound or parameters used to generate an audio signal. A reverberant sound is a sound that arrives at a listener as reverberation after a direct sound. As an example, a reverberant sound is a sound that arrives at a listener after a relatively late stage (e.g., about 150 ms after the arrival of the direct sound) after an early reflected sound (described later) arrives at the listener, and after having been reflected more times (e.g., several tens of times) than an early reflected sound.
[0384] The reverberation processor 1311 refers to the audio signal and spatial information contained in the input signal, and calculates the reverberation sound using a predetermined function prepared in advance as a function for generating the reverberation sound.
[0385] The reverberation processor 1311 may generate reverberant sounds by applying a known reverberation generation method to the audio signal included in the input signal. An example of a known reverberation generation method is the Schroeder method, but known reverberation generation methods are not limited to the Schroeder method. Furthermore, when applying a known reverberation generation method, the reverberation processor 1311 uses the shape and acoustic characteristics of the sound reproduction space indicated by the spatial information. This allows the reverberation processor 1311 to calculate parameters for generating reverberant sounds.
[0386] In the early reflection process, the early reflection processor 1312 calculates parameters for generating early reflection sounds based on spatial information. The early reflection sounds are reflected sounds that arrive at the listener after one or more reflections at a relatively early stage after a direct sound from a sound source object arrives at the listener (for example, about several tens of milliseconds after the direct sound arrives).
[0387] The early reflection processing unit 1312 refers to, for example, the audio signal and metadata, and calculates the path of the reflected sound that travels from the sound source object to the listener after being reflected by the reflecting object. For example, the path calculation may use the shape of the three-dimensional sound field (space), the size of the three-dimensional sound field, the positions of reflecting objects such as structures, and the reflectance of the reflecting object.
[0388] The early reflection processing unit 1312 may also calculate the path of the direct sound. Information about the path may be used as a parameter by which the early reflection processing unit 1312 generates the early reflected sound, or may be used as a parameter by which the selection unit 1314 selects the reflected sound.
[0389] In the distance attenuation process, the distance attenuation processor 1313 calculates the volume of the direct sound and the reflected sound that reach the listener based on the path lengths of the direct sound and the reflected sound. The volume of the direct sound and the reflected sound that reach the listener attenuates in proportion to the distance of the path to the listener (inversely proportional to the distance) relative to the volume of the sound source. Therefore, the distance attenuation processor 1313 can calculate the volume of the direct sound by dividing the volume of the sound source by the path length of the direct sound, and can calculate the volume of the reflected sound by dividing the volume of the sound source by the path length of the reflected sound.
[0390] In the selection process, the selection unit 1314 selects a generation target reflected sound based on parameters calculated before the selection process. Any of the selection methods disclosed herein may be used to select the generation target reflected sound.
[0391] Furthermore, when the selection process is included in the pipeline process, the processing after the selection process may not be executed for the reflected sounds that are not selected in the selection process. By not executing the processing after the selection process for the reflected sounds that are not selected, it is possible to reduce the computational load of the audio signal processing device 1001 more than when only the binaural processing is not executed.
[0392] Furthermore, when the selection process is included in the pipeline process, by assigning an earlier order to the selection process among the multiple processes in the pipeline process, it becomes possible to omit more processes and reduce the amount of calculations more.
[0393] In the binaural processing, the binaural processing unit 1316 performs signal processing so that the audio signal of the direct sound is perceived by the listener as a sound arriving from the direction of the sound source object. Furthermore, the binaural processing unit 1316 performs signal processing so that the reflected sound selected by the selection unit 1314 is perceived by the listener as a sound arriving from the reflecting object.
[0394] For example, the binaural processing unit 1316 performs processing to apply the HRIR DB based on the position and orientation of the listener in the sound space so that sound arrives at the listener from the position of a sound source object or the position of an obstacle object.
[0395] HRIR (Head-Related Impulse Responses) is a response characteristic when one impulse is generated. Specifically, HRIR is a response characteristic obtained by converting a head-related transfer function, which represents changes in sound caused by surrounding objects including the auricle, the human head, and shoulders, from a frequency domain representation to a time domain representation by Fourier transform. The HRIR DB is a database containing such information.
[0396] Furthermore, the position and orientation of the listener in the sound space are, for example, the position and orientation of the virtual listener in the virtual sound space. The position and orientation of the virtual listener in the virtual sound space may change in accordance with the movement of the listener's head. The position and orientation of the virtual listener in the virtual sound space may also be determined based on information acquired from the sensor 1405.
[0397] The programs, spatial information, HRIR DB, threshold data, and other parameters used in the above processing are obtained from the memory 1404 provided in the audio signal processing device 1001 or from outside the audio signal processing device 1001.
[0398] The pipeline processing may also include other processes. The rendering unit 1300 may also include processing units (not shown) for performing other processes included in the pipeline processing. For example, the rendering unit 1300 may include a diffraction processing unit and an occlusion processing unit.
[0399] The diffraction processing unit executes processing to generate an audio signal representing a sound including diffracted sound caused by an obstacle object between the listener and the sound source object in a three-dimensional sound field (space). When an obstacle object exists between the sound source object and the listener, the diffracted sound is a sound that travels from the sound source object to the listener, going around the obstacle object.
[0400] The diffraction processing unit calculates a path of the diffracted sound from the sound source object to the listener, bypassing the obstacle object, and generates the diffracted sound based on the path, for example, by referring to the audio signal and metadata. The path calculation may use the positions of the sound source object, the listener, and the obstacle object in the three-dimensional sound field (space), as well as the shape and size of the obstacle object.
[0401] When a sound source object is present on the other side of an obstacle object, the occlusion processing unit generates an audio signal of sound that leaks from the sound source object and passes through the obstacle object based on spatial information and information such as the material of the obstacle object.
[0402] (Example of a Sound Source Object) In the above, the position information assigned to the sound source object indicates a "point" in the virtual space as the position of the sound source object. That is, in the above, the sound source is defined as a "point sound source."
[0403] On the other hand, a sound source in a virtual space may be defined as an object having length, size, shape, etc., i.e., as a spatially extended sound source rather than a point sound source. In this case, the distance between the listener and the sound source and the direction from which the sound is coming are not determined. Therefore, reflected sounds caused by such sound sources may be limited to those selected by the selection unit 1302 without being analyzed by the analysis unit 1301 or regardless of the analysis results. This makes it possible to avoid deterioration in sound quality that may occur when reflected sounds are not selected.
[0404] Alternatively, a representative point such as the center of gravity of the object may be determined, and the processing of the present disclosure may be applied on the assumption that the sound is generated from that representative point. In this case, the threshold may be adjusted according to information on the spatial extent of the sound source.
[0405] (Examples of Direct Sound and Reflected Sound) For example, direct sound is sound that is not reflected by a reflecting object, and reflected sound is sound that is reflected by a reflecting object. Direct sound may be sound that arrives at the listener from a sound source without being reflected by a reflecting object, or reflected sound may be sound that arrives at the listener from a sound source after being reflected by a reflecting object.
[0406] Furthermore, the direct sound and the reflected sound are not limited to sounds that have arrived at the listener, but may be sounds that have not yet arrived at the listener. For example, the direct sound may be sounds output from a sound source, or in other words, sounds from the sound source.
[0407] (Example of Bitstream Structure) A bitstream includes, for example, an audio signal and metadata. The audio signal is sound data that expresses sound, and indicates information about the frequency and intensity of the sound. The metadata includes spatial information about the sound space, which is the space of the sound field.
[0408] For example, the spatial information is information about a space in which a listener who listens to a sound based on an audio signal is located. Specifically, the spatial information is information about a predetermined position (localization position) for localizing a sound image at a predetermined position in a sound space (e.g., a three-dimensional sound field), that is, for allowing the listener to perceive a sound arriving from a direction corresponding to the predetermined position. The spatial information includes, for example, sound source object information and position information indicating the position of the listener.
[0409] The sound source object information is information about a sound source object that generates a sound based on an audio signal. That is, the sound source object information is information about an object (sound source object) that reproduces an audio signal, and is information about a virtual sound source object that is placed in a virtual sound space. Here, the virtual sound space may correspond to a real space in which an object that generates a sound is placed, and the sound source object in the virtual sound space may correspond to an object that generates a sound in the real space.
[0410] The sound source object information may indicate the position of the sound source object arranged in the sound space, the orientation of the sound source object, the directivity of the sound emitted by the sound source object, whether the sound source object belongs to a living thing or not, whether the sound source object is a moving object or not, etc. For example, the audio signal is associated with one or more sound source objects indicated by the sound source object information.
[0411] The bitstream has a data structure that is made up of, for example, metadata (control information) and an audio signal.
[0412] The audio signal and metadata may be contained in a single bitstream or in separate bitstreams, or may be contained in a single file or in separate files.
[0413] A bitstream may exist for each sound source or for each playback time. Even if a bitstream exists for each playback time, multiple bitstreams may be processed in parallel at the same time.
[0414] Metadata may be assigned to each bitstream, or may be assigned to multiple bitstreams together as information for controlling multiple bitstreams. In this case, multiple bitstreams may share the same metadata. Metadata may also be assigned for each playback time.
[0415] When multiple bitstreams or multiple files exist, one or more of the bitstreams or files may contain information indicating the associated bitstreams or files, or alternatively, each of all of the bitstreams or each of all of the files may contain information indicating the associated bitstreams or files.
[0416] Here, the related bitstreams or related files are, for example, bitstreams or files that may be used simultaneously during audio processing, and may also include bitstreams or files that collectively describe information indicating related bitstreams or related files.
[0417] Here, the information indicating the related bitstream or related file may be, for example, an identifier indicating the related bitstream or related file. Alternatively, the information indicating the related bitstream or related file may be, for example, a file name indicating the related bitstream or related file, a URL (Uniform Resource Locator), or a URI (Uniform Resource Identifier).
[0418] In this case, the acquisition unit identifies and acquires the related bitstream or related file based on the information indicating the related bitstream or related file. Alternatively, the bitstream or file may contain information indicating the related bitstream or related file, and another bitstream or another file may contain information indicating the related bitstream or related file.
[0419] Here, the file containing information indicating the associated bitstream or associated file may be a control file such as a manifest file used for content distribution.
[0420] Note that all or part of the metadata may be obtained from sources other than the bitstream of the audio signal. For example, either the metadata for controlling the sound or the metadata for controlling the video may be obtained from sources other than the bitstream, or both may be obtained from sources other than the bitstream.
[0421] Furthermore, metadata for controlling the video may be included in the bitstream acquired by the stereophonic sound reproduction system 1000. In this case, the stereophonic sound reproduction system 1000 may output the metadata for controlling the video to a display device that displays images or a stereophonic video reproduction device that reproduces the stereophonic video.
[0422] (Examples of Information Included in Metadata) Metadata may be information used to describe a scene expressed in sound space, where a scene is a term that refers to a collection of all elements that represent three-dimensional video and sound events in sound space that are modeled by the stereophonic sound reproduction system 1000 using the metadata.
[0423] That is, the metadata may include not only information for controlling audio processing but also information for controlling video processing. The metadata may include only one of information for controlling audio processing and information for controlling video processing, or may include both.
[0424] The stereophonic sound reproduction system 1000 generates virtual sound effects by performing sound processing on audio signals using metadata included in the bitstream and additionally acquired interactive listener position information, etc. Among the sound effects, early reflection processing, obstacle processing, diffraction processing, blocking processing, and reverberation processing may be performed, and other sound processing may be performed using the metadata. For example, sound effects such as distance attenuation, localization, or Doppler effect may be added.
[0425] Furthermore, information on switching on / off all or part of the sound effects, or priority information for multiple sound effect processes may be added to the metadata.
[0426] As an example, the metadata includes information about a sound space including sound source objects and obstacle objects, and information about a positioning position for localizing a sound image at a predetermined position within the sound space (i.e., allowing the listener to perceive sound coming from a predetermined direction).
[0427] Here, an obstacle object is an object that may affect the sound perceived by the listener by, for example, blocking or reflecting the sound emitted by the sound source object before it reaches the listener. Obstacle objects may include not only stationary objects but also moving objects such as animals or machines. The animal may also be a person, etc.
[0428] Furthermore, when multiple sound source objects exist in a sound space, other sound source objects can be obstacle objects for any sound source object. In other words, both non-sound-emitting objects, such as building materials or inanimate objects, which do not emit sound, and sound source objects that emit sound can be obstacle objects.
[0429] The metadata includes information that represents all or part of the shape of the sound space, the shape and position of obstacle objects in the sound space, the shape and position of sound source objects in the sound space, and the position and orientation of the listener in the sound space.
[0430] The sound space may be either a closed space or an open space. The metadata may also include information indicating the reflectance of obstacle objects that may reflect sound in the sound space. For example, the floor, walls, or ceiling that form the boundaries of the sound space may also constitute obstacle objects.
[0431] The reflectance is the ratio of the energy of reflected sound to incident sound, and may be set for each frequency band of sound. Of course, the reflectance may be set uniformly regardless of the frequency band of sound. Note that when the sound space is an open space, parameters such as a uniform attenuation rate, diffracted sound, and early reflected sound may be used.
[0432] The metadata may include information other than reflectance as a parameter related to an obstacle object or a sound source object. For example, the metadata may include information related to the material of the object as a parameter related to both a sound source object and a non-sound-producing object. Specifically, the metadata may include information such as diffusion rate, transmittance, and sound absorption rate.
[0433] The information about the sound source object may include information indicating the volume, radiation characteristics (directivity), playback conditions, the number and type of sound sources in one object, and the sound source area in the object. The playback conditions may determine, for example, whether the sound is a continuous sound or an event-triggering sound. The sound source area in the object may be determined based on the relative relationship between the position of the listener and the position of the object, or may be determined using the object as a reference.
[0434] For example, if the sound source area is defined relative to the position of the listener and the position of the object, it is possible for the listener to perceive sound A coming from the right side of the object and sound B coming from the left side of the object.
[0435] Furthermore, when a sound source region is defined using an object as a reference, it is possible to fix which region of the object will emit which sound. For example, when a listener views an object from the front, it is possible for the listener to perceive a high-pitched sound from the right side of the object and a low-pitched sound from the left side of the object. When a listener views an object from the back, it is possible for the listener to perceive a low-pitched sound from the right side of the object and a high-pitched sound from the left side of the object.
[0436] The spatial metadata may include the time to early reflections, the reverberation time, the ratio of direct sound to diffuse sound, etc. If the ratio of direct sound to diffuse sound is zero, the listener will perceive only direct sound.
[0437] (Supplementary Note) The aspects grasped based on the present disclosure are not limited to the embodiments, and may be implemented with various modifications.
[0438] For example, a process performed by a specific component in the embodiment may be performed by another component instead of the specific component. Also, the order of multiple processes may be changed, or multiple processes may be performed in parallel.
[0439] Furthermore, ordinal numbers such as first and second used in the description may be rearranged, removed, or newly added as appropriate. These ordinal numbers do not necessarily correspond to a meaningful order, but may be used to identify elements.
[0440] Furthermore, for example, in comparison with a threshold, "greater than or equal to the threshold" and "greater than the threshold" may be interpreted interchangeably. Similarly, "equal to or less than the threshold" and "smaller than the threshold" may be interpreted interchangeably. Furthermore, for example, "time" and "hour" may be interpreted interchangeably.
[0441] Furthermore, in the process of selecting one or more processing target sounds from a plurality of sounds, if there is no sound that satisfies the conditions, then none of the sounds may be selected as the processing target sounds. In other words, the process of selecting one or more processing target sounds from a plurality of sounds may include cases in which no processing target sounds are selected.
[0442] Also, for example, reference to at least one of a first element, a second element, and a third element may correspond to the first element, the second element, the third element, or any combination thereof.
[0443] Furthermore, for example, in the embodiments, the cases where the aspects understood based on the present disclosure are implemented as an audio processing device, an encoding device, or a decoding device are described. However, the aspects understood based on the present disclosure are not limited to these, and may be implemented as software for executing an audio processing method, an encoding method, or a decoding method.
[0444] For example, a program for executing the above-described acoustic processing method, encoding method, or decoding method may be stored in advance in a ROM, and the CPU may operate in accordance with the program.
[0445] Furthermore, a program for executing the above-described acoustic processing method, encoding method, or decoding method may be stored in a computer-readable recording medium, and the computer may then record the program stored in the recording medium into its RAM and operate in accordance with the program.
[0446] Each of the above components may be realized as an LSI, which is typically an integrated circuit having input and output terminals. These may be individually integrated into a single chip, or a single chip may include all or some of the components of the embodiments. The LSI may be expressed as an IC, a system LSI, a super LSI, or an ultra LSI depending on the degree of integration.
[0447] Furthermore, the present invention is not limited to LSIs, and dedicated circuits or general-purpose processors may also be used. Furthermore, FPGAs, which can be programmed after LSI manufacturing, or reconfigurable processors, which allow the connection or settings of circuit cells within the LSI to be reconfigured, may also be used. Furthermore, if an integrated circuit technology that replaces LSIs emerges due to advances in semiconductor technology or other derived technologies, that technology may naturally be used to integrate components. The application of biotechnology, etc., is also a possibility.
[0448] Furthermore, the FPGA, CPU, etc. may download all or part of the software for realizing the acoustic processing method, encoding method, or decoding method described in the present disclosure via wireless or wired communication. Furthermore, all or part of the software for updating may be downloaded via wireless or wired communication. The FPGA, CPU, etc. may then store the downloaded software in memory and operate based on the stored software to perform the digital signal processing described in the present disclosure.
[0449] In this case, the device equipped with the FPGA or CPU may be connected to the signal processing device wirelessly or via a wire, or may be connected to the signal processing server via a network, and this device and the signal processing device or the signal processing server may perform the acoustic processing method, encoding method, or decoding method described in the present disclosure.
[0450] For example, the sound processing device, encoding device, or decoding device in the present disclosure may include an FPGA, a CPU, etc. Furthermore, the sound processing device, encoding device, or decoding device may include an interface for externally obtaining software for operating the FPGA, CPU, etc., and a memory for storing the obtained software. Then, the FPGA, CPU, etc. may operate based on the stored software to perform the signal processing described in the present disclosure.
[0451] A server may provide software related to the acoustic processing, encoding processing, or decoding processing of the present disclosure. Then, a terminal or device may operate as the acoustic processing device, encoding device, or decoding device described in the present disclosure by installing the software. Note that the terminal or device may connect to the server via a network and install the software.
[0452] Furthermore, a device other than the terminal or device may connect to a server via a network to acquire data for installing the software, and the other device may provide the data for installing the software to the terminal or device, thereby installing the software in the terminal or device. Note that an example of the software may be VR software or AR software for causing a terminal or device to execute the acoustic processing method described using the embodiment.
[0453] In the above-described embodiments, each component may be configured with dedicated hardware, or may be realized by executing a software program suitable for each component. Each component may be realized by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.
[0454] Although the devices and the like according to one or more aspects have been described above based on the embodiments, the aspects grasped based on the present disclosure are not limited to the embodiments. As long as they do not deviate from the spirit of the present disclosure, forms obtained by applying various modifications conceivable by a person skilled in the art to the embodiments, and forms constructed by combining components in different modifications, may also be included within the scope of one or more aspects.
[0455] (Additional Notes) The above description of the embodiments discloses the following techniques.
[0456] (Technology 1) A sound processing device comprising a circuit and a memory, wherein the circuit uses the memory to acquire sound space information including information on a sound source in a sound space, information on objects in the sound space, and information on the position of a listener in the sound space, and the sound space information is used to calculate an evaluation value of a reflected sound generated in response to a sound generated from the sound source.
[0457] (Technology 2) The sound processing device according to Technology 1, wherein the circuit controls whether or not to select the reflected sound based on the evaluation value.
[0458] (Technology 3) The sound processing device according to Technology 2, wherein the circuit does not perform binaural processing on the reflected sound if the reflected sound is not selected.
[0459] (Technology 4) A sound processing device according to any one of technologies 1 to 3, wherein the circuit calculates the volume of the reflected sound and calculates the evaluation value when the volume exceeds a predetermined threshold.
[0460] (Technology 5) The sound processing device described in Technology 2, wherein when the reflected sound is selected based on the evaluation value, the circuit calculates the total computational load of one or more selected reflected sounds including the reflected sound, and cancels the selection of the reflected sound if the total computational load exceeds a predetermined upper limit.
[0461] (Technology 6) The sound processing device according to Technology 5, wherein the total calculation load is defined by the number of the one or more selectively reflected sounds or the processing amount of the one or more selectively reflected sounds.
[0462] (Technology 7) An audio processing device described in any of Technologies 1 to 6, wherein the circuit calculates the volume of each of a plurality of reflected sounds that occur as the reflected sound in the sound space, and calculates the evaluation value of each of one or more reflected sounds among the plurality of reflected sounds that have a volume equal to or greater than a predetermined threshold.
[0463] (Technology 8) The circuit calculates the total computational load of the one or more reflected sounds, and if the total computational load exceeds a predetermined upper limit, calculates the evaluation value of each of the one or more reflected sounds.
[0464] (Technology 9) A sound processing device described in any of technologies 1 to 8, wherein the circuit calculates the evaluation value of each of a plurality of reflected sounds generated as the reflected sound in the sound space, adds the computational load of each of the plurality of reflected sounds to a total computational load in descending order of the evaluation value, compares the total computational load with a predetermined upper limit each time the computational load of the reflected sound is added to the total computational load, and selects the reflected sound if the total computational load obtained by adding the computational loads of the reflected sounds does not exceed the predetermined upper limit, and does not select one or more remaining reflected sounds after the reflected sound from among the plurality of reflected sounds if the total computational load obtained by adding the computational loads of the reflected sounds exceeds the predetermined upper limit.
[0465] (Technology 10) A sound processing device described in any of technologies 1 to 9, wherein the evaluation value is the sum of at least one of an index value related to volume, a visual index value, an index value related to the object, and an index value indicating the relationship between the reflected sound and a direct sound corresponding to the reflected sound.
[0466] (Technology 11) The sound processing device according to Technology 10, wherein the circuit increases the index value relating to the volume as the volume of the sound generated from the sound source increases.
[0467] (Technology 12) An audio processing device described in Technology 10 or 11, wherein the circuit increases the visual index value when the sound source is within the listener's field of view compared to when the sound source is not within the listener's field of view.
[0468] (Technology 13) The sound processing device according to any one of techniques 10 to 12, wherein the circuit increases the visual index value as the moving speed of the sound source decreases.
[0469] (Technology 14) A sound processing device according to any one of techniques 10 to 13, wherein the index value relating to the object is assigned to each object in the sound space and is included in the sound space information.
[0470] (Technology 15) An audio processing device described in any of Technologies 10 to 14, wherein the circuit increases the index value indicating the relationship between the direct sound and the reflected sound the larger the angle between the direction from which the direct sound arrives and the direction from which the reflected sound arrives.
[0471] (Technology 16) An acoustic processing device described in any of Technologies 10 to 15, wherein the circuit increases an index value indicating the relationship between the direct sound and the reflected sound the greater the difference between the distance the direct sound takes from the sound source to the listener and the distance the reflected sound takes from the sound source to the listener after reflection.
[0472] (Technology 17) An audio processing device described in any of Technologies 10 to 16, wherein the circuit increases the index value indicating the relationship between the direct sound and the reflected sound the more the amplitude value of the reflected sound exceeds a temporal masking threshold, which is the threshold of a temporal masking phenomenon in which the reflected sound is masked by the direct sound when the amplitude value of the reflected sound is below a threshold.
[0473] (Technology 18) An audio processing device described in any of Technologies 10 to 17, wherein the circuit reduces an index value for the object related to a selected reflected sound among multiple reflected sounds generated as the reflected sound in the sound space, calculates the evaluation value for reflected sounds that have not yet been selected, and repeatedly performs a process of selecting reflected sounds in descending order of the evaluation value, and terminates the repeatedly performed process when the total computational load of one or more selected reflected sounds from the multiple reflected sounds exceeds a predetermined upper limit.
[0474] (Technology 19) An audio processing device comprising a circuit and a memory, wherein the circuit uses the memory to acquire information on the volume of a sound output from a sound source, and uses the volume information to correct an evaluation value of a reflected sound corresponding to the sound, and controls whether or not to select the reflected sound based on the corrected evaluation value.
[0475] (Technology 20) The sound processing device according to Technology 19, wherein the volume has a transition.
[0476] (Technology 21) An acoustic processing method including the steps of: acquiring sound space information including information on a sound source in a sound space, information on objects in the sound space, and information on the position of a listener in the sound space; and using the sound space information to calculate an evaluation value of reflected sound that occurs in response to sound emitted from the sound source.
[0477] (Technology 22) A program for causing a computer to execute the acoustic processing method according to Technology 21.
[0478] The present disclosure includes aspects that are applicable to, for example, an audio processing device, an encoding device, a decoding device, or a terminal or device that includes any of these devices.
[0479] 1000 Stereophonic sound reproduction system 1001 Audio signal processing device (audio processing device) 1002 Audio presentation device 1100, 1120, 1500 Encoding device 1101, 1113 Input data 1102 Encoder 1103 Encoded data 1104, 1114, 1404, 1503 Memory 1110, 1130 Decoding device 1111 Audio signal 1112, 1200, 1210 Decoder 1121 Transmitting unit 1122 Transmitted signal 1131 Receiving unit 1132 Received signal 1201, 1211 Spatial information management unit 1202 Audio data decoder 1203, 1213, 1300 Rendering unit 1301 Analysis unit 1302, 1314 Selection unit 1303 Synthesis unit 1311 Reverberation processing unit 1312 Early reflection processing unit 1313 Distance attenuation processing unit 1315 Generation unit 1316 Binaural processing unit 1401 Speaker 1402, 1501 Processor 1403, 1502 Communication IF 1405 Sensor
Claims
1. A circuit and a memory, The circuit uses the memory to Acquire sound space information including information on a sound source in a sound space, information on an object in the sound space, and information on a position of a listener in the sound space; Using the sound space information, an evaluation value of a reflected sound generated in response to the sound generated from the sound source is calculated. Sound processing equipment.
2. The circuit controls whether or not to select the reflected sound based on the evaluation value. The sound processing device according to claim 1 .
3. the circuit does not perform binaural processing on the reflected sound if the reflected sound is not selected; The sound processing device according to claim 2 .
4. The circuit comprises: Calculating the volume of the reflected sound; Calculating the evaluation value when the volume exceeds a predetermined threshold value. The sound processing device according to any one of claims 1 to 3.
5. When the reflected sound is selected based on the evaluation value, the circuit Calculating a total calculation load of one or more selective reflected sounds including the reflected sound; When the total calculation load exceeds a predetermined upper limit, the selection of the reflected sound is canceled. The sound processing device according to claim 2 .
6. The total calculation load is defined by the number of the one or more selective reflected sounds or the processing amount of the one or more selective reflected sounds. The sound processing device according to claim 5 .
7. The circuit comprises: Calculating a volume of each of a plurality of reflected sounds generated as the reflected sound in the sound space; Calculating the evaluation value of each of one or more reflected sounds having a volume equal to or greater than a predetermined threshold value among the plurality of reflected sounds; The sound processing device according to any one of claims 1 to 3.
8. The circuit comprises: Calculating a total computational load of the one or more reflected sounds; When the total calculation load exceeds a predetermined upper limit, the evaluation value of each of the one or more reflected sounds is calculated. The sound processing device according to claim 7 .
9. The evaluation value is a total value of at least one of an index value related to a volume, a visual index value, an index value related to the object, and an index value indicating a relationship between a direct sound corresponding to the reflected sound and the reflected sound. The sound processing device according to any one of claims 1 to 3.
10. The circuit increases the index value relating to the volume as the volume of the sound generated from the sound source increases. The sound processing device according to claim 9 .
11. the circuitry increases the visual indicator value when the sound source is within the field of view of the listener more than when the sound source is not within the field of view of the listener. The sound processing device according to claim 9 .
12. The circuitry increases the visual indicator value the slower the sound source is moving. The sound processing device according to claim 9 .
13. The index value related to the object is assigned to each object in the sound space and is included in the sound space information. The sound processing device according to claim 9 .
14. The circuit increases an index value indicating a relationship between the direct sound and the reflected sound as an angle between the direction from which the direct sound arrives and the direction from which the reflected sound arrives increases. The sound processing device according to claim 9 .
15. The circuit increases an index value indicating a relationship between the direct sound and the reflected sound as a difference between a distance from the sound source to the listener and a distance from the sound source to the listener through reflection increases. The sound processing device according to claim 9 .
16. the circuit increases the index value indicating the relationship between the direct sound and the reflected sound as the amplitude value of the reflected sound exceeds a temporal masking threshold, which is the threshold of a temporal masking phenomenon in which the reflected sound is masked by the direct sound when the amplitude value of the reflected sound is equal to or less than a threshold. The sound processing device according to claim 9 .
17. A circuit and a memory, The circuit uses the memory to Acquire information on the volume of the sound output from the sound source, Using the volume information, correcting an evaluation value of a reflected sound corresponding to the sound; controlling whether or not to select the reflected sound based on the corrected evaluation value; Sound processing equipment.
18. The volume has a transition. The sound processing device according to claim 17.
19. obtaining sound space information including information of a sound source in a sound space, information of an object in the sound space, and information of a position of a listener in the sound space; and calculating an evaluation value of a reflected sound generated in response to the sound generated from the sound source using the sound space information. Acoustic processing methods.
20. 20. A method for causing a computer to execute the acoustic processing method according to claim 19, program.