Audio signal processing method, computer program, and audio signal processing device

By merging and reproducing sound signals based on the direction of sound arrival in virtual reality and augmented reality, the problem of excessive computational load and workload is solved, achieving a more immersive audio experience.

CN121970374APending Publication Date: 2026-05-01PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
Filing Date
2024-10-04
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively reduce the computational load and complexity of sound signal processing in virtual reality and augmented reality, especially when processing multiple reflected sounds, leading to excessive consumption of computing resources.

Method used

The sound signal processing device determines whether to merge sound signals based on the direction of sound arrival, and generates and reproduces the merged sound signal after merging. It uses simple indicators and center of gravity position to determine the direction of arrival of the merged sound, reducing the amount of computation and load.

Benefits of technology

It achieves a more immersive audio experience by reducing computational load and providing a more realistic sense of presence without compromising sound localization and spatial perception.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121970374A_ABST
    Figure CN121970374A_ABST
Patent Text Reader

Abstract

A sound signal processing method performed by a sound signal processing device includes: an acquisition step of acquiring a first sound signal indicating a first sound and including first attribute information for specifying an attribute of the first sound signal, and a second sound signal indicating a second sound and including second attribute information for specifying an attribute of the second sound signal; the second sound signal includes second attribute information for determining an attribute of the second sound signal. A determination step for determining whether or not to merge the acquired first sound signal and the acquired second sound signal on the basis of an index corresponding to a first arrival direction in which the first sound reaches a listening position, which is the position of the listener, and a second arrival direction in which the second sound reaches the listening position; a merging step for generating a merged sound signal in which the first sound signal and the second sound signal are merged when it is determined to merge the first sound signal and the second sound signal; and a reproduction step of outputting an output signal based on the generated combined sound signal.
Need to check novelty before this filing date? Find Prior Art

Description

Sound signal processing methods, computer programs, and sound signal processing devices Technical Field

[0001] This disclosure relates to methods for processing sound signals, etc. Background Technology

[0002] In recent years, goods and services utilizing ER (Extended Reality), including VR (Virtual Reality), AR (Augmented Reality), and MR (Mixed Reality), have become increasingly popular. Consequently, the importance of sound signal processing technologies—which assign acoustic effects corresponding to the environment of a virtual sound source to the sound emitted in virtual or real space, thereby providing listeners with immersive audio—has increased.

[0003] Additionally, the listener can also be a listener or a user. Furthermore, technologies related to the sound signal processing method disclosed herein are shown in Patent Document 1, Patent Document 2, Patent Document 3, and Non-Patent Document 1.

[0004] Prior Art Documents: Patent Documents: Patent Document 1: Japanese Patent No. 6288100; Patent Document 2: Japanese Unexamined Patent Publication No. 2019-22049; Patent Document 3: International Publication No. 2021 / 180938; Non-Patent Documents: Non-Patent Document 1: BCJ Moore, *An Outline of Auditory Psychology*, Chengxin Bookstore, April 20, 1994, Chapter 6: Spatial Perception, p. 225; Non-Patent Document 2: Kazuhiro Iida, *Spatial Acoustics (Acoustic Science Series)*, Corona Corporation, July 2010, Figure 2.26. Summary of the Invention

[0005] The problem to be solved by the invention: However, in the technology shown in Patent Document 1, it is sometimes difficult to reduce the amount of computation and the computational load.

[0006] Therefore, the purpose of this disclosure is to provide a sound signal processing method that can reduce computational load and computational complexity.

[0007] The present disclosure discloses a sound signal processing method performed by a sound signal processing device, comprising: an acquisition step, acquiring a first sound signal and a second sound signal, wherein the first sound signal represents a first sound and includes first attribute information for determining attributes of the first sound signal, and the second sound signal represents a second sound and includes second attribute information for determining attributes of the second sound signal; a decision step, determining whether to merge the acquired first sound signal and the acquired second sound signal based on indices corresponding to a first arrival direction and a second arrival direction, wherein the first arrival direction is the direction in which the first sound arrives at the location of the listener, i.e., the listening position, and the second arrival direction is the direction in which the second sound arrives at the listening position; a merging step, generating a merged sound signal obtained by merging the first sound signal and the second sound signal if it is determined that the first sound signal and the second sound signal should be merged; and a reproduction step, outputting an output signal based on the generated merged sound signal.

[0008] Additionally, one aspect of the audio signal processing method disclosed herein is an audio signal processing method performed by an audio signal processing apparatus, comprising: an acquisition step, acquiring a first audio signal and a second audio signal, wherein the first audio signal represents a first sound and includes first attribute information for determining attributes of the first audio signal, and the second audio signal represents a second sound and includes second attribute information for determining attributes of the second audio signal; a merging step, wherein, if it is determined that the acquired first audio signal and the acquired second audio signal are to be merged, generating a merged audio signal obtained by merging the first audio signal and the second audio signal; and a reproduction step, outputting a merged audio signal based on the generated audio signal. The output signal of the merged sound signal includes the first attribute information contained in the first sound signal, which includes first position information indicating the position of the first sound and first volume information indicating the volume of the first sound. The second attribute information contained in the second sound signal includes second position information indicating the position of the second sound and second volume information indicating the volume of the second sound. In the merging step, based on the first position information and the first volume information, and the second position information and the second volume information, the direction of arrival of the merged sound represented by the merged sound signal at the listener's location, i.e., the listening location, is determined.

[0009] Additionally, one aspect of the sound signal processing method disclosed herein is a sound signal processing method executed by a sound signal processing device, comprising: an acquisition step, acquiring M sound signals, the sound signals representing predetermined tones and including attribute information for determining the attributes of the sound signals, wherein M is an integer greater than or equal to 2; a decision step, based on the direction of arrival of the M predetermined tones at the location of the listener, i.e., the listening location, deciding whether to merge N sound signals from the acquired M sound signals, wherein N is an integer greater than or equal to 1 and less than or equal to M; a merging step, if it is decided to merge the N sound signals, generating a merged sound signal obtained by merging the N sound signals; and a reproduction step, outputting an output signal based on the generated merged sound signal.

[0010] Additionally, one aspect of the audio signal processing method disclosed herein is an audio signal processing method executed by an audio signal processing device, comprising: an acquisition step, acquiring a first audio signal and a second audio signal, wherein the first audio signal represents a first sound and includes first attribute information for determining attributes of the first audio signal, and the second audio signal represents a second sound and includes second attribute information for determining attributes of the second audio signal; a decision step, determining whether to merge the acquired first audio signal and the acquired second audio signal based on a first direction of arrival of the first sound to the listener's location (i.e., the listening location) and a second direction of arrival of the second sound to the listening location; a merging step, generating a merged audio signal obtained by merging the first audio signal and the second audio signal if it is determined that the first audio signal and the second audio signal should be merged; and a reproduction step, outputting an output signal based on the generated merged audio signal, wherein in the decision step, multiple signals representing the direction of arrival of the sound to the listening location and a head-related transfer function based on the direction of arrival are used. In the direction of arrival information, an arrival direction information corresponding to the first arrival direction and an arrival direction information corresponding to the second arrival direction are determined. In the merging step, if the determined arrival direction information corresponding to the first arrival direction and the determined arrival direction information corresponding to the second arrival direction are the same, the merged sound signal is generated. In the reproduction step, the output signal generated by applying the head-related transfer function represented by the arrival direction information corresponding to the first arrival direction to the generated merged sound signal is output. If the determined arrival direction information corresponding to the first arrival direction and the determined arrival direction information corresponding to the second arrival direction are different, the output signal generated by applying the head-related transfer function represented by the arrival direction information corresponding to the first arrival direction to the first sound signal and the output signal generated by applying the head-related transfer function represented by the arrival direction information corresponding to the second arrival direction to the second sound signal are output.

[0011] Furthermore, a computer program of one aspect of this disclosure enables a computer to perform the aforementioned sound signal processing method.

[0012] Additionally, one aspect of the audio signal processing apparatus disclosed herein includes: an acquisition unit that acquires a first audio signal and a second audio signal, wherein the first audio signal represents a first sound and includes first attribute information for determining attributes of the first audio signal, and the second audio signal represents a second sound and includes second attribute information for determining attributes of the second audio signal; a determination unit that decides whether to merge the acquired first audio signal and the acquired second audio signal based on indicators corresponding to a first arrival direction and a second arrival direction, wherein the first arrival direction is the direction in which the first sound arrives at the location of the listener, i.e., the listening position, and the second arrival direction is the direction in which the second sound arrives at the listening position; a merging unit that, if it is determined that the first audio signal and the second audio signal should be merged, generates a merged audio signal obtained by merging the first audio signal and the second audio signal; and a reproduction unit that outputs an output signal based on the generated merged audio signal.

[0013] Furthermore, these inclusive or specific methods can also be implemented by non-transitory recording media such as systems, devices, methods, integrated circuits, computer programs, or computer-readable CD-ROMs, or by any combination of systems, devices, methods, integrated circuits, computer programs, and recording media.

[0014] The effect of the invention is that the sound signal processing method according to one aspect of the present disclosure can reduce the amount of computation and the computational load. Attached Figure Description

[0015] Figure 1 is a diagram showing an example of direct and reflected tones generated in sound space.

[0016] Figure 2 is a diagram showing an example of the stereo sound reproduction system of Embodiment 1.

[0017] Figure 3A is a block diagram showing an example of the configuration of the encoding device according to Embodiment 1.

[0018] Figure 3B is a block diagram showing an example of the configuration of the decoding device according to Embodiment 1.

[0019] Figure 3C is a block diagram showing another configuration example of the encoding device according to Embodiment 1.

[0020] Figure 3D is a block diagram showing another configuration example of the decoding device according to Embodiment 1.

[0021] Figure 4A is a block diagram showing an example of the configuration of the decoder in Embodiment 1.

[0022] Figure 4B is a block diagram showing another configuration example of the decoder in Embodiment 1.

[0023] Figure 5 is a diagram showing an example of the physical configuration of the sound signal processing device according to Embodiment 1.

[0024] Figure 6 is a diagram showing an example of the physical configuration of the encoding device according to Embodiment 1.

[0025] Figure 7 is a block diagram showing an example of the configuration of the rendering unit in Embodiment 1.

[0026] Figure 8 is a flowchart illustrating an example of the operation of the sound signal processing device according to Embodiment 1.

[0027] Figure 9 is a diagram showing the relative positional relationship between the listener and the obstacle object.

[0028] Figure 10 is a diagram showing the relatively close positional relationship between the listener and the obstacle object.

[0029] Figure 11 is a graph showing the relationship between the time difference and threshold of direct and reflected tones.

[0030] Figure 12A is a partial diagram illustrating an example of a method for setting threshold data.

[0031] Figure 12B is a partial diagram illustrating an example of how threshold data is set.

[0032] Figure 12C is a partial diagram illustrating an example of how threshold data is set.

[0033] Figure 13 is a diagram illustrating an example of a method for setting a threshold.

[0034] Figure 14 is a flowchart illustrating an example of selection processing.

[0035] Figure 15 is a graph showing the relationship between the direction of the direct tone, the direction of the reflected tone, the time difference, and the threshold.

[0036] Figure 16 is a graph showing the relationship between angle difference, time difference, and threshold.

[0037] Figure 17 is a block diagram showing another example of the configuration of the rendering unit.

[0038] Figure 18 is a flowchart illustrating another example of selection processing.

[0039] Figure 19 is a flowchart illustrating yet another example of selection processing.

[0040] Figure 20 is a flowchart illustrating a first variation of the operation of the sound signal processing apparatus according to Embodiment 1.

[0041] Figure 21 is a flowchart illustrating a second variation of the operation of the sound signal processing apparatus of Embodiment 1.

[0042] Figure 22 is a diagram showing an example of the configuration of avatars, sound source objects, and obstacle objects.

[0043] Figure 23 is a flowchart illustrating yet another example of the selection process.

[0044] Figure 24 is a block diagram showing a configuration example for pipeline processing in the rendering unit.

[0045] Figure 25 is a diagram showing the transmission and diffraction of sound.

[0046] Figure 26 is a diagram showing an example of the positional relationship between the listener and the obstacle object in Implementation 1.

[0047] Figure 27 is a diagram showing another example of the positional relationship between the listener and the obstacle object in Implementation 1.

[0048] Figure 28 is an example of the echo detection limit threshold of Implementation 1.

[0049] Figure 29 is a diagram showing an example of the positional shift of the sound image of the reflected sound in the positional relationship shown in Figure 27.

[0050] Figure 30 is a diagram showing an example of direct and indirect tones arriving at the listener from the same direction.

[0051] Figure 31 is a diagram showing examples of direct and indirect tones reaching the listener from one sound source and another sound source, respectively.

[0052] Figure 32 is a diagram illustrating an example of sound, i.e., transmitted sound and diffracted sound, reaching a listener from a sound source.

[0053] Figure 33 is a block diagram showing an example of the configuration of the rendering unit in Embodiment 2.

[0054] Figure 34 is a flowchart illustrating an example 1 of the operation of the sound signal processing device in Embodiment 2.

[0055] Figure 35 is a flowchart showing an example of the processing performed by the selection unit and the reproduction unit in Operation Example 1 of Embodiment 2.

[0056] Figure 36 is a flowchart showing an example of the processing performed by the selection unit and the reproduction unit in Operation Example 2 of Embodiment 2.

[0057] Figure 37 is a diagram showing the SOFA of the HRTF in Implementation Method 2.

[0058] Figure 38 is a diagram of a cone used to illustrate the relationship between the position of the HRTF and the listening position as defined in Embodiment 2.

[0059] Figure 39 is a flowchart showing an example of the processing performed by the selection unit and the reproduction unit in Operation Example 3 of Embodiment 2.

[0060] Figure 40 is a diagram illustrating the direction of arrival of the combined sound used in Operation Example 3 of Embodiment 2.

[0061] Figure 41 is a flowchart showing an example of the processing performed by the selection unit and the reproduction unit in Operation Example 4 of Embodiment 2.

[0062] Figure 42 is a flowchart showing an example of the processing performed by the selection unit and the reproduction unit in Operation Example 5 of Embodiment 2.

[0063] Figure 43 is a flowchart showing an example of the processing performed by the selection unit and the reproduction unit in Operation Example 6 of Embodiment 2.

[0064] Figure 44 is a flowchart showing an example of the processing performed by the selection unit and the reproduction unit in Operation Example 7 of Embodiment 2.

[0065] Figure 45 is a block diagram showing an example of the configuration of the rendering unit in Embodiment 3.

[0066] Figure 46 is a diagram showing an example of the merged region in Implementation Method 3.

[0067] Figure 47 is a flowchart illustrating an example 1 of the operation of the sound signal processing device in Embodiment 3.

[0068] Figure 48 is a flowchart showing an example of the processing performed by the selection unit and the reproduction unit in Operation Example 1 of Embodiment 3.

[0069] Figure 49 is a diagram illustrating the merged region in Implementation Method 3.

[0070] Figure 50 is another diagram illustrating the merged region of Embodiment 3.

[0071] Figure 51 is a diagram illustrating the method for calculating the direction of arrival of the combined sound in Embodiment 3.

[0072] Figure 52 is a diagram showing another first example of the merged region in Implementation 3.

[0073] Figure 53 is a diagram showing another second example of the merged region in Implementation 3.

[0074] Figure 54 is a diagram showing another third example of the merged region in Implementation 3.

[0075] Figure 55 is a diagram showing another fourth example of the merged region in Implementation 3.

[0076] Figure 56 is a diagram showing another fifth example of the merged region in Implementation 3.

[0077] Figure 57 is a diagram showing another sixth example of the merged region in Implementation 3.

[0078] Figure 58 is a diagram illustrating an example of the phase difference between the direct tone and the reflected tone in Embodiment 4.

[0079] Figure 59 is a diagram showing another example of the phase difference between the direct tone and the reflected tone in Embodiment 4.

[0080] Figure 60 is a flowchart showing an example of the processing performed by the selection unit in Operation Example 1 of Embodiment 4.

[0081] Figure 61 is a flowchart of the first example showing the details of step S703 of embodiment 4.

[0082] Figure 62 is a diagram showing the direct sound, the reflected sound, and the sound of adding the direct sound and the reflected sound together in Embodiment 4.

[0083] Figure 63 is a flowchart of the third example showing the details of step S703 in embodiment 4.

[0084] Figure 64 is a diagram showing the first example of SOFA for HRTF in other implementations.

[0085] Figure 65 is a diagram showing a second example of SOFA for HRTF in other implementations.

[0086] Figure 66 is a diagram showing a third example of SOFA for HRTF in other implementations. Detailed Implementation

[0087] (The insights that form the basis of this disclosure) In the past, sound signal processing techniques have been studied to impart acoustic effects to sounds emitted by virtual sound sources in virtual or real spaces based on the environment of that space and to provide immersive audio to listeners.

[0088] Such a sound signal processing technique is disclosed in Patent Document 1. More specifically, Patent Document 1 discloses a technique for detecting the importance of audio signals (sound signals) and not outputting audio signals with low detected importance. In this way, by not outputting audio signals with low importance, this sound signal processing technique is expected to reduce the amount of computation and the computational load.

[0089] Additionally, reflected sound sometimes becomes important in sound space (virtual or real space).

[0090] Figure 1 is an example of direct and reflected sounds generated in a sound space. In sound processing that uses sound to represent the characteristics of a virtual space, in order to represent the breadth of the space and the material of the walls, as well as to accurately grasp the location of the sound source (sound image localization), it is effective to reproduce not only direct sounds but also reflected sounds.

[0091] For example, when listening to sound in a cuboid-shaped room as shown in Figure 1, a sound source produces six primary reflections corresponding to the six walls. The reproduction of these reflections provides clues for a proper understanding of the space and the sound image. Furthermore, for each reflection, secondary reflections are produced on surfaces other than the one that produced the reflection. These reflections also serve as perceptually valid clues.

[0092] However, even considering only the case of secondary reflection, a sound source will produce 1 direct tone and 36 (6+6×5) reflected tones, resulting in 37 vocal lines. Processing these vocal lines requires a considerable amount of computation.

[0093] Furthermore, in recent years, applications related to the concept of the Metaverse, such as virtual meetings, virtual shopping, or virtual concerts, have inevitably involved multiple sound sources, thus requiring a much larger amount of computation.

[0094] Furthermore, listeners in virtual spaces use headphones or VR goggles to receive audio. To provide stereo sound to such listeners, binaural processing is performed on each sound ray, assigning sound pressure levels and phase differences between the two ears to reproduce the direction of arrival and the sense of distance. Therefore, the computational load is extremely high if all the reflected sounds are to be reproduced.

[0095] On the other hand, small rechargeable batteries are sometimes used as batteries for VR goggles worn by listeners experiencing virtual space, due to their convenience. To extend battery life, the computational load required for the processing described above should ideally be low. Therefore, it is desirable to reduce the number of sound lines generated on a scale of several hundred without compromising sound localization and spatial accuracy.

[0096] Furthermore, in systems that reproduce sound, there are sometimes 6 degrees of freedom (DoF) allowed for the listener's position (i.e., the listening position as the listener's location) and orientation. In such cases, the positional relationship between the listener, the sound source, and the object reflecting the sound cannot be determined if it is not during reproduction (rendering). Therefore, the reflected sound also cannot be determined if it is not during reproduction. Consequently, it is difficult to predetermine the reflected sound of the object being processed.

[0097] Therefore, the process of selecting and outputting (reproducing) one or more reflected sounds from multiple reflected sounds generated in the sound space during reproduction, whether they are the processed or unprocessed objects, is beneficial for the appropriate reduction of computational load and computational complexity.

[0098] Furthermore, controlling whether to select a sound corresponds to determining whether to select a sound, and more specifically, to determining whether to select and output (reproduce) a sound. Additionally, selecting a sound can mean selecting it as the target sound for processing, or selecting it as a non-target sound.

[0099] Furthermore, while Patent Document 1 detects the importance of an audio signal, and more specifically the importance of the direct tone represented by that audio signal, it does not investigate the importance of reflected tones. Therefore, when indirect tones such as reflected tones are generated as shown in Figure 1, the computational load and computational complexity increase, making it sometimes difficult to reduce the computational load and computational complexity.

[0100] Therefore, there is a need for sound signal processing methods that can reduce computational load in the sound space.

[0101] Therefore, the first aspect of the sound signal processing method disclosed herein is a sound signal processing method executed by a sound signal processing device, comprising: an acquisition step, acquiring a first sound signal and a second sound signal, wherein the first sound signal represents a first sound and includes first attribute information for determining attributes of the first sound signal, and the second sound signal represents a second sound and includes second attribute information for determining attributes of the second sound signal; a determination step, determining whether to merge the acquired first sound signal and the acquired second sound signal based on indices corresponding to a first arrival direction and a second arrival direction, wherein the first arrival direction is the direction in which the first sound arrives at the location of the listener, i.e., the listening position, and the second arrival direction is the direction in which the second sound arrives at the listening position; a merging step, generating a merged sound signal obtained by merging the first sound signal and the second sound signal if it is determined that the first sound signal and the second sound signal should be merged; and a reproduction step, outputting an output signal based on the generated merged sound signal.

[0102] Therefore, based on the indicators corresponding to the first arrival direction of the first sound and the second arrival direction of the second sound, it is determined whether to merge the first and second sound signals. If it is determined that the first and second sound signals should be merged, an output signal based on the merged sound signal is output in the reproduction step. If it is determined that the first and second sound signals should not be merged, an output signal based on the first sound signal and an output signal based on the second sound signal are output in the reproduction step. Compared to the case where only one output signal based on the first sound signal and one output signal based on the second sound signal are output, the number of output signals is reduced when outputting an output signal based on the merged sound signal, thus reducing the computational load. In other words, a sound signal processing method that reduces computational load can be implemented.

[0103] The second aspect of the audio signal processing method disclosed herein, in the first aspect of the audio signal processing method, wherein the index is an index composed of two orthogonal axes.

[0104] Therefore, a simpler index can be used, that is, compared with the use of complex indexes, it does not require a large amount of computation and a large computational load, thus enabling a sound signal processing method that can reduce the amount of computation and the computational load.

[0105] The third aspect of the audio signal processing method disclosed herein, in the first or second aspect of the audio signal processing method, includes first attribute information in the first audio signal that includes first position information indicating the position of the first sound and first volume information indicating the volume of the first sound, and second attribute information in the second audio signal that includes second position information indicating the position of the second sound and second volume information indicating the volume of the second sound. In the merging step, based on the first position information and the first volume information, and the second position information and the second volume information, the direction of arrival of the merged sound represented by the merged audio signal at the listening position is determined.

[0106] Therefore, in the merging step, the direction of arrival of the merged tone can be determined based on the first position information, the first volume information, the second position information, and the second volume information.

[0107] The fourth aspect of the sound signal processing method disclosed herein, in the third aspect of the sound signal processing method, in the merging step, when the volume of the first sound represented by the first volume information is regarded as the weight corresponding to the position of the first sound represented by the first position information, and the volume of the second sound represented by the second volume information is regarded as the weight corresponding to the position of the second sound represented by the second position information, the position of the merged sound is determined to be the position of the center of gravity of the first sound and the second sound, and the direction from the determined position of the center of gravity toward the listening position is determined as the direction of arrival of the merged sound.

[0108] Therefore, the direction of arrival of the merged tone can be determined based on the position of the centroids of the first and second sounds, making it easy to determine the direction of arrival of the merged tone. In other words, determining the direction of arrival of the merged tone does not require a large amount of computation or a large computational load, thus enabling a sound signal processing method that reduces computational load.

[0109] The fifth aspect of the sound signal processing method disclosed herein is a sound signal processing method executed by a sound signal processing apparatus, comprising: an acquisition step, acquiring a first sound signal and a second sound signal, wherein the first sound signal represents a first sound and includes first attribute information for determining attributes of the first sound signal, and the second sound signal represents a second sound and includes second attribute information for determining attributes of the second sound signal; a merging step, wherein, if it is determined that the acquired first sound signal and the acquired second sound signal are to be merged, generating a merged sound signal obtained by merging the first sound signal and the second sound signal; and a reproduction step, outputting a sound signal based on the generated sound signal. The output signal of the merged sound signal, wherein the first attribute information contained in the first sound signal includes first position information indicating the position of the first sound and first volume information indicating the volume of the first sound, and the second attribute information contained in the second sound signal includes second position information indicating the position of the second sound and second volume information indicating the volume of the second sound, in the merging step, based on the first position information and the first volume information, and the second position information and the second volume information, determines the direction of arrival of the merged sound represented by the merged sound signal at the listener's location, i.e., the listening location.

[0110] Therefore, if it is decided to merge the first and second audio signals, the reproduction step outputs an output signal based on the merged audio signal. If it is assumed that the first and second audio signals are not merged, the reproduction step outputs an output signal based on the first audio signal and an output signal based on the second audio signal. Compared to outputting both an output signal based on the first and second audio signals, outputting an output signal based on the merged audio signal reduces the number of output signals, thus reducing computational complexity and workload. In other words, an audio signal processing method that reduces computational complexity and workload can be implemented.

[0111] The sixth aspect of the audio signal processing method of this disclosure, in any of the second to fifth aspects of the audio signal processing method, when it is determined that the first audio signal and the second audio signal are to be merged, outputs the output signal generated by applying a head-related transfer function based on the determined direction of arrival of the merged sound to the generated merged audio signal in the reproduction step.

[0112] This enables a sound signal processing method that allows listeners to hear more immersive combined sounds.

[0113] The seventh aspect of the audio signal processing method of this disclosure, in any of the first to fourth aspects of the audio signal processing method, if it is determined that the first audio signal and the second audio signal will not be merged, in the reproduction step, outputs the output signal generated by applying a head-related transfer function based on the first arrival direction to the first audio signal, and the output signal generated by applying a head-related transfer function based on the second arrival direction to the second audio signal.

[0114] This enables a sound signal processing method that allows listeners to hear a more immersive first and second sound.

[0115] The eighth aspect of the sound signal processing method disclosed herein is a sound signal processing method executed by a sound signal processing device, comprising: an acquisition step, acquiring M sound signals, the sound signals representing predetermined tones and including attribute information for determining the attributes of the sound signals, wherein M is an integer greater than or equal to 2; a determination step, based on the direction of arrival of the M predetermined tones at the location of the listener, i.e., the listening location, deciding whether to merge N sound signals from the acquired M sound signals, wherein N is an integer greater than or equal to 1 and less than or equal to M; a merging step, if it is decided to merge the N sound signals, generating a merged sound signal obtained by merging the N sound signals; and a reproduction step, outputting an output signal based on the generated merged sound signal.

[0116] Therefore, the decision to merge N sound signals is based on the direction of arrival of the specified sound. If merging N sound signals is determined, the reproduction step outputs an output signal based on the merged sound signals. If the decision is not to merge N sound signals, the reproduction step outputs an output signal based on each of the N sound signals. Compared to outputting an output signal based on each of the N sound signals, outputting an output signal based on the merged sound signals reduces the number of output signals, thus reducing computational complexity and workload. In other words, a sound signal processing method that reduces computational complexity and workload can be implemented.

[0117] The ninth aspect of the sound signal processing method disclosed herein, in the eighth aspect of the sound signal processing method, the attribute information contained in each of the M sound signals obtained includes specified tone position information indicating the position of the specified tone and specified tone volume information indicating the volume of the specified tone. In the merging step, based on the N specified tone position information and the N specified tone volume information, the direction of arrival of the merged tone represented by the merged sound signal to the listening position is determined.

[0118] Therefore, in the merging step, the direction of arrival of the merged tone can be determined based on the specified tone position information and the specified tone volume information.

[0119] The tenth aspect of the sound signal processing method disclosed herein, in the ninth aspect of the sound signal processing method, in the merging step, when the volume of the specified tone represented by the specified tone volume information is regarded as the weight corresponding to the position of the specified tone represented by the specified tone position information about the specified tone, determines that the position of the merged tone is the position of the center of gravity of N specified tones, and determines the direction from the determined position of the center of gravity toward the listening position as the direction of arrival of the merged tone.

[0120] Therefore, the direction of arrival of the merged tone can be determined based on the position of the center of gravity of N specified tones, which means that the direction of arrival of the merged tone can be determined easily. That is, in order to determine the direction of arrival of the merged tone, a large amount of computation and a large computational load are not required, thus enabling a sound signal processing method that can reduce the amount of computation and the computational load.

[0121] The eleventh aspect of this disclosure is a sound signal processing method executed by a sound signal processing apparatus, comprising: an acquisition step, acquiring a first sound signal and a second sound signal, wherein the first sound signal represents a first sound and includes first attribute information for determining attributes of the first sound signal, and the second sound signal represents a second sound and includes second attribute information for determining attributes of the second sound signal; a determination step, based on a first direction of arrival of the first sound to the location of the listener (i.e., the listening location) and a second direction of arrival of the second sound to the listening location, determining whether to merge the acquired first sound signal and the acquired second sound signal; a merging step, wherein, if it is determined that the first sound signal and the second sound signal should be merged, generating a merged sound signal obtained by merging the first sound signal and the second sound signal; and a reproduction step, outputting an output signal based on the generated merged sound signal, wherein, in the determination step, multiple arrival signals representing the direction of arrival of the sound to the listening location and a head-related transfer function based on the direction of arrival are used. In the direction information, an arrival direction information corresponding to the first arrival direction and an arrival direction information corresponding to the second arrival direction are determined. In the merging step, if the determined arrival direction information corresponding to the first arrival direction and the determined arrival direction information corresponding to the second arrival direction are the same, the merged sound signal is generated. In the reproduction step, the output signal generated by applying the head-related transfer function represented by the arrival direction information corresponding to the first arrival direction to the generated merged sound signal is output. If the determined arrival direction information corresponding to the first arrival direction and the determined arrival direction information corresponding to the second arrival direction are different, the output signal generated by applying the head-related transfer function represented by the arrival direction information corresponding to the first arrival direction to the first sound signal and the output signal generated by applying the head-related transfer function represented by the arrival direction information corresponding to the second arrival direction to the second sound signal are output.

[0122] Therefore, based on the arrival direction information corresponding to the first arrival direction of the first sound and the arrival direction information corresponding to the second arrival direction of the second sound, it is determined whether to merge the first sound signal and the second sound signal. If it is determined that the first sound signal and the second sound signal should be merged, the reproduction step outputs an output signal based on the merged sound signal. If it is assumed that the first sound signal and the second sound signal should not be merged, the reproduction step outputs an output signal based on the first sound signal and an output signal based on the second sound signal. Compared to the case where both an output signal based on the first sound signal and an output signal based on the second sound signal are output, the number of output signals is reduced when an output signal based on the merged sound signal is output, thus reducing the computational load. In other words, a sound signal processing method that reduces computational load can be implemented.

[0123] The computer program of the 12th aspect of this disclosure is a computer program for causing a computer to perform the sound signal processing method of any one of the 1st to 11th aspects.

[0124] Therefore, the computer can execute the above-mentioned sound signal processing method according to the computer program.

[0125] The audio signal processing apparatus of the 12th aspect of this disclosure includes: an acquisition unit that acquires a first audio signal and a second audio signal, the first audio signal representing a first sound and including first attribute information for determining attributes of the first audio signal, and the second audio signal representing a second sound and including second attribute information for determining attributes of the second audio signal; a determination unit that decides whether to merge the acquired first audio signal and the acquired second audio signal based on indicators corresponding to a first arrival direction and a second arrival direction, the first arrival direction being the direction in which the first sound arrives at the location of the listener, i.e., the listening position, and the second arrival direction being the direction in which the second sound arrives at the listening position; a merging unit that, if it is determined that the first audio signal and the second audio signal should be merged, generates a merged audio signal obtained by merging the first audio signal and the second audio signal; and a reproduction unit that outputs an output signal based on the generated merged audio signal.

[0126] Therefore, based on indicators corresponding to the first arrival direction of the first sound and the second arrival direction of the second sound, it is determined whether to merge the first sound signal and the second sound signal. If it is determined that the first sound signal and the second sound signal should be merged, the reproduction unit outputs an output signal based on the merged sound signal. If it is assumed that the first sound signal and the second sound signal should not be merged, the reproduction unit outputs an output signal based on the first sound signal and an output signal based on the second sound signal. Compared to the case where both an output signal based on the first sound signal and an output signal based on the second sound signal are output, the number of output signals is reduced when an output signal based on the merged sound signal is output, thus reducing the computational load. In other words, a sound signal processing apparatus capable of reducing computational load can be realized.

[0127] (Embodiment 1) (Example of a stereo sound reproduction system) FIG2 is a diagram showing an example of a stereo sound reproduction system 1000. Specifically, FIG2 shows a stereo sound reproduction system 1000 as an example of a system capable of applying the sound processing or decoding processing of the present disclosure. Stereo sound is also represented as immersive audio. The stereo sound reproduction system 1000 includes a sound signal processing device 1001 and a sound prompting device 1002.

[0128] The sound signal processing device 1001 is also manifested as an audio processing device, which performs audio processing on the sound signal emitted by the virtual sound source to generate an audio-processed sound signal for the listener. The sound signal is not limited to speech; any audible sound is acceptable. Audio processing, for example, is signal processing performed on the sound signal to reproduce the effects that the sound undergoes from the sound source to the listener.

[0129] The sound signal processing device 1001 performs sound processing based on spatial information describing the reasons for the aforementioned effects. Spatial information includes, for example, information indicating the location of the sound source, the listener, and surrounding objects; information indicating the shape of the space; and parameters related to sound propagation. The sound signal processing device 1001 may be, for example, a PC (Personal Computer), a smartphone, a tablet computer, or a game console.

[0130] The processed audio signal is presented to the listener by the audio prompt device 1002. The audio prompt device 1002 is connected to the audio signal processing device 1001 via wireless or wired communication. The processed audio signal generated by the audio signal processing device 1001 is transmitted to the audio prompt device 1002 via wireless or wired communication.

[0131] When the sound prompting device 1002 is composed of multiple devices, such as a device for the right ear and a device for the left ear, the multiple devices simultaneously provide prompts through communication between the multiple devices or through communication between each of the multiple devices and the sound signal processing device 1001. The sound prompting device 1002 may be, for example, a headset, earplugs, head-mounted display, or a surround sound system composed of multiple fixed speakers, worn on the listener's head.

[0132] Furthermore, the stereo sound reproduction system 1000 can also be used in combination with image prompting devices or stereoscopic image prompting devices that visually provide ER experiences, including AR / VR. For example, the space processed by spatial information is a virtual space, where the positions of sound sources, listeners, and objects are virtual positions of virtual sound sources, virtual listeners, and virtual objects within the virtual space. This space can also be represented as a sound space. Additionally, spatial information can also be represented as sound spatial information.

[0133] Furthermore, Figure 2 shows a system configuration example where the sound signal processing device 1001 and the sound prompting device 1002 are different devices. However, the stereo sound reproduction system 1000, which can apply the sound processing method (sound signal processing method) or decoding method of this disclosure, is not limited to the configuration shown in Figure 2. For example, the sound signal processing device 1001 may be included in the sound prompting device 1002, and the sound prompting device 1002 may perform both sound processing and sound prompting.

[0134] Alternatively, the sound signal processing device 1001 and the sound prompting device 1002 may share the implementation of the sound processing described in this disclosure. Furthermore, a portion or all of the sound processing described in this disclosure may be implemented via a server connected to the sound signal processing device 1001 or the sound prompting device 1002 through a network.

[0135] Furthermore, the audio signal processing apparatus 1001 can also perform audio processing by decoding a bitstream generated by encoding at least a portion of the audio signal and spatial information data used for audio processing. Therefore, the audio signal processing apparatus 1001 can also be exemplified as a decoding device.

[0136] (Example of an encoding device) FIG3A is a block diagram showing an example of the configuration of the encoding device 1100. Specifically, FIG3A shows the configuration of the encoding device 1100 as an example of the encoding device of this disclosure.

[0137] Input data 1101 is encoded object data containing spatial information and / or audio signals input to encoder 1102. Details regarding the spatial information will be explained later.

[0138] Encoder 1102 encodes the input data 1101 to generate encoded data 1103. Encoded data 1103 is, for example, a bitstream generated through encoding processing.

[0139] The memory 1104 stores the encoded data 1103. The memory 1104 may be, for example, a hard disk or an SSD (Solid-State Drive), or other types of memory.

[0140] Furthermore, in the above description, the bitstream generated through encoding processing was listed as an example of encoded data 1103 stored in memory 1104, but the encoded data 1103 can also be data other than a bitstream. For example, the encoding device 1100 may also store transformed data generated by converting the bitstream into a specified data format in memory 1104. The transformed data may, for example, be a file or multiplexed stream corresponding to more than one bitstream.

[0141] Here, the file is a file with a file format such as ISOBMFF (ISO Base Media File Format). Furthermore, the encoded data 1103 can also be in the form of multiple packets generated by splitting the aforementioned bitstream or file.

[0142] For example, the bitstream generated by encoder 1102 can be transformed into data different from the bitstream. In this case, encoding device 1100 has a transformation unit (not shown), which can perform the transformation processing either by the transformation unit or by a CPU (Central Processing Unit), which is an example of a processor described later.

[0143] (Example of a decoding device) FIG3B is a block diagram showing an example of the configuration of the decoding device 1110. Specifically, FIG3B shows the configuration of the decoding device 1110 as an example of the decoding device of this disclosure.

[0144] Memory 1114 stores, for example, the same data as the encoded data 1103 generated by encoding device 1100. The stored data is read from memory 1114 and input as input data 1113 into decoder 1112. Input data 1113 is, for example, a bitstream intended for decoding. Memory 1114 can be, for example, a hard disk or SSD, or other type of storage.

[0145] Alternatively, the decoding device 1110 may not input the data read from the memory 1114 as input data 1113 into the decoder 1112 as is, but instead transform the read data and input the transformed data as input data 1113 into the decoder 1112. The data before transformation may be, for example, multiplexed data containing more than one bitstream. Here, the multiplexed data may also be a file with a file format such as ISOBMFF.

[0146] Furthermore, the data before conversion can also be multiple packets generated by splitting the aforementioned bitstream or file. Alternatively, data different from the bitstream can be read from memory 1114 and converted into a bitstream. In this case, the decoding device 1110 may also include a conversion unit (not shown), and the conversion processing may be performed by the conversion unit, or by a CPU, as an example of a processor described later.

[0147] Decoder 1112 decodes the input data 1113 and generates an audio signal 1111 representing a prompt to the listener.

[0148] (Another example of an encoding device) FIG3C is a block diagram showing another configuration example of an encoding device. Specifically, FIG3C shows the configuration of encoding device 1120 as another example of the encoding device of this disclosure. In FIG3C, the same reference numerals as those in FIG3A are assigned to the same components as those in FIG3A, and descriptions of these components are omitted.

[0149] Encoding device 1100 stores encoded data 1103 in memory 1104. On the other hand, encoding device 1120 differs from encoding device 1100 in that it has a transmitting unit 1121 that transmits encoded data 1103 to the outside.

[0150] The transmitting unit 1121 transmits a transmission signal 1122 generated based on encoded data 1103 or data transformed from encoded data 1103 into other data formats to other devices or servers. The data used in generating the transmission signal 1122 may be, for example, a bit stream, multiplexed data, file, or packet as described in the encoding device 1100.

[0151] (Another Example of a Decoding Apparatus) FIG3D is a block diagram showing another configuration example of a decoding apparatus. Specifically, FIG3D shows the configuration of decoding apparatus 1130 as another example of a decoding apparatus of this disclosure. In FIG3D, the same reference numerals as those in FIG3B are assigned to the same components, and descriptions of these components are omitted.

[0152] Decoding device 1110 reads input data 1113 from memory 1114. On the other hand, decoding device 1130 differs from decoding device 1110 in that it has a receiving unit 1131 that receives input data 1113 from the outside.

[0153] The receiving unit 1131 receives the received signal 1132 to obtain received data, and outputs the input data 1113 input to the decoder 1112. The received data can be the same as the input data 1113 input to the decoder 1112, or it can be data in a different format than the input data 1113.

[0154] If the format of the received data differs from the format of the input data 1113, the receiving unit 1131 may convert the received data into the input data 1113. Alternatively, the receiving data may be converted into the input data 1113 by a conversion unit (not shown) or CPU of the decoding device 1130. The received data may be, for example, a bit stream, multiplexed data, a file, or a packet as described in the encoding device 1120.

[0155] (Example of a decoder) Figure 4A is a block diagram showing an example of the configuration of decoder 1200. Specifically, Figure 4A shows the configuration of decoder 1200 as an example of decoder 1112 in Figure 3B or Figure 3D.

[0156] Input data 1113 is the encoded bitstream, which contains encoded audio data as the encoded audio signal and metadata used in audio processing.

[0157] The Spatial Information Management Unit 1201 acquires the metadata contained in the input data 1113 and parses the metadata. The metadata contains information describing the elements that act on sound and are configured in the sound space. The Spatial Information Management Unit 1201 manages the spatial information used in sound processing obtained by parsing the metadata and provides the spatial information to the Rendering Unit 1203.

[0158] Furthermore, in this disclosure, the information used in sound processing is represented as spatial information, but other representations may also be used. For example, the information used in sound processing may be represented as sound spatial information or scene information. In addition, when the information used in sound processing changes over time, the spatial information input to the rendering unit 1203 may also be represented as information such as spatial state, sound spatial state, or scene state.

[0159] Furthermore, spatial information can be managed on a per-sound-space or per-scene basis. For example, when multiple distinct rooms are represented as virtual spaces, these rooms can be managed as separate scenes. Additionally, even within the same space, spatial information can be managed as different scenes depending on the presented conditions.

[0160] Therefore, multiple spatial information can be managed for multiple sound spaces or multiple scenes. In the management of multiple spatial information, each spatial information can be assigned an identifier to distinguish between the multiple spatial information.

[0161] Spatial information data can also be included in the bitstream, which is one example of input data 1113. Alternatively, the bitstream may contain identifiers for spatial information, and the spatial information data may be obtained from an information source outside the bitstream. Specifically, when the bitstream only contains identifiers for spatial information, the identifiers can be used during rendering to obtain spatial information data stored in the device's memory or an external server as input data 1113.

[0162] Furthermore, the information managed by the Spatial Information Management Department 1201 is not limited to the information contained in the bitstream. For example, the input data 1113 may include data on the characteristics and structure of the representation space obtained from software or servers providing VR or AR, as data not included in the bitstream.

[0163] Furthermore, the input data 1113 may also include data representing the characteristics and location of the listener or object. Additionally, the input data 1113 may include information about the listener's location obtained by sensors possessed by the terminal, including the decoding devices (1110, 1130), and may also include information representing the terminal's location inferred based on the information obtained from the sensors.

[0164] That is, the spatial information management unit 1201 can also communicate with external systems or servers to obtain spatial information and the listener's location (i.e., listening location). The spatial information management unit 1201 can obtain clock synchronization information from external systems and perform clock synchronization processing with the rendering unit 1203.

[0165] Furthermore, the space described above can be a virtually formed space, i.e., VR space, or a real space or a virtual space corresponding to a real space, i.e., AR space or MR space. Additionally, the virtual space can also be represented as a sound field or acoustic space. Furthermore, the information indicating position described above can be information such as coordinate values ​​representing the position within the space, information representing the relative position with respect to a defined reference position, or information representing the motion or acceleration of the position within the space.

[0166] The audio data decoder 1202 decodes the encoded audio data contained in the input data 1113 to obtain the audio signal.

[0167] The encoded audio data acquired by the stereo sound reproduction system 1000 is, for example, a bitstream encoded in a format specified by MPEG-H 3D Audio (ISO / IEC 23008-3). Furthermore, MPEG-H 3D Audio is merely one example of an encoding method that can be used when generating encoded audio data contained within a bitstream. The encoded audio data can also be a bitstream encoded in other encoding methods.

[0168] For example, the encoding method can also be an irreversible codec such as MP3 (MPEG-1 Audio Layer-3), AAC (Advanced Audio Coding), WMA (Windows Media Audio), AC3 (Audio Codec-3), or Vorbis. Alternatively, the encoding method can be a reversible codec such as ALAC (Apple Lossless Audio Codec) or FLAC (Free Lossless Audio Codec).

[0169] Alternatively, any encoding method other than those described above can be used. For example, PCM (pulse code modulation) data can be used as encoded audio data. In this case, for example, if the number of quantization bits of the PCM data is N, the decoding process can be a process of converting the N-bit binary number into a number form (e.g., floating-point form) that the rendering unit 1203 can process.

[0170] The rendering unit 1203 acquires the sound signal and spatial information, uses the spatial information to perform sound processing on the sound signal, and outputs the sound processed sound signal (sound signal 1111).

[0171] Before rendering begins, the Spatial Information Management Unit 1201 reads the metadata of the input signal, detects the rendering items such as objects and sounds defined by the spatial information, and sends them to the Rendering Unit 1203. After rendering begins, the Spatial Information Management Unit 1201 monitors the changes in spatial information and the listener's location over time, updates and manages the spatial information, and sends the updated spatial information to the Rendering Unit 1203.

[0172] The rendering unit 1203 generates and outputs a sound signal with added audio processing based on the sound signal contained in the input data 1113 and the spatial information received from the spatial information management unit 1201.

[0173] Spatial information update processing and sound signal output processing with added audio processing can also be performed by the same thread. Spatial information management unit 1201 and rendering unit 1203 can assign processing to their respective independent threads. When spatial information update processing and sound signal output processing with added audio processing are performed in different threads, the start frequency of each thread can be set separately, or the processing can be performed in parallel.

[0174] When the spatial information management unit 1201 and the rendering unit 1203 perform processing in different independent threads, computing resources can be allocated preferentially to the rendering unit 1203. As a result, sound output processing that does not allow for even small delays, such as generating noise like a popping sound with a delay of 1 sample (0.02 msec), can be performed safely.

[0175] At this time, the allocation of computing resources to the spatial information management unit 1201 is restricted. However, compared to the output processing of audio signals, the updating of spatial information is a low-frequency process (e.g., updating the listener's facial orientation), and therefore does not need to be instantaneous like the output processing of audio signals. Therefore, even if the allocation of computing resources is restricted, it will not have a significant impact on the sound quality.

[0176] Spatial information updates can be performed periodically at preset times or intervals, or they can be performed when preset conditions are met. Furthermore, spatial information updates can be performed manually by the listener or the administrator of the sound space, or they can be triggered by changes in external systems.

[0177] For example, the listener can operate the controller to update the spatial information when their avatar's standing position is instantly distorted, or when it moves forward or backward. Alternatively, the spatial information can be updated when the virtual space administrator implements a performance that suddenly changes the environment of the venue. In these cases, the thread used to update the spatial information managed by the Spatial Information Management Department 1201 can be started either periodically or as a single interrupt.

[0178] The information update thread, which performs spatial information update processing, handles tasks such as updating the position or orientation of the listener's avatar within the virtual space based on the position or orientation of the VR goggles worn by the listener, and updating the position of moving objects within the virtual space. This is provided within a relatively low-frequency processing thread, typically around tens of Hz. Processing reflecting the properties of direct tones can also be performed within this low-frequency processing thread. This is because the frequency of changes in the properties of direct tones is lower than the frequency of audio processing frames used for audio output. This approach reduces the computational load on the processing and avoids the risk of impulsive noise that would result from updating information at unnecessarily high frequencies.

[0179] Figure 4B is a block diagram illustrating another configuration example of the decoder. Specifically, Figure 4B illustrates the configuration of decoder 1210, which is another example of decoder 1112 in Figure 3B or Figure 3D.

[0180] Figure 4B differs from Figure 4A in that the input data 1113 contains unencoded audio signals instead of encoded audio data. Input data 1113 includes a bitstream containing metadata and the audio signal.

[0181] Since the Spatial Information Management Unit 1211 is the same as the Spatial Information Management Unit 1201 in Figure 4A, the description is omitted.

[0182] Since the rendering unit 1213 is the same as the rendering unit 1203 in Figure 4A, its description is omitted.

[0183] Alternatively, decoders 1112, 1200, and 1210 can also be represented as audio processing units that perform audio processing. Furthermore, decoders 1110 and 1130 can also be audio signal processing devices 1001, or can be represented as audio processing devices.

[0184] (Physical Configuration of the Sound Signal Processing Apparatus) Figure 5 is a diagram showing an example of the physical configuration of the sound signal processing apparatus 1001. Alternatively, the sound signal processing apparatus 1001 of Figure 5 could also be the decoding device 1110 of Figure 3B or the decoding device 1130 of Figure 3D. The multiple components shown in Figure 3B or Figure 3D can also be installed using the multiple components shown in Figure 5. Furthermore, a portion of the configuration described herein can also be incorporated into the sound prompting device 1002.

[0185] The audio signal processing device 1001 in Figure 5 includes a processor 1402, a memory 1404, a communication interface 1403, a sensor 1405, and a speaker 1401.

[0186] The processor 1402 is, for example, a CPU, a DSP (Digital Signal Processor), or a GPU (Graphics Processing Unit). The audio processing or decoding processing of this disclosure can also be implemented by the CPU, DSP, or GPU executing a program stored in memory 1404. Furthermore, the processor 1402 is, for example, a circuit that performs information processing. The processor 1402 can also be a dedicated circuit that performs signal processing of sound signals, including the audio processing of this disclosure.

[0187] The memory 1404 may be composed of, for example, RAM (Random Access Memory) or ROM (Read Only Memory). The memory 1404 may also include magnetic recording media such as a hard disk or semiconductor memory such as an SSD. Furthermore, the memory 1404 may be internal memory built into the CPU or GPU. In addition, the memory 1404 may store spatial information managed by the spatial information management units 1201 and 1211. Furthermore, it may also store threshold data, which will be described later.

[0188] The communication IF1403 is, for example, a communication module corresponding to communication methods such as Bluetooth (registered trademark) or WIGIG (registered trademark). The audio signal processing device 1001 communicates with other communication devices, for example, via the communication IF1403, to obtain the bitstream of the decoded object. The obtained bitstream is stored, for example, in the memory 1404.

[0189] The communication IF1403, for example, consists of signal processing circuitry and an antenna corresponding to the communication method. The communication method is not limited to Bluetooth (registered trademark) and WIGIG (registered trademark), but can also be LTE (Long Term Evolution), NR (New Radio), or Wi-Fi (registered trademark), etc.

[0190] Furthermore, the communication method is not limited to the wireless communication methods mentioned above. The communication method can also be a wired communication method such as Ethernet (registered trademark), USB (Universal Serial Bus), or HDMI (registered trademark) (High-Definition Multimedia Interface).

[0191] Sensor 1405 performs sensing to infer the listener's position and orientation. Specifically, sensor 1405 infers the listener's position and / or orientation based on the detection results of one or more of the following: position, orientation, movement, velocity, angular velocity, and acceleration of a part or part of the body or the whole body, and generates position / or orientation information representing the listener's position and / or orientation.

[0192] Alternatively, the sensor 1405 may be an external device of the sound signal processing device 1001. A part of the body may also be the listener's head, etc. The position / or orientation information may be information indicating the listener's position and / or orientation in real space, or it may be information indicating the displacement of the listener's position and / or orientation based on a predetermined point in time. Furthermore, the position / or orientation information may also be information indicating the relative position and / or orientation to the stereo sound reproduction system 1000 or the external device equipped with the sensor 1405.

[0193] Sensor 1405 can be, for example, a camera or a ranging device such as LiDAR (Light Detection and Ranging). Sensor 1405 can also capture images of the listener's head movements and detect these movements by processing the captured images. Furthermore, a device that uses wireless communication in any frequency band, such as millimeter waves, to perform position estimation can also be used as sensor 1405.

[0194] Furthermore, the sound signal processing device 1001 can also obtain position information from an external device equipped with sensor 1405 via communication IF 1403. In this case, the sound signal processing device 1001 may not include sensor 1405. Here, the external device is, for example, the sound prompting device 1002 illustrated in FIG2 or a stereoscopic image playback device worn on the listener's head. In this case, sensor 1405 is configured by combining various sensors such as gyroscope sensors and accelerometer sensors.

[0195] For example, as the speed of the listener's head movement, sensor 1405 can detect the angular velocity of rotation about at least one of three mutually orthogonal axes in the sound space, and can also detect the acceleration of displacement about at least one of the three axes.

[0196] For example, as a measure of head movement in the listener, sensor 1405 can detect rotational motion about at least one of three mutually orthogonal axes in the sound space, and displacement about at least one of these three axes. Specifically, sensor 1405 detects 6DoF position (x, y, z) and angle (yaw, pitch, roll) as the listener's position. Sensor 1405 is constructed by combining various sensors used for motion detection, such as gyroscopes and accelerometers.

[0197] Alternatively, sensor 1405 can be implemented using a camera or a GPS (Global Positioning System) receiver, which detects the listener's location. Location information obtained by inferring location using LiDAR or similar devices as sensor 1405 can also be used. For example, sensor 1405 can be built into a smartphone in the case where the stereo sound reproduction system 1000 is implemented in a smartphone.

[0198] Furthermore, sensor 1405 may also include a temperature sensor such as a thermocouple that detects the temperature of the sound signal processing device 1001. Additionally, sensor 1405 may also include a battery included in the sound signal processing device 1001, or a sensor that detects the remaining amount of the battery connected to the sound signal processing device 1001.

[0199] The loudspeaker 1401 includes, for example, a drive mechanism and an amplifier, such as a diaphragm, a magnet, or a voice coil, which transmits the processed sound signal as a sound prompt to the listener. The loudspeaker 1401 activates the drive mechanism based on the sound signal amplified by the amplifier (more specifically, a waveform signal representing the waveform of the sound), causing the diaphragm to vibrate. The diaphragm, vibrating in response to the sound signal, generates sound waves, which propagate through the air and reach the listener's ear, allowing the listener to perceive the sound.

[0200] Additionally, an example is given here of a sound signal processing device 1001 that includes a speaker 1401 and prompts the sound signal after sound processing via the speaker 1401, but the sound signal prompting mechanism is not limited to the above configuration.

[0201] For example, an audio-processed sound signal can be output to an external sound prompting device 1002 connected via a communication module. Communication via the communication module can be either wired or wireless. Furthermore, as another example, the sound signal processing device 1001 has a terminal for outputting an analog sound signal, to which a cable for an earphone or similar device is connected, and a prompting sound signal is emitted from the earphone or similar device.

[0202] In the above-described cases, the sound prompting device 1002 may also be a headset, earphone, head-mounted display, neck speaker, or wearable speaker worn on the head or part of the listener's body. Alternatively, the sound prompting device 1002 may also be a surround sound system composed of multiple fixed speakers. Furthermore, the sound prompting device 1002 can also reproduce sound signals.

[0203] (Physical configuration of the encoding device) Figure 6 is a diagram showing an example of the physical configuration of the encoding device 1500. The encoding device 1500 of Figure 6 may also be the encoding device 1100 of Figure 3A or the encoding device 1120 of Figure 3C, or the multiple components shown in Figure 3A or Figure 3C may be installed by the multiple components shown in Figure 6.

[0204] The encoding device 1500 in Figure 6 includes a processor 1501, a memory 1503, and a communication IF 1502.

[0205] The processor 1501 is, for example, a CPU, DSP, or GPU. The encoding processing of this disclosure can also be implemented by the CPU, DSP, or GPU executing a program stored in memory 1503. Furthermore, the processor 1501 is, for example, a circuit that performs information processing. The processor 1501 can also be a dedicated circuit that performs signal processing on an audio signal, including the encoding processing of this disclosure.

[0206] The memory 1503 may be composed of, for example, RAM or ROM. The memory 1503 may also include magnetic recording media, such as a hard disk, or semiconductor memory, such as an SSD. Furthermore, the memory 1503 may also be internal memory embedded in a CPU or GPU.

[0207] The communication IF1502 is, for example, a communication module corresponding to communication methods such as Bluetooth (registered trademark) or WIGIG (registered trademark). The encoding device 1500 communicates with other communication devices, for example, via the communication IF1502, and transmits the encoded bit stream.

[0208] The communication IF1502, for example, consists of a signal processing circuit and an antenna corresponding to the communication method. The communication method is not limited to Bluetooth (registered trademark) and WIGIG (registered trademark), but can also be LTE, NR, or Wi-Fi (registered trademark), etc. Furthermore, the communication method is not limited to wireless communication. The communication method can also be wired communication methods such as Ethernet (registered trademark), USB, or HDMI (registered trademark).

[0209] The communication module, for example, consists of a signal processing circuit and an antenna corresponding to the communication method. In the examples above, Bluetooth (registered trademark) or WIGIG (registered trademark) are cited as communication methods, but it can also correspond to communication methods such as LTE (Long Term Evolution), NR (New Radio), or Wi-Fi (registered trademark). Furthermore, the communication IF may not be a wireless communication method as described above, but rather a wired communication method such as Ethernet (registered trademark), USB (Universal Serial Bus), or HDMI (registered trademark) (High-Definition Multimedia Interface).

[0210] [Structure of the rendering unit] Figure 7 is a block diagram showing an example of the structure of the rendering unit 1300. Specifically, Figure 7 shows an example of the detailed structure of the rendering unit 1300 corresponding to the rendering units 1203 and 1213 in Figures 4A and 4B.

[0211] The rendering unit 1300 consists of a resolution unit 1301, a selection unit 1302, and a reproduction unit 1303. It performs additional audio processing on the audio data contained in the input signal and outputs it.

[0212] The input signal may consist of spatial information, sensor information, and sound data. The input signal may also contain a bitstream of sound data and metadata (control information), in which case spatial information may also be included in the metadata.

[0213] Spatial information is information related to the sound space (three-dimensional sound field) formed by the stereo sound reproduction system 1000. It consists of information related to the objects contained in the sound space and information related to the listener. Among the objects, there are sound source objects that emit sound and become sound sources, and non-sound-emitting objects that do not emit sound. Sound source objects can also be simply represented as sound sources.

[0214] Non-sound-producing objects can act as obstacles reflecting the sound emitted by a sound source object, but there are also cases where a sound source object acts as an obstacle reflecting the sound emitted by other sound source objects. Obstacle objects can also be represented as reflecting objects.

[0215] As information shared by both the sound source object and the non-sound-producing object, it includes location information, shape information, and the attenuation rate of the volume when the object reflects the sound.

[0216] Position information is represented by coordinate values ​​along three axes in Euclidean space, such as the X, Y, and Z axes, but it doesn't necessarily have to be three-dimensional. For example, position information can also be two-dimensional, represented by coordinate values ​​along only the X and Y axes. The position information of an object is determined by the representative position of its shape, represented by a mesh or voxels.

[0217] Shape information can also include information related to the material of the surface.

[0218] The attenuation rate can be represented by a real number above 0 and below 1, or by a negative decibel value. Since volume is not amplified by reflection in real space, the attenuation rate is set to a negative decibel value. However, for example, to present the sense of horror in an unreal space, an attenuation rate above 1, i.e., a positive decibel value, can be deliberately set.

[0219] Furthermore, the attenuation rate can be set to a different value for each of the multiple frequency bands, or it can be set independently for each frequency band. Additionally, when setting the attenuation rate for each type of material on the object's surface, the corresponding attenuation rate value can be used based on information related to the surface material.

[0220] Furthermore, spatial information can also include information indicating whether an object is a living organism or whether it is a moving object. If the object is a moving object, the position indicated by the location information can also change over time. In this case, information about the changed position or the amount of change is transmitted to the rendering unit 1300.

[0221] Information related to the sound source object includes not only the information shared by the sound source object and the non-sound-producing object, but also the sound data and the information required to project the sound data into the sound space. Sound data represents information related to the frequency and intensity of the sound, and is data that reflects the sound perceived by the listener.

[0222] The audio data is typically a PCM signal, but it can also be data compressed using encoding methods such as MP3. In this case, the signal needs to be decoded at least before it reaches the reproduction unit 1303, so the rendering unit 1300 may also include a decoding unit (not shown). Alternatively, the signal can also be decoded by the audio data decoder 1202.

[0223] For a sound source object, one sound data point or multiple sound data points can be set. Additionally, recognition information can be assigned to each sound data point, and information related to the sound source object can also include this recognition information.

[0224] Information required to project sound data into the sound space may include, for example, reference volume information used as a reference in the reproduction of sound data, information representing the properties (also called characteristics) of the sound data, information related to the location of the sound source object, and information related to the orientation of the sound source object (i.e., information related to the directivity of the sound emitted by the sound source object).

[0225] Reference volume information can be, for example, the effective value of the amplitude of the sound data at the sound source location when the sound data is radiated into the sound space, or it can be represented as a decibel (dB) value in floating point.

[0226] For example, when the reference volume is 0dB, it can also mean that the volume of the signal level represented by the sound data is not increased or decreased, but the sound is radiated into the sound space at the original volume from the position indicated by the information related to the location of the sound source object. Alternatively, when the reference volume is -6dB, it can also mean that the volume of the signal level represented by the sound data is set to approximately half, and the sound is radiated into the sound space from the position indicated by the information related to the location of the sound source object.

[0227] The reference volume information can be assigned to each sound data point individually, or it can be assigned to multiple sound data points uniformly.

[0228] Information representing the properties of sound data can be, for example, information related to the volume of the sound source, and information representing the time-series variation of the sound source's volume.

[0229] For example, in a virtual conference room where the sound space is a speaker and the sound source is the speaker, the volume changes intermittently over a short period of time. That is, the audible and silent parts alternate. Conversely, in a concert hall where the sound space is a concert hall and the sound source is a performer, the volume is maintained for a certain duration. Furthermore, in a battlefield where the sound space is an explosive device and the sound source is an explosive device, the volume of the explosion sound only increases momentarily, then remains silent or low.

[0230] In this way, the volume information of the sound source includes not only the magnitude of the sound, but also information about changes in the magnitude of the sound. This information can also be used to represent the properties of the sound data.

[0231] Transition information can also be represented by time-series data showing frequency characteristics. Transition information can also be represented by data showing the duration of the sound interval. Transition information can also be represented by time-series data showing the duration of both the sound and silent intervals. Transition information can also be represented by multiple sets of time-series data listing the duration during which the amplitude of the sound signal can be considered stationary (approximately constant) and the amplitude values ​​of the signal during that period.

[0232] Transition information can also be represented by data that allows the frequency characteristics of a sound signal to be considered as a stationary duration. Transition information can also be represented by listing multiple sets of data, in time series, of the duration during which the frequency characteristics of the sound signal can be considered stationary and the frequency characteristics during that period. Transition information can also be represented, for example, by data showing the approximate shape of a spectrogram.

[0233] Alternatively, the volume used as the reference for the aforementioned frequency characteristics can also be the aforementioned reference volume. Information about the reference volume and information representing the properties of the sound data can be used for calculating the volume of direct or reflected sounds perceived by the listener, or for selecting whether or not to allow the listener to perceive them. Other examples and methods of using information representing the properties of the sound data will be described later.

[0234] Furthermore, the reflected sound in this embodiment is an example of an indirect sound. An indirect sound can be a reflected sound, a diffracted sound, or the like. In this embodiment, a reflected sound, as an example of an indirect sound, is used for explanation, but the same treatment is performed even if an indirect sound is used instead of a reflected sound.

[0235] Information related to the orientation of a sound source object (orientation information) is typically represented by yaw, pitch, and roll. Alternatively, the roll rotation can be omitted, and the orientation information of the sound source object can be represented by azimuth (yaw) and pitch (pitch). The orientation information of the sound source object can also change over time, and any changes are transmitted to the rendering unit 1300.

[0236] Information relevant to the listener relates to their position and orientation within the sound space. Position information is represented by the XYZ axes of Euclidean space, but it doesn't necessarily have to be three-dimensional; it can also be two-dimensional. Orientation information is typically represented by yaw, pitch, and roll. Alternatively, the roll can be omitted, and the listener's orientation information can be represented by azimuth (yaw) and pitch (pitch).

[0237] The listener's location and orientation information can also change over time, and if changes occur, they will be transmitted to the rendering unit 1300.

[0238] The sensor information includes information such as the amount of rotation or displacement detected by the sensor 1405 worn by the listener, as well as the listener's position and orientation. The sensor information is transmitted to the rendering unit 1300, which updates the listener's position and orientation information based on the sensor information. The sensor information may also include, for example, position information obtained by the portable terminal using GPS, a camera, or LiDAR for self-position estimation.

[0239] Alternatively, instead of sensor 1405, information obtained from an external source via a communication module can be used as sensor information for detection. Information indicating the temperature of the sound signal processing device 1001 and the remaining battery level can also be obtained from sensor 1405. Furthermore, the computing resources (CPU capacity, memory resources, or PC performance, etc.) of the sound signal processing device 1001 or the sound prompting device 1002 can be obtained in real time.

[0240] The analysis unit 1301 analyzes the sound signal contained in the input signal and the spatial information received from the spatial information management units 1201 and 1211, and calculates the information required for generating direct sound and reflected sound in the reproduction unit 1303, as well as the information required for selecting whether to generate reflected sound.

[0241] The information required for the generation of direct and reflected tones includes values ​​related to the path taken to reach the listening position, the time taken to reach the position, and the volume upon arrival for each direct and reflected tone. These values ​​represent, for example, the path taken to reach the listening position, the time taken to reach the position, and the volume upon arrival.

[0242] The information required for selecting the output reflected tone is information representing the relationship between the direct tone and the reflected tone, such as values ​​related to the time difference between the direct tone and the reflected tone, and values ​​related to the volume ratio of the direct tone and the reflected tone at the listening position. For example, the values ​​related to the time difference between the direct tone and the reflected tone, and the values ​​related to the volume ratio of the direct tone and the reflected tone at the listening position, are respectively values ​​representing the time difference between the direct tone and the reflected tone, and values ​​representing the volume ratio of the direct tone and the reflected tone at the listening position.

[0243] Furthermore, when volume is expressed in decibels on a logarithmic axis (in the case of representing volume in the decibel range), the volume ratio of two signals is naturally represented by the difference in decibel values. Specifically, the volume ratio of two signals can be the difference between the amplitude values ​​of each signal expressed in the decibel range. This value can also be calculated based on energy values ​​or power values, etc. Moreover, this difference can be referred to as the gain difference or simply the gain difference in the decibel range.

[0244] That is, the volume ratio in this disclosure is essentially the ratio of signal amplitudes, so it can also be expressed as Soundvolume ratio, Volume ratio, Amplitude ratio, Soundlevel ratio, Sound intensity ratio, or Gain ratio, etc. Furthermore, when the unit of volume is decibels, the volume ratio in this disclosure can of course also be referred to as volume difference.

[0245] In this disclosure, "volume ratio" typically refers to the gain difference between two sounds expressed in decibels. In examples of embodiments, the threshold data is also typically defined by the gain difference expressed in decibels. However, the volume ratio is not limited to the gain difference in decibels. When using a volume ratio expressed outside the decibel range, the threshold data defined in the decibel range can be converted to the units of the calculated volume ratio for use. Alternatively, the threshold data defined in each unit can be stored in memory in advance.

[0246] That is, even if a ratio such as energy value or power value is used instead of volume ratio, it is obvious that the algorithm in this disclosure can be applied to the solution of the problem of this disclosure.

[0247] The time difference between the arrival of the direct tone and the reflected tone is, for example, the time difference between the arrival time of the direct tone (arrival moment) and the arrival time of the reflected tone (arrival moment). Alternatively, for simplicity, the time difference between the arrival of the direct tone and the reflected tone is sometimes recorded as the time difference between the direct tone and the reflected tone. The time difference between the direct tone and the reflected tone can also be the time difference between the moments when the direct tone and the reflected tone arrive at the listening position, the difference in time required for the direct tone and the reflected tone to arrive at the listening position, or the time difference between the moment the direct tone ends and the moment the reflected tone arrives at the listening position. The methods for calculating these values ​​will be described later.

[0248] The selection unit 1302 uses the information calculated by the analysis unit 1301 and the threshold data to select whether the reproduction unit 1303 will generate the reflected sound. In other words, the selection unit 1302 determines whether to select the reflected sound as the target reflected sound for generation. In other words, the selection unit 1302 selects which of the multiple reflected sounds generated by the reproduction unit 1303 will be the reflected sound.

[0249] Threshold data, for example, can be represented in a graph where the horizontal axis represents the time difference between the direct and reflected sounds, and the vertical axis represents the volume ratio of the direct and reflected sounds. This represents the boundary (threshold) at which the reflected sound is perceived or not. Threshold data can also be represented by an approximation using the time difference between the direct and reflected sounds as a variable, or by an arrangement of values ​​indexed by the time difference between the direct and reflected sounds and corresponding thresholds.

[0250] Selection unit 1302 selects to generate a reflected sound if, for example, the ratio of the volume of the direct sound at arrival to the volume of the reflected sound at arrival is greater than a threshold set based on reference threshold data, among the values ​​of the time difference between the arrival time of the direct sound and the arrival time of the reflected sound. Furthermore, the arrival volume refers to the volume of the sound when it reaches the listening position.

[0251] The time difference between the arrival time of the direct tone and the arrival time of the reflected tone is, in other words, the difference in time required for the direct tone and the reflected tone to reach the listening position respectively. Alternatively, the time difference between the end of the direct tone's emission and the arrival time of the reflected tone at the listening position can also be used as the time difference between the direct tone and the reflected tone. In this case, different threshold data can be used than the threshold data set based on the time difference between the arrival times of the direct tone and the reflected tone, or a common threshold data can be used.

[0252] The threshold data can be obtained from the memory 1404 of the audio signal processing device 1001, or from an external storage device via the communication module. The method for storing the threshold data and the method for setting the threshold will be described later.

[0253] The reproduction unit 1303 synthesizes the direct sound signal with the reflected sound signal selected and generated by the selection unit 1302.

[0254] Specifically, the reproduction unit 1303 processes the input sound signal to generate a direct tone based on the information about the arrival time and volume of the direct tone calculated by the analysis unit 1301. Furthermore, the reproduction unit 1303 processes the input sound signal to generate a reflected tone based on the information about the arrival time and volume of the reflected tone selected by the selection unit 1302. Then, the reproduction unit 1303 combines the generated direct tone and reflected tone and outputs them.

[0255] [Example of Rendering Unit Operation] Figure 8 is a flowchart showing an example of the operation of the sound signal processing device 1001. Figure 8 shows the processing mainly performed by the rendering unit 1300 of the sound signal processing device 1001.

[0256] In the input signal analysis process (S101 in FIG8), the analysis unit 1301 analyzes the input signal input to the sound signal processing device 1001 and detects direct tones and reflected tones that can be generated in the sound space. The reflected tones detected here are candidate reflected tones selected by the selection unit 1302 as the reflected tones to be generated by the reproduction unit 1303. In addition, the analysis unit 1301 analyzes the input signal and calculates the information needed for the generation of direct tones and reflected tones and the information needed for the selection of the reflected tones to be generated.

[0257] First, the characteristics of direct and reflected tones are calculated. Specifically, the arrival time and volume of the direct and reflected tones when they reach the listener are calculated. If multiple objects exist in the sound space as reflection targets, the characteristics of the reflected tones are calculated for each object separately.

[0258] The direct arrival time (td) is calculated based on the direct arrival path (pd). The direct arrival path (pd) is the path connecting the location information S(xs, ys, zs) of the sound source object with the location information A1(xa, ya, za) of the listener. The direct arrival time (td) is the value obtained by dividing the length of the path connecting the location information S(xs, ys, zs) and the location information A1(xa, ya, za) by the speed of sound (approximately 340 m / s).

[0259] For example, the path length (X) is calculated using (xs-xa)^2 + (ys-ya)^2 + (zs-za)^2)^0.5. Volume decreases inversely with distance. Therefore, given the volume as N in the location information S(xs, ys, zs) of the sound source object and the unit distance as U, the volume (ld) when the direct sound arrives is calculated using ld = N. Find U / X.

[0260] The volume N at the sound source location can also be the reference volume described earlier.

[0261] The arrival time (tr) of the reflected sound is calculated based on the arrival path (pr). The arrival path (pr) is the path that connects the position of the sound image of the reflected sound with the position information A1 (xa, ya, za).

[0262] Furthermore, the location of the sound image of reflected sound can be derived using methods such as the "mirror method" or the "ray tracing method," or any other method. The mirror method assumes that the reflected wave on the wall of a room has a mirror image at a position symmetrical to the sound source relative to the wall, and simulates the sound image by assuming that sound waves are emitted from that mirror image location. The ray tracing method simulates the image (sound image) observed at a certain point by tracing waves that propagate in straight lines, such as light rays or sound rays.

[0263] Figure 9 shows the relative positions of the listener and the obstacle. Figure 10 shows the relative positions of the listener and the obstacle. In other words, Figures 9 and 10 respectively illustrate examples of sound images formed at positions symmetrical to the sound source location, separated by a wall. By determining the position of the sound image of the reflected sound on the xyz axes based on this relationship, the arrival time of the reflected sound can be calculated in the same way as the arrival time of the direct sound.

[0264] The arrival time (tr) of the reflected sound is obtained by dividing the length (Y) of the path connecting the position of the sound image of the reflected sound to the position information A1 (xa, ya, za) by the speed of sound (approximately 340 m / s). Volume decreases inversely with distance. Therefore, when the volume at the sound source is N, the unit distance is U, and the attenuation rate of the volume in the reflection is G, the volume (lr) when the reflected sound arrives decreases by lr = N. G Find U / Y.

[0265] As explained above, the attenuation rate G can be represented by a real number greater than or equal to 0 and less than 1, or by a negative decibel value. In this case, the overall volume attenuation of the signal corresponds to the amount of G. Furthermore, the attenuation rate can also be set for each of the multiple frequency bands. In this case, the analysis unit 1301 applies the specified attenuation rate for each frequency component of the signal. Additionally, to reduce computational complexity, the analysis unit 1301 can use representative values ​​or average values ​​of multiple attenuation rates from multiple frequency bands as the overall attenuation rate, thereby correspondingly attenuating the overall volume of the signal.

[0266] Next, the analysis unit 1301 calculates the volume ratio (L) required in the selection of the reflected sound of the generated object, namely the ratio of the volume (ld) when the direct sound arrives to the volume (lr) when the reflected sound arrives, and the time difference (T) between the direct sound and the reflected sound.

[0267] The ratio of the volume (ld) of the direct tone to the volume (lr) mentioned above, i.e., the volume ratio (L), is, for example, the value obtained by dividing the volume (lr) of the reflected tone by the volume (ld) of the direct tone, given by L = (N G U / Y) / (N) U / X) = G X / Y is calculated. Since the calculated value is a volume ratio, the values ​​of N and U can be any pre-set values.

[0268] The time difference (T) between the direct tone and the reflected tone can also be the time difference between the direct tone and the reflected tone when they arrive at the listening position. For example, the time difference (T) between the direct tone and the reflected tone when they arrive at the listening position can be calculated by T = tr - td.

[0269] Furthermore, the time difference (T) can also be the difference between the arrival times of the direct tone and the reflected tone at the listening position. Alternatively, the time difference (T) can also be the time difference between the end of the direct tone's speech and the arrival time of the reflected tone at the listening position. In other words, the time difference (T) can also be the time difference between the end of the direct tone and the beginning of the reflected tone at the listening position.

[0270] Next, in the selection process for the reflected sound (S102 in FIG8), the selection unit 1302 selects whether the reflected sound calculated by the analysis unit 1301 is generated by the reproduction unit 1303. In other words, the selection unit 1302 determines whether to select the reflected sound as the target reflected sound for generation. If there are multiple reflected sounds, the selection unit 1302 selects whether to generate each reflected sound. The result of the selection unit 1302's selection of whether to generate each reflected sound can be either selecting more than one target reflected sound from multiple reflected sounds, or selecting none of the target reflected sounds.

[0271] Furthermore, the selection unit 1302 is not limited to generation processing; it can also select reflected sounds that are the application targets of other processing. For example, the selection unit 1302 can also select reflected sounds that are the application targets of binaural processing. In addition, the selection unit 1302 generally selects only one or more reflected sounds that are the processing targets. However, the selection unit 1302 can also select only one or more reflected sounds that are not the processing targets. Furthermore, processing can be applied to one or more reflected sounds that are not selected.

[0272] For example, the selection of reflected tones is based on the volume ratio (L) and time difference (T) calculated by the analysis unit 1301. By performing selection processing based on the time difference (T) between the direct tone and the reflected tone, it is possible to more appropriately select reflected tones that have a greater impact on the listener's perception compared to the case where selection processing is based solely on the volume difference between the direct tone and the reflected tone.

[0273] Specifically, the decision to generate a reflected tone is made, for example, by comparing the volume ratio of the direct tone to the reflected tone, corresponding to the time difference between the direct tone and the reflected tone, with a pre-set threshold. The threshold is set with reference to threshold data. The threshold data is an indicator representing the boundary at which the reflected tone of a direct tone is perceived by the listener, defined by the ratio of the volume (Id) when the direct tone arrives to the volume (lr) when the reflected tone arrives.

[0274] Furthermore, a threshold corresponds to a value expressed as a numerical value set in relation to a time difference (T). Threshold data corresponds to the relationship between the time difference (T) and the threshold, and to tabular data or formulas used to determine or calculate the threshold under the time difference (T). The form and type of threshold data are not limited to tabular data or formulas.

[0275] Figure 11 is a graph showing the relationship between the time difference between the direct tone and the reflected tone and a threshold. For example, the threshold data for the volume ratio preset according to each value of the time difference between the direct tone and the reflected tone, as shown in Figure 11, can also be referenced. Alternatively, threshold data can be obtained from the threshold data shown in Figure 11 through interpolation or extrapolation.

[0276] Furthermore, the threshold for the volume ratio under the time difference (T) calculated by the analysis unit 1301 is determined based on the threshold data. The selection unit 1302 then decides whether to select the reflected sound as the target reflected sound based on whether the volume ratio (L) of the direct sound to the reflected sound calculated by the analysis unit 1301 is higher than this threshold.

[0277] By using threshold data of volume ratios preset according to each value of the time difference between direct and reflected tones, selection processing that takes into account backmasking or priority effects can be achieved. Detailed explanations of the types, formats, storage methods, and setting methods of the threshold data will be provided later.

[0278] Next, in the direct tone and reflected tone generation process (S103 in FIG8), the reproduction unit 1303 generates the sound signal of the direct tone and the sound signal of the reflected tone selected by the selection unit 1302 as the target of generation, and synthesizes them.

[0279] The direct tone sound signal is generated by applying the direct tone arrival time (td) and the direct tone arrival volume (ld) calculated by the analysis unit 1301 to the sound data of the sound source object contained in the input information. Specifically, the sound data is delayed by the direct tone arrival time (td) and multiplied by the direct tone arrival volume (ld). The sound data delay process is a process of shifting the position of the sound data back and forth on the time axis. For example, the sound data delay process disclosed in Patent Document 2 without degrading the sound quality can also be applied.

[0280] The sound signal of the reflected sound is generated in the same way as the direct sound by applying the sound data of the sound source object to the arrival time (tr) of the reflected sound and the volume (lr) of the reflected sound when it arrives, which is calculated by the analysis unit 1301.

[0281] However, the volume (lr) of the reflected sound during its arrival differs from the volume (ld) of the direct sound. It is affected by the attenuation rate G applied to the volume of the reflected sound. G can be an attenuation rate applied across the entire frequency band. Alternatively, it can be a reflection rate specified for each defined frequency band to reflect the bias in the frequency components generated by reflection. In this case, the application of the volume (lr) of the reflected sound can also be implemented as a frequency equalizer process, multiplying the volume by the attenuation rate for each frequency band.

[0282] In the example above, the path lengths of the direct and reflected tone candidates as they reach the listener are calculated. Then, the arrival time and volume are calculated based on each path length. Finally, the reflected tone candidates are selected based on their time difference and volume ratio.

[0283] Alternatively, as another example, selection can be performed based on the path lengths of the direct and reflected sounds as they reach the listener, omitting the calculations of arrival time and volume, as well as time difference and volume ratio. In this case, a threshold corresponding to the path length difference can be pre-set for the path length ratio. Furthermore, selection can be performed based on whether the calculated path length ratio is above the threshold corresponding to the calculated path length difference. Thus, selection can be performed based on the path length difference corresponding to the time difference while reducing computational load.

[0284] In addition to the path length difference, the value of a parameter representing the speed of sound propagation, or the value of a parameter that affects the speed of sound propagation, can also be used.

[0285] (Details of the selection process) This section explains the details of the selection process for whether or not to generate reflected sounds.

[0286] The selection of the reflected sound is performed by comparing a threshold value for the volume ratio of the direct sound to the reflected sound at a given time difference (T), i.e., a volume ratio threshold, with a volume ratio (L) calculated by the analysis unit 1301. For example, among the volume ratio thresholds preset for each value of the time difference between the direct sound and the reflected sound, the threshold value for the volume ratio at the time difference (T) between the direct sound and the reflected sound calculated by the analysis unit 1301 is referenced. Furthermore, whether the reflected sound is selected as the target reflected sound is determined based on whether the volume ratio (L) calculated by the analysis unit 1301 is higher than the threshold value.

[0287] The time difference (T) can be, for example, the difference in the time when the direct tone and the reflected tone arrive at the listening position, the time difference in the time required for the direct tone and the reflected tone to arrive at the listening position, or the time difference between the time when the direct tone ends and the time when the reflected tone arrives at the listening position. Here, the end time of the direct tone can also be obtained, for example, by adding the duration of the direct tone to the time of its arrival.

[0288] Regarding threshold data, it can also be determined through auditory nerve action or cognitive function of the brain. More specifically, it can be determined based on the minimum time difference between two sounds that the listener can perceive and detect, through the prioritization effect, the time-dependent masking phenomenon, or a combination thereof (described later). Specific values ​​can be derived from known research findings on time-dependent masking effects, prioritization effects, or echo detection limits, or they can be determined through audiovisual experiments applied to this virtual space.

[0289] Figures 12A, 12B, and 12C are diagrams illustrating examples of methods for setting threshold data. As shown in Figures 12A, 12B, and 12C, the threshold data is represented by the boundary (threshold) at which the reflected sound is perceived or not in a graph where the horizontal axis represents the time difference between the direct sound and the reflected sound, and the vertical axis represents the volume ratio between the direct sound and the reflected sound.

[0290] The threshold data can also be represented by an approximation that uses the time difference between the direct tone and the reflected tone as a variable. Furthermore, the threshold data can also be stored in a region of memory 1404 as an index of the time difference between the direct tone and the reflected tone, as shown in Figure 11, and an arrangement of thresholds corresponding to those indices.

[0291] Furthermore, in Example 4 of Figure 12C, where the height of the line parallel to the horizontal axis (the minimum audible limit) is used as the threshold, the volume ratio (L) of the direct tone to the reflected tone is not compared with the threshold; rather, the volume of the reflected tone itself is compared with the threshold. This is because the threshold represents the volume boundary of whether a sound can be perceived by a listener, and it is the threshold used to determine sounds with volumes lower than this threshold as unreproducible sounds. That is, the threshold corresponding to the minimum audible limit is not a threshold for the ratio of the volume of the reflected tone to the volume of the direct tone.

[0292] When the minimum audible limit is used as the threshold, the threshold is constant and independent of the time difference (T), so the time difference (T) can be disregarded.

[0293] Furthermore, when multiple reflected sounds are generated through analytical processing (S101 in Figure 8), selection processing can be performed on all reflected sounds, or selection processing can be performed only on reflected sounds with high evaluation values ​​derived from each reflected sound using a pre-set evaluation method. Here, the evaluation value of a reflected sound corresponds to its perceived importance. Additionally, a high evaluation value corresponds to a large evaluation value, and these distinctions can be interchanged.

[0294] The selection unit 1302 may, for example, calculate the evaluation value of the reflected sound using a pre-set evaluation method corresponding to the volume of the sound source, the visuality of the sound source, the locality of the sound source, the visuality of the reflecting object (obstacle object), or the geometric relationship between the direct sound and the reflected sound.

[0295] Specifically, a higher volume of the sound source can result in a higher evaluation value. Furthermore, to ensure consistency between visual and acoustic localization, a higher evaluation value can also be achieved when the sound source object or its reflection (obstacle) is visible to the listener, or when the sound source object has high localization.

[0296] Furthermore, the opening of the angles of arrival of the direct and reflected tones, as well as the difference in their arrival times, significantly impact spatial perception. Therefore, a higher evaluation value can be achieved when the openings of the angles of arrival of the direct and reflected tones are large, or when the difference in their arrival times is significant.

[0297] The volume information of the audio source can also represent the base volume set for each content, the time transition of the volume, or both.

[0298] For example, in a virtual conference room where the virtual space is a virtual meeting room and the direct audio is conversational sound, the volume changes intermittently over a short period of time. That is, the audible and silent parts alternate. Conversely, in a concert hall where the virtual space is a concert hall and the direct audio is a musical performance, the volume is maintained for a certain duration. Finally, in a battlefield where the virtual space is a battlefield and the direct audio is an explosion, the volume increases only momentarily, then remains silent or low.

[0299] In this way, the volume information of the sound source not only includes information about the reference volume that corresponds to the volume setting when the sound is radiated into the virtual space, but also information about the changes in the volume of the sound.

[0300] Transition information can also be represented by time-series data showing frequency characteristics. Transition information can also be represented by data showing the duration of the sound interval. Transition information can also be represented by time-series data showing the duration of both the sound and silent intervals. Transition information can also be represented by multiple sets of time-series data listing the duration during which the amplitude of the sound signal can be considered stationary (approximately constant) and the amplitude values ​​of the signal during that period.

[0301] Information about the transition can also be represented by data that shows the duration of a stationary frequency response of a sound signal. Alternatively, it can be represented by a time series listing multiple sets of data showing the duration of a stationary frequency response of a sound signal and the frequency response during that period.

[0302] Furthermore, there has been a long-standing and widespread practice of using temporal shifts in the frequency characteristics of signals for audio processing in virtual space (Patent Document 1, etc.). Given this prior art, the aforementioned group could also be a group of time lengths with constant frequency characteristics and their frequency characteristics.

[0303] Geometric relationships can also refer to the relationships between the positions of a sound source, a listener, and a reflecting object within a virtual space. Through these relationships, the path lengths of both direct and reflected sounds can be calculated geometrically. Therefore, by utilizing the inverse relationship between volume and distance, the reference volume of the reflected sound relative to the reference volume of the direct sound can be calculated.

[0304] In calculating the reference volume of reflected sound, the reflection coefficient of the reflecting object can also be used. Alternatively, a typical value commonly used can be used as the reflection coefficient. On the other hand, in special cases such as when the reflecting object is covered by sound-absorbing material, a specially assigned reflection coefficient can be used as the reflection coefficient of the reflecting object.

[0305] Reflected sound can also be evaluated based on its volume. The volume of the reflected sound can be determined based on the geometric relationship between the direct sound and the reflected sound as described above, as well as the index assigned to the reflecting object. The volume can also be compared to a pre-set threshold to evaluate the reflected sound.

[0306] Furthermore, information representing the temporal change in the volume of the sound source can also be reflected in the evaluation. For example, if the information representing the temporal change in the volume of the sound source represents the duration of the sound interval, the evaluation value of the reflected sound can remain unchanged when the time is within the sound interval. On the other hand, if the time is outside the sound interval, even if the reference volume of the reflected sound exceeds a threshold, the evaluation value of the reflected sound can be reduced or set to zero.

[0307] Alternatively, information representing the temporal shift in the volume of a sound source can also be data that lists multiple groups of amplitude values ​​of the signal over a duration during which the amplitude of the sound signal is considered approximately constant, in a time series. In this case, the processing of the reflected sound can also be evaluated by changing the reference volume of the reflected sound in conjunction with the changes in the amplitude values ​​in the data.

[0308] Alternatively, the volume information representing the direct tone can be obtained by using both the reference volume information and the volume information that changes over time. For example, after calculating an evaluation value based on the reference volume information, the evaluation value can be corrected using the volume information that changes over time.

[0309] In evaluating reflected sounds, all of the above methods can be performed, or only some of them can be performed. For example, reflected sounds can be evaluated using multiple evaluation methods, or they can be evaluated using only one evaluation method.

[0310] When evaluating reflected sounds using multiple evaluation methods, the decision to select a reflected sound can be based on the combined evaluation value determined by the multiple evaluation methods, or it can be based on the individual evaluation values ​​of the multiple evaluation methods.

[0311] The sound signal processing device 1001 may select a sound if, when deciding whether to select a reflected sound based on each of multiple evaluation methods, all evaluation results based on multiple evaluation methods indicate that sound should be selected. Alternatively, the sound signal processing device 1001 may select a sound if one of the evaluation results based on multiple evaluation methods indicates that sound should be selected.

[0312] Alternatively, priorities can be set for the first to third evaluation methods. Furthermore, the sound signal processing device 1001 can also ultimately determine that sound is not selected if the first evaluation method determines that sound is not selected, regardless of the determination results in the second and third evaluation methods. Additionally, the sound signal processing device 1001 can also ultimately determine that sound is selected if one of the second and third evaluation methods determines that sound is not selected, but the other determines that sound is selected.

[0313] Furthermore, selection processing and evaluation processing can be performed independently, or only one of them can be performed. Alternatively, evaluation processing can be performed only on reflected sounds that were determined to be selected in the selection processing, and the evaluation processing can then determine whether to select the reflected sound again. Or, evaluation processing can be performed only on reflected sounds that were determined not to be selected in the selection processing, and the evaluation processing can then determine whether to select the reflected sound again.

[0314] The selection process described above can be interpreted as selecting reflected sounds based on the properties of the direct sound. For example, in selecting reflected sounds based on the properties of the direct sound, a threshold used in the selection of reflected sounds is set or adjusted according to the properties of the direct sound. Alternatively, an evaluation value used in the selection of reflected sounds can be calculated based on one or more of the following: the volume of the sound source, the visibility of the sound source, the localization of the sound source, the visibility of the reflecting object (obstacle object), and the geometric relationship between the direct sound and the reflected sound.

[0315] Furthermore, the process of selecting reflected tones based on the properties of direct tones is not limited to setting or adjusting thresholds based on the properties of direct tones, or calculating evaluation values ​​used in selecting reflected tones of the processed object; other processes may also be performed. Moreover, when performing processes such as setting or adjusting thresholds based on the properties of direct tones, or calculating evaluation values ​​used in selecting reflected tones of the processed object, a portion of the process may be modified, or new processes may be added.

[0316] In addition, setting thresholds can also include adjusting thresholds and changing thresholds.

[0317] [Threshold setting method] The threshold data used in the selection process can be set by referring to known values ​​of echo detection limits based on priority effects or masking thresholds based on backmasking effects.

[0318] The priority effect refers to the phenomenon where, when two sounds are heard from different locations, the listener perceives which sound source was heard first in time. If two short sounds merge and sound like one, the overall sound's perceived location (locality) is largely determined by the location of the initial sound. The echo detection limit, a phenomenon occurring through the priority effect, is the minimum time difference that allows a listener to perceive the discrepancy between two sounds.

[0319] In Example 2 of Figure 12C, the horizontal axis corresponds to the arrival time of the reflected sound (echo), specifically, the delay time from the arrival time of the direct sound to the arrival time of the reflected sound. The vertical axis corresponds to the volume ratio of the detectable reflected sound to the direct sound, specifically, the threshold for whether the reflected sound arriving with the delay time can be detected.

[0320] Figure 13 is a diagram illustrating an example of a threshold setting method. The horizontal axis in Figure 13 corresponds to the arrival time of the reflected sound, specifically the time difference (T) between the direct sound and the reflected sound. The vertical axis in Figure 13 corresponds to the volume of the reflected sound. Specifically, the vertical axis in Figure 13 can correspond either to the volume of the reflected sound (volume ratio) set relatively to the volume of the direct sound, or to the volume of the reflected sound determined absolutely regardless of the volume of the direct sound.

[0321] For example, when the listener is relatively far from the obstacle, as shown in Figure 9, the arrival time of the reflected sound is later, as shown in Figure 13C, and the threshold is set low. As a result, a reflected sound is generated in the case of Figure 9. On the other hand, when the listener is relatively close to the obstacle, as shown in Figure 10, the arrival time of the reflected sound is earlier than in the case of Figure 9, as shown in Figure 13B, and the threshold is set high. As a result, no reflected sound is generated in the case of Figure 10.

[0322] In addition, threshold data can also be stored in memory 1404 and retrieved from memory 1404 for use in selection processing.

[0323] Figure 14 is a flowchart illustrating an example of the selection process. First, the selection unit 1302 selects the reflected tone detected by the analysis unit 1301 (S201). Next, the selection unit 1302 detects the volume ratio (L) of the direct tone and the reflected tone, as well as the time difference (T) between the direct tone and the reflected tone (S202 and S203).

[0324] The time difference (T) can be, for example, the time difference between the arrival time of the direct tone and the reflected tone at the listening position, the time difference between the arrival time of the direct tone and the arrival time of the reflected tone, or the time difference between the end of the direct tone's sound and the arrival time of the reflected tone at the listening position. An example based on the time difference between the arrival time of the direct tone and the arrival time of the reflected tone is given here.

[0325] Specifically, the selection unit 1302 calculates the difference between the length of the direct sound path and the length of the reflected sound path based on the location information of the sound source object and the listener, as well as the location and shape information of the obstacle object. Furthermore, the selection unit 1302 detects the time difference (T) between the time when the direct sound arrives at the listener's position and the time when the reflected sound arrives at the listener's position by dividing this length by the speed of sound.

[0326] The volume reaching the listener decreases proportionally (inversely proportional to distance) to the volume of the sound source. Therefore, the volume of the direct tone is obtained by dividing the volume of the sound source by the length of the direct tone's path. The volume of the reflected tone is obtained by dividing the volume of the sound source by the length of the reflected tone's path and then multiplying by the attenuation rate assigned to the virtual obstacle object. The selection unit 1302 detects the volume ratio by calculating the ratio of their volumes.

[0327] Furthermore, the selection unit 1302 uses threshold data to determine a threshold corresponding to the time difference (T) (S204). Next, the selection unit 1302 determines whether the detected volume ratio (L) is above the threshold (S205).

[0328] When the volume ratio (L) is above the threshold ("Yes" in S205), the selection unit 1302 selects the reflected sound as the reflected sound of the generation target (S206). When the volume ratio (L) is below the threshold ("No" in S205), the selection unit 1302 does not select the reflected sound as the reflected sound of the generation target (S207). That is, in this case, the selection unit 1302 determines the reflected sound as a reflected sound other than the generation target.

[0329] Then, the selection unit 1302 determines whether there is an unspecified reflected sound (S208). If there is an unspecified reflected sound ("Yes" in S208), the selection unit 1302 repeats the above process (S201 to S207). If there is no unspecified reflected sound ("No" in S208), the selection unit 1302 ends the process.

[0330] This selection process can be performed on all reflected sounds generated in the parsing process, or only on reflected sounds with high evaluation values.

[0331] [Details of the Threshold Storage Method] The threshold data in this embodiment is stored in the memory 1404 of the sound signal processing apparatus 1001. The form and type of the stored threshold data can be arbitrary. When multiple forms and types of thresholds are stored, it is possible to determine which form and type of threshold to use for the selection processing of reflected sounds during the selection process. The method for determining which threshold data to use for the selection process will be described later.

[0332] Furthermore, threshold data of multiple forms and types can be combined and stored. The combined threshold data can also be read from the spatial information management units 1201 and 1211 and set as the threshold used in the selection process. Alternatively, threshold data stored in the memory 1404 can be stored in the spatial information management units 1201 and 1211.

[0333] Threshold data can also be stored, for example, as thresholds at various time differences, to depict lines of thresholds as shown in [Example 1] and [Example 2] of Figure 12C.

[0334] Furthermore, the threshold data can also be stored as a table data that establishes a correspondence between the threshold and the time difference (T), as shown in Figure 11. That is, the threshold data can also be stored as table data with the time difference (T) as an index. Of course, the threshold shown in Figure 11 is an example, and the threshold is not limited to the example in Figure 11. In addition, the threshold itself can be approximated by a function with the time difference (T) as a variable, and the coefficients of the function can be stored. Furthermore, multiple approximations can be combined and stored.

[0335] For example, the time difference (T) can be set as timeDiff, and the threshold can be set as gainThresh, and the threshold data can be represented by the following formula.

[0336] [Mathematical Expression 1] The threshold is defined solely by the time range in which the priority effect occurs. When the time difference is outside this time range (a value less than 1 ms or greater than 40 ms in the above formula), a judgment based on gainThresh may not be performed, and the judgment may be made solely by the threshold representing the minimum volume reproduced in the virtual space, as described later.

[0337] Through experiments, the inventors have determined that, within the timeframe in which the preference effect occurs, the threshold is preferably approximated by an upwardly convex function. The above formula is an example of an approximation generated based on this experiment.

[0338] The memory 1404 may also store information related to the relational formulas representing the relationship between time difference (T) and threshold. That is, it may also store formulas that take time difference (T) as a variable. The threshold for each time difference (T) may also be approximated by a straight line or curve, and parameters representing the geometric shape of the straight line or curve may be stored. For example, if the geometric shape is a straight line, the starting point and slope of the straight line may also be stored.

[0339] Furthermore, the type and format of threshold data can be set and stored according to each property of the direct tone. Additionally, parameters used to adjust the threshold based on the property of the direct tone and for selection processing can be stored. The process of adjusting the threshold based on the property of the direct tone and for selection processing will be described later as a variation of the threshold setting method.

[0340] As an example of storing multiple threshold data in combination, the larger of the masking threshold and the echo detection limit threshold can be stored for each time difference (T), as shown in Example 3 of Figure 12C. Alternatively, the larger of the minimum volume reproduced in virtual space and the echo detection limit threshold can be stored for each time difference (T), as shown in Example 4 of Figure 12C.

[0341] The combination of multiple types of threshold data is not limited to this. For example, information on the maximum value can also be stored in multiple threshold data for each time difference (T).

[0342] Furthermore, in the above, the information related to the threshold has time items as a one-dimensional index. The information related to the threshold can also have a two-dimensional or three-dimensional index that also includes variables related to the direction of arrival.

[0343] Figure 15 is a diagram showing the relationship between the direction of the direct tone, the direction of the reflected tone, the time difference, and the threshold. For example, as shown in Figure 15, the threshold calculated in advance according to the relationship between the direction of the direct tone (θ), the direction of the reflected tone (γ), the time difference (T), and the volume ratio (L) can also be stored.

[0344] The direction of the direct tone (θ) corresponds to the angle relative to the direction in which the direct tone arrives at the listener. The direction of the reflected tone (γ) corresponds to the angle relative to the direction in which the reflected tone arrives at the listener. Here, the direction the listener is facing is set to 0 degrees. The time difference (T) corresponds to the difference between the arrival time of the direct tone and the arrival time of the reflected tone towards the listening position. The volume ratio (L) corresponds to the volume ratio of the volume of the direct tone arrival to the volume of the reflected tone arrival.

[0345] Of course, the threshold shown in Figure 15 is an example, and the threshold is not limited to the example in Figure 15. Furthermore, Figure 15 mainly illustrates the threshold when the angle (θ) of the direction of arrival of the direct tone is 0 degrees. However, the threshold for cases where the direction of arrival of the direct tone (θ) is other than 0 degrees is also stored in memory 1404.

[0346] Furthermore, in the above, the threshold is stored as an arrangement where the direction of the direct tone (θ) (more specifically, the angle (θ) of the direction of arrival of the direct tone) and the direction of the reflected tone (γ) (more specifically, the angle (γ) of the direction of arrival of the reflected tone) are treated as independent variables or indices. However, the angle (θ) of the direction of arrival of the direct tone and the angle (γ) of the direction of arrival of the reflected tone may also not be used as independent variables.

[0347] For example, the angle difference between the direction of arrival of the direct tone (θ) and the direction of arrival of the reflected tone (γ) can also be used. This angle difference corresponds to the angle formed by the direction of arrival of the direct tone and the direction of arrival of the reflected tone, and can also be expressed as the arrival angles of the direct tone and the reflected tone.

[0348] Figure 16 is a graph showing the relationship between angular difference, time difference, and threshold. For example, as shown in the example in Figure 16, the threshold can also be pre-calculated by using the angular difference (Φ) between the angle of arrival of the direct sound (θ) and the angle of arrival of the reflected sound (γ) as a variable. Of course, the threshold shown in Figure 16 is an example, and the threshold is not limited to the example in Figure 16.

[0349] In the example of Figure 16, the number of variables used in deriving the threshold can be reduced. Therefore, the number of thresholds stored in memory 1404 can be reduced. Consequently, the amount of data stored in memory 1404 can be reduced.

[0350] Furthermore, when using the angle difference (Φ) between the angle (θ) of the direction of arrival of the direct tone and the angle (γ) of the direction of arrival of the reflected tone, the threshold data can also be stored in a two-dimensional arrangement. Additionally, in the selection process, a three-dimensional arrangement can be used to calculate the difference between the angle (θ) of the direction of arrival of the direct tone and the angle (γ) of the direction of arrival of the reflected tone.

[0351] The method of selecting reflected sounds using a threshold corresponding to the direction of arrival will be described later.

[0352] [First Variation of the Threshold Setting Method] In the examples of Figures 12A, 12B, and 12C, multiple forms and types of thresholds can also be stored in the spatial information management units 1201 and 1211. Furthermore, it can be determined which form and type of threshold among the multiple forms and types will be used for the selection processing of the reflected sound. Specifically, as shown in Example 3 of Figure 12C, the highest threshold can be used at the time difference (T) corresponding to the arrival time of the reflected sound.

[0353] Alternatively, as shown in Example 4, a masking threshold, an echo detection limit threshold, and a threshold representing the minimum volume reproduced in the virtual space can be stored. Furthermore, the highest threshold can be used for the time difference (T) corresponding to the arrival time of the reflected sound.

[0354] [Second variation of the threshold setting method] As another example of the threshold setting method, a method for setting the threshold based on the properties of direct sounds will be described.

[0355] Figure 17 is a block diagram showing another configuration example of the rendering unit 1300 shown in Figure 7. The rendering unit 1300 in Figure 17 differs from the rendering unit 1300 in Figure 7 in that it includes a threshold adjustment unit 1304. The description of the part other than the threshold adjustment unit 1304 is the same as that described in Figure 7, so it is omitted.

[0356] The threshold adjustment unit 1304 selects a threshold to be used by the selection unit 1302 from the threshold data based on information representing the nature of the sound signal. Alternatively, the threshold adjustment unit 1304 may also adjust the threshold included in the threshold data based on information representing the nature of the sound signal.

[0357] Information indicating the nature of the sound signal can also be included in the input signal. Furthermore, the threshold adjustment unit 1304 can also obtain information indicating the nature of the sound signal from the input signal. Alternatively, the analysis unit 1301 can analyze the sound signal contained in the received input signal, derive the nature of the sound signal, and output information indicating the nature of the sound signal to the threshold adjustment unit 1304.

[0358] Information representing the nature of a sound signal can be obtained either before rendering begins or at any time during rendering.

[0359] Furthermore, the threshold adjustment unit 1304 may not be included in the sound signal processing device 1001, or it may function as a threshold adjustment unit 1304 in another communication device. In this case, the parsing unit 1301 or the selection unit 1302 may also obtain information representing the nature of the sound signal, threshold data corresponding to the nature, or information for adjusting the threshold data according to the nature from other communication devices via the communication IF 1403.

[0360] Figure 18 is a flowchart illustrating another example of the selection process. Figure 19 is a flowchart illustrating yet another example of the selection process. In Figures 18 and 19, a threshold is set based on the properties of the direct sound. Specifically, in Figure 18, the threshold adjustment unit 1304 determines the threshold from the threshold data based on the time difference (T) and the properties of the sound signal. In Figure 19, the threshold adjustment unit 1304 adjusts the threshold determined from the threshold data based on the time difference (T) based on the properties of the sound signal.

[0361] The actions for each example are explained below. Additionally, explanations of processes common to the examples in Figure 14 are omitted.

[0362] First, an example of the process shown in Figure 18 will be explained. Here, threshold data is pre-stored in memory 1404 according to each property of the direct tone. Thus, multiple threshold data corresponding to multiple properties are pre-stored in memory 1404. Furthermore, the threshold adjustment unit 1304 determines the threshold data to be used in the selection process of the reflected tone from the multiple threshold data.

[0363] For example, the threshold adjustment unit 1304 obtains the properties of the direct tone based on the input signal (S211). The threshold adjustment unit 1304 may also obtain the properties of the direct tone that are associated with the input signal. Next, the threshold adjustment unit 1304 determines a threshold corresponding to the time difference (T) and the properties of the direct tone (S212).

[0364] Furthermore, as shown in FIG19, the threshold adjustment unit 1304 may also adjust the threshold determined by the selection unit 1302 based on the properties of the direct tone (S221).

[0365] In any case, the input signal may include information representing the nature of the sound signal, information for adjusting the threshold according to the nature of the sound signal, or both. The threshold adjustment unit 1304 may also use one or both of these to adjust the threshold.

[0366] Furthermore, information indicating the nature of the sound signal, information used to adjust the threshold, or both, can be transmitted via an input signal different from the input signal containing the sound signal. In this case, information relating to an input signal different from the input signal can also be included in the input signal containing the sound signal, or the information relating to the input signal different from the input signal can be stored in memory 1404 along with information about the threshold.

[0367] In the examples of Figures 18 and 19, the threshold used in selecting the reflected tone is set according to the properties of the direct tone, i.e., the properties of the sound signal. Either pre-set threshold data for each property can be used, as in Figure 18, or the threshold can be adjusted according to the properties of the sound signal, as in Figure 19. Furthermore, the parameters of the threshold data can also be adjusted according to the properties of the sound signal.

[0368] Furthermore, the operation performed by the threshold adjustment unit 1304 can also be performed by the analysis unit 1301 or the selection unit 1302. For example, the analysis unit 1301 may obtain the properties of the sound signal. Alternatively, the selection unit 1302 may set the threshold based on the properties of the sound signal.

[0369] Next, the relationship between the properties of the sound signal and the threshold will be explained.

[0370] Two short sounds arriving at a listener's ear consecutively, if the time interval between them is sufficiently short, are perceived as a single sound. This phenomenon is called the priority effect. It is known that the priority effect occurs only for discontinuous, i.e., transient sounds (Non-Patent Document 1). Therefore, when the sound signal represents a stationary tone, the echo detection limit can be set lower compared to when the sound signal represents a non-stationary tone.

[0371] That is, based on the characteristics of this priority effect, for example, when the direct tone is a stable sound, the threshold is set to be relatively small. Alternatively, the higher the stability, the smaller the threshold can be set.

[0372] An example of processing when the sound signal is stationary will be explained. First, the threshold adjustment unit 1304 or the analysis unit 1301 determines the stationarity based on the amount of change in the frequency components of the sound signal over time. For example, if the amount of change is small, the stationarity is determined to be high. Conversely, if the amount of change is large, the stationarity is determined to be low. The result of the determination can be used to set a flag representing the level of stationarity, or a parameter representing stationarity can be set based on the amount of change.

[0373] Next, the threshold adjustment unit 1304 may adjust the threshold data or threshold based on information indicating the stability of the sound signal, such as a flag or parameter, and set the adjusted threshold data or threshold as the threshold data or threshold used in the selection unit 1302.

[0374] Alternatively, parameters for setting threshold data based on information representing the stability of the direct tone can be pre-stored in the memory 1404. In this case, the threshold adjustment unit 1304 can also determine the stability of the sound signal and set the threshold data used in selecting the reflected tone based on the information and parameters representing the stability.

[0375] Alternatively, multiple parameters of the threshold data may be pre-stored in the memory 1404 corresponding to multiple patterns of the direct tone's stability. In this case, the threshold adjustment unit 1304 may also determine the stability of the sound signal, select parameters of the threshold data based on the pattern of the direct tone's stability, and set the threshold data used in the selection of reflected tones based on the parameters of the threshold data.

[0376] In addition, the stability of a sound signal can be determined based on the amount of change in the frequency components of the sound signal each time a sound signal is input.

[0377] Alternatively, the stationarity of the sound signal can be determined based on information representing stationarity that has been pre-associated with the sound signal. That is, information representing the stationarity of the sound signal can be pre-associated with the sound signal and stored in the memory 1404. The parsing unit 1301 can also obtain the information representing stationarity associated with the sound signal each time an input sound signal is received. Furthermore, the threshold adjustment unit 1304 can adjust the threshold based on the information representing stationarity associated with the sound signal.

[0378] As another example of setting a threshold based on the properties of the sound signal, the application range of the echo detection limit can be set shorter when the sound signal represents a shorter sound (such as a click) compared to when the sound signal represents a longer sound. This processing is based on the characteristics of the priority effect.

[0379] It is known that, through the priority effect, two short sounds arriving consecutively at a listener's ear are perceived as a single sound if the time interval between them is sufficiently short. The upper limit of this time interval depends on the length of the sound. For example, the upper limit of this time interval is approximately 5 ms for a click sound, and sometimes as high as 40 ms for complex sounds such as human voices or music (Non-Patent Document 1).

[0380] Based on this priority effect, for example, in the case of sounds with shorter direct tone durations, a shorter duration threshold is set. Furthermore, the shorter the direct tone duration, the shorter the duration threshold is set.

[0381] Setting a shorter time threshold means setting a threshold corresponding to the echo detection limit based on the priority effect characteristics within a range where the time difference (T) between the direct tone and the reflected tone is small. Outside this range, no threshold corresponding to the echo detection limit based on the priority effect characteristics is set. That is, outside this range, the threshold is small. Therefore, setting a shorter time threshold for shorter sounds corresponds to setting a smaller threshold for shorter sounds.

[0382] As another example of setting a threshold based on the nature of the direct sound, the threshold can be set lower when the direct sound is an intermittent sound (speech, etc.) compared to when the direct sound is a continuous sound (music, etc.).

[0383] For example, when the direct sound corresponds to speech, there are repeated vocal and non-vocal parts, and as a masking effect, only the aftermasking effect occurs in the non-vocal part. On the other hand, when the direct sound is a continuous sound like musical content, both the aftermasking effect and the simultaneous masking effect based on the sound produced at that time occur. Therefore, the comprehensive masking effect is higher in the case of music, etc., than in the case of speech, etc.

[0384] Based on the masking effect characteristics described above, the threshold can be set higher in the case of music, etc., compared to the case of speech, etc. Conversely, the threshold can be set lower in the case of speech, etc., compared to the case of music, etc. That is, the threshold can be set lower when there are many interruptions in the direct sound.

[0385] As mentioned above, information indicating the properties of a direct tone can also include information about its smoothness, discontinuity, and duration. Furthermore, information indicating the properties of a direct tone can be any combination of these characteristics. Additionally, information indicating the properties of a direct tone can be information about the temporal variation of any one of these characteristics, or information about the temporal variation of any combination thereof. In other words, information indicating the properties of a direct tone can also be information about the temporal variation of the direct tone.

[0386] For example, as shown in the explanation of stationarity determination, information representing the properties of a direct tone can also be time-series data of frequency characteristics. Here, frequency characteristics can also be represented in conventional forms such as gain values ​​for each frequency band, Fourier series of the signal over the time axis, or LPC coefficients or cepstral coefficients used to calculate the frequency envelope.

[0387] Furthermore, information representing the properties of a direct tone can also be presented as information representing the discontinuity of the direct tone, listing multiple sets of information (the approximate shape of the amplitude envelope) of the signal's amplitude stability over a time series. Here, the amplitude value can also be expressed as a ratio relative to a reference volume.

[0388] Furthermore, the information representing the properties of a direct tone can also be information related to the frequency characteristics of the direct tone. For example, the information representing the properties of a direct tone can also be information representing the stability of the frequency characteristics of the direct tone. Specifically, the information representing the properties of a direct tone can also be information that lists multiple groups of frequency characteristics of the signal during a period of small frequency variation in a time series (approximate shape of a spectrum). Here, the volume used as a reference for the aforementioned frequency characteristics can also be the aforementioned reference volume.

[0389] For example, information representing the temporal variation of a direct tone is information representing the envelope of the direct tone. Information representing the temporal variation of a direct tone can also be used when the "minimum audible limit" is a threshold value as shown in [Example 4] of Figure 12C. The signal compared to the minimum audible limit is the volume of the reflected tone.

[0390] The volume of the reflected sound is obtained through geometric calculations based on the positions of the sound source, the listener, and the reflecting object. Specifically, a reference volume of the reflected sound is obtained relative to a reference volume of the sound source. By using information about changes in the volume of the sound source as information representing the properties of the direct sound, the reference volume of the reflected sound is increased or decreased, allowing for an accurate determination of the volume of the reflected sound at any given moment. This is because changes in the volume of the sound source are reflected in changes in the volume of the reflected sound.

[0391] After adjusting the volume of the reflected sound, by comparing the volume of the reflected sound with the threshold, it is possible to more accurately and appropriately select the reflected sound that is audibly desired.

[0392] Of course, the same result can be obtained by adjusting the threshold based on the reciprocal of the change in the volume of the sound source, without adjusting the reference volume of the reflected sound, and then comparing the adjusted threshold with the reference volume of the reflected sound. That is, the reference volume of the reflected sound can be adjusted using information about the change in the volume of the sound source, and the threshold can also be adjusted using the same information. The adjustment of the reference volume of the reflected sound and the adjustment of the threshold are mutually corresponding.

[0393] Depending on the surface composition of the object reflecting the sound, the reflectivity of the sound (and the attenuation rate of the reflected sound) varies in each frequency band. Therefore, as described later, the reflectivity (attenuation rate) of the sound can also be associated with each frequency band for the object reflecting the sound. Based on this reflectivity information and the information from the spectrogram, it is possible to more accurately determine whether to select the reflected sound. For example, the following process can be performed.

[0394] Specifically, for example, information from the spectrogram indicates that frequency components in the high-frequency band are more dominant than those in the low-frequency band within a certain time interval. Additionally, for example, information from the reflectivity of sound indicates that the reflectivity of frequency components in the high-frequency band is extremely low compared to that in the low-frequency band.

[0395] In this case, even if the signal amplitude of the sound source is large on the time axis, the volume of the reflected sound is reduced by multiplying the frequency components represented by the information in the spectrogram with the attenuation rate of each frequency band represented by the information in the reflectivity, and the reflected sound may not be selected.

[0396] As mentioned above, information representing the properties of a direct tone can also be information representing the temporal variation of the direct tone. For example, information representing the properties of a direct tone can also represent values ​​obtained by analyzing the direct tone over a predetermined time period.

[0397] Specifically, information representing the properties of a direct tone can also be obtained by calculating the average energy or average amplitude of the direct tone for each predetermined time length. Alternatively, information representing the properties of a direct tone can also be obtained by calculating the energy or average amplitude of the direct tone for each short time analysis length and then taking a weighted average of the energy or average amplitude for each long time analysis length that is longer than the short time analysis length.

[0398] More specifically, for example, information representing the temporal variation of the direct tone can also be obtained by calculating the energy or average amplitude of the direct tone for each pre-defined short time length (e.g., 5 ms, hereinafter referred to as an analysis frame). Alternatively, information representing the temporal variation of the direct tone can also be represented by a weighted average of the energy or average amplitude calculated over the past N-1 analysis frames.

[0399] Assuming the energy of the nth analysis frame is represented by E(n), the information I(n) representing the properties of the direct tone is obtained according to the following formula.

[0400] [Mathematical Expression 2] Here, the parameter a(i) represents the weighting coefficient. Typically, a(i) is set such that a(i) ≥ 0 and the sum of a(i) is 1. However, the method of setting a(i) is not limited to this.

[0401] Furthermore, information I(n) representing the properties of the direct tone is calculated every 5ms since the direct tone was received. That is, the temporal variation of information I(n) representing the properties of the direct tone can be calculated with low latency. Therefore, this method is suitable for applications requiring real-time performance.

[0402] Alternatively, the information I(n) representing the properties of direct sounds can be obtained using the following formula.

[0403] [Mathematical Expression 3] Here, the parameter b(i) represents the weighting coefficient. Typically, b(i) is set such that b(i) ≥ 0 and the sum of b(i) is 1. However, the method of setting b(i) is not limited to this.

[0404] In this formula, information I(n) representing the properties of a direct tone is recursively obtained. Therefore, the average energy over a long time period can be calculated with minimal computation.

[0405] Equations 1 and 2 above can be considered as filters where E(n) is the input signal and I(n) is the output signal. In this case, Equation 1 is a filter for the moving average (MA) model, and Equation 2 is a filter for the autoregressive (AR) model, both exhibiting the characteristics of low-pass filters. Alternatively, a filter for the ARMA model, which combines the two, can also be used.

[0406] Furthermore, the method for deriving information representing the temporal variation of the direct tone is not limited to the aforementioned calculation formula or filter; other known methods can also be used. As mentioned above, the information representing the temporal variation of the direct tone represents a value obtained by analyzing the direct tone over a predetermined time length. The direct tone can also be analyzed from a perspective other than average energy.

[0407] Furthermore, as mentioned above, the information representing the properties of a direct tone can also be information related to the frequency characteristics of the direct tone. This information related to the frequency characteristics of the direct tone can also be information calculated using those characteristics. For example, information related to the frequency characteristics of the direct tone can be obtained by averaging the low-frequency components of the direct tone over a predetermined analysis length to obtain the average energy of the low-frequency components.

[0408] Specifically, the low-frequency components of the direct tone are determined by applying a low-pass filter to the direct tone contained in the analysis frame length. Based on the energy or average amplitude of this low-frequency component, information representing the properties of the direct tone is derived in the same manner as in Equation 1 above.

[0409] Assume the energy of the low-frequency components in the nth analysis frame is determined by E L In the case of (n), the information I(n) representing the nature of the direct sound is obtained according to the following formula.

[0410] [Mathematical Expression 4] Here, the parameter c(i) represents the weighting coefficient. Typically, c(i) is set such that c(i) ≥ 0 and the sum of c(i) is 1. However, the method of setting c(i) is not limited to this.

[0411] Furthermore, information I(n) representing the properties of the direct tone is calculated every 5ms since the direct tone was received. That is, the temporal variation of information I(n) representing the properties of the direct tone can be calculated with low latency. Therefore, this method is suitable for applications requiring real-time performance.

[0412] In addition, similar to Equation 2, the information I(n) representing the properties of direct sounds can also be obtained according to the following formula.

[0413] [Mathematical Expression 5] Here, the parameter d(i) represents the weight coefficient. Typically, d(i) is set such that d(i) ≥ 0 and the sum of d(i) is 1. However, the method of setting d(i) is not limited to this.

[0414] In this formula, information I(n) representing the properties of a direct tone is recursively obtained. Therefore, the average energy over a long time period can be calculated with minimal computation.

[0415] Equations 3 and 4 above can be considered as filters with E(n) as the input signal and I(n) as the output signal. In this case, Equation 3 is a filter for the moving average (MA) model, and Equation 4 is a filter for the autoregressive (AR) model, both of which have the characteristics of low-pass filters. Alternatively, a filter for the ARMA model, which combines the two, can also be used.

[0416] In the above method for determining the low-frequency components of a direct tone, a filter with low-pass characteristics is used; however, the method for determining the low-frequency components of a direct tone is not limited to this. Furthermore, the method for deriving information representing the time variation of the direct tone is not limited to the aforementioned calculation formula or filter; other known methods can also be used. For example, the spectrum of a direct tone can be calculated by performing a frequency transformation on the direct tone. Furthermore, the energy or average amplitude of the low-frequency components of the spectrum can be calculated.

[0417] Furthermore, in the above, MA or AR models are used to derive information representing the temporal variation of direct tones. The coefficients of these models can be pre-set fixed values ​​or time-varying values.

[0418] In addition, the relationship between the analysis frame length and the information update thread generation interval can also be as follows.

[0419] For example, when the analysis frame duration is TA (msec) and the information update thread generation interval is TU (msec), the value of N in Equations (1) and (3) above in the MA filter can also be approximately equal to the value given by TU / TA. Additionally, b(i) and d(i) (1≤i<N) in Equations (2) and (4) above in the AR filter can also be approximately equal to the filter's time constant, which is approximately equal to TU (msec).

[0420] The reason for this setting is that the filter is expected to converge during the information update interval.

[0421] On the other hand, in the above-described configuration, if the value of the information representing the temporal variation of the direct tone changes too drastically, I(n) can be pre-calculated. Furthermore, the pre-calculated I(n) can be applied to the selection processing of reflected tones. For example, in the processing of the frame at time t, I(t+tau) can be used. Here, tau is a value determined based on the convergence characteristics of the filter. In cases of slow convergence, the value of tau is larger compared to cases of fast convergence.

[0422] Furthermore, auditory masking (frequency masking) information calculated based on the direct tone can also be used as information representing the characteristics of the direct tone. The auditory masking information represents a threshold value for the amplitude in the frequency domain that is masked by the direct tone. It is also possible to compare the amplitude value of reflected tones in the same frequency domain with the threshold value, without selecting reflected tones with amplitude values ​​smaller than the threshold. The amplitude value of reflected tones in the frequency domain can also be obtained by the analysis unit 1301 as information representing the characteristics of the reflected tones.

[0423] By setting the threshold used in selecting reflected sounds based on the properties of the direct sound, the reflected sounds that are audibly required can be appropriately selected, and auditory characteristics can be effectively reflected in the stereo sound reproduction system 1000. The processing of detecting the properties of the direct sound, determining the threshold based on the properties, and adjusting the threshold based on the properties can be performed either during the rendering process or before the rendering process begins.

[0424] For example, these processes can occur during virtual space creation (when the software is created), at the start of virtual space processing (when the software starts or rendering begins), or at timed intervals in information update threads that occur periodically during virtual space processing. Furthermore, virtual space creation can be timed to build the virtual space before the start of sound processing, or it can occur when virtual space information (spatial information) is acquired, or it can occur when the software acquires the information.

[0425] Here, in the information update thread, processing is performed to update the spatial information managed by the Spatial Information Management Departments 1201 and 1211.

[0426] The information update thread is responsible for tasks such as updating the position and orientation of the listener's avatar configured in the virtual space based on the position and orientation of the VR goggles worn by the listener, or updating the position of moving objects in the virtual space. This processing is provided within a processing thread that starts at a relatively low frequency of around tens of Hz.

[0427] In such low-frequency processing threads, the processing of information representing the properties of direct tones can still be performed. This is because the frequency of changes in the properties of direct tones is lower than the frequency of audio processing frames used for audio output. Therefore, the computational load of this processing can be relatively reduced. Furthermore, updating information at an unnecessarily fast frequency carries the risk of generating impulse noise. This risk can be avoided by updating information at a low frequency.

[0428] [Third Variation of the Threshold Setting Method] As another example of the threshold setting method, the threshold can also be set based on the computing resources (CPU capacity, memory resources, PC performance, or remaining battery power, etc.) used to process the reproduction of the virtual space. More specifically, the sensor 1405 of the sound signal processing device 1001 detects the amount of computing resources and sets a higher threshold when the amount of computing resources is low. As a result, the volume of more reflected sounds is lower than the threshold, thus reducing the reflected sounds processed by both ears and reducing the amount of computing power.

[0429] Alternatively, in cases where signal processing is performed by battery-powered devices such as smartphones or VR headsets, it is desirable to prioritize long processing times and conserve computing resources. In such cases, the threshold can be set high without checking the amount or remaining amount of computing resources.

[0430] [Fourth variation of the threshold setting method] As another example of the threshold setting method, the threshold can also be set by the administrator or listener of the virtual space by having a threshold setting unit (not shown) in the sound signal processing device 1001 or the sound prompting device 1002.

[0431] For example, the listener wearing the sound prompt device 1002 could choose between an "energy-saving mode" (fewer reflected sounds and less computation) and a "high-performance mode" (more reflected sounds and more computation). Alternatively, the administrator of the stereo sound reproduction system 1000 or the producer of the stereo sound content could choose the mode. Furthermore, instead of a mode, a threshold or threshold data could be directly selected.

[0432] [First Variation of the Operation of the Rendering Unit] Figure 20 is a flowchart showing a first variation of the operation of the sound signal processing apparatus 1001. Figure 20 shows the processing mainly performed by the rendering unit 1300 of the sound signal processing apparatus 1001. In this variation, volume compensation processing is added to the operation of the rendering unit 1300.

[0433] For example, the analysis unit 1301 acquires data (input signal) (S301). Next, the analysis unit 1301 analyzes the data (S302). Next, the selection unit 1302 determines whether to select reflected sounds based on the analysis results (S303). Next, the reproduction unit 1303 performs volume compensation processing based on the unselected reflected sounds (S304). Next, the reproduction unit 1303 performs audio processing on both direct and reflected sounds (S305). Finally, the reproduction unit 1303 outputs both direct and reflected sounds as audio (S306).

[0434] In the processes described above (S301 to S306), the processes other than the volume compensation process (S304) are common to the other examples described above, so their descriptions are omitted.

[0435] Volume compensation processing is performed for reflected sounds that were not selected in the selection process. For example, by not selecting reflected sounds in the selection process, a lack of volume perception occurs. Volume compensation processing can suppress the unpleasantness that accompanies this lack of volume perception. Two methods are disclosed as examples of methods for compensating for volume perception. Either method can be used.

[0436] First, the method of compensating for the sense of volume by increasing the volume of the direct tone will be explained. The reproduction unit 1303 generates the direct tone by increasing its volume by an amount corresponding to the volume of the unselected reflected tone. As a result, the sense of volume lost due to the lack of reflected tone generation is compensated.

[0437] When the reproduction unit 1303 increases the volume, it can also increase the volume according to the frequency characteristics of the reflected sound, one frequency component at a time. To enable this process, a predetermined attenuation rate for the volume of the reflected sound can be assigned to each frequency band. Thus, the frequency characteristics of the reflected sound can be derived.

[0438] Next, a method for compensating for the sense of volume by synthesizing reflected sounds into direct sounds will be explained. In this method, the reproduction unit 1303 adds unselected reflected sounds to the direct sounds to generate a direct sound, thereby compensating for the sense of volume caused by the lack of generated reflected sounds. The generated direct sound reflects the volume (amplitude), frequency, and delay of the unselected reflected sounds.

[0439] In the case of increasing the volume of the direct tone, the computational workload of the compensation process is very small, but only the volume is compensated. In the case of synthesizing the reflected tone into the direct tone, the computational workload of the compensation process is larger compared to the method of increasing the volume of the direct tone, but the characteristics of the reflected tone are compensated more accurately.

[0440] In both cases, no reflected tones are generated, only direct tones, thus reducing the overall computational load. In particular, the computational load required for binaural processing, which includes convolutional HRTF processing, is reduced, resulting in a significant reduction in overall computational load. This is because the computational load required for binaural processing is far greater than that required for the aforementioned compensation processing.

[0441] In addition, if the reason for not selecting the reflected sound is that the volume of the reflected sound is lower than the masking threshold, since the sense of volume will not be lost, the reflected sound can be removed without compensation.

[0442] [Second Variation of the Operation of the Rendering Unit] Figure 21 is a flowchart illustrating a second variation of the operation of the sound signal processing apparatus 1001. Figure 21 primarily illustrates the processing performed by the rendering unit 1300 of the sound signal processing apparatus 1001. In this variation, left and right volume difference adjustment processing is added to the operation of the rendering unit 1300.

[0443] For example, the analysis unit 1301 analyzes the input signal (S401). Next, the analysis unit 1301 detects the direction of sound arrival (S402). Next, the selection unit 1302 adjusts the volume difference of the sound perceived by the left and right ears (S403). Furthermore, the selection unit 1302 adjusts the time difference (delay) of the sound arrival perceived by the left and right ears (S404). Based on the adjusted sound information, the selection unit 1302 determines whether to select the reflected sound (S405).

[0444] In the above-described processes (S401 to S405), the processes other than the left and right volume difference adjustment process (S403) and the delay adjustment process (S404) are common to the other examples described above, so their descriptions are omitted.

[0445] Figure 22 is a diagram showing an example of the configuration of the avatar, the sound source object, and the obstacle object. For example, when the listener's facing direction is 0 degrees, as shown in Figure 22, if the polarity (e.g., positive or negative) of the direction of arrival of the direct sound (θ) and the direction of arrival of the reflected sound (γ) (the direction of the reflected sound (γ)) is different, the volume difference generated between the two ears is corrected.

[0446] Specifically, when the polarities of θ and γ are different, the ear that primarily (first) perceives the sound in the direct tone and the reflected tone is different. In this case, the selection unit 1302 performs a left-right volume difference adjustment process (S403) to adjust the volume of the direct tone according to the position of the ear that primarily perceives the reflected tone. For example, the selection unit 1302 attenuates the volume of the direct tone when it reaches the listener by multiplying the volume by (1.0 - 0.3sin(θ)) (0 ≤ θ ≤ 180°).

[0447] The selection unit 1302 determines whether to select the reflected sound by calculating the volume ratio of the corrected direct sound volume to the reflected sound volume as described above, and comparing the calculated volume ratio with a threshold. This corrects the volume difference between the two ears, more accurately derives the volume of the direct sound that affects the reflected sound, and more accurately determines whether to select the reflected sound.

[0448] In addition to adjusting the left and right volume difference (S403), the selection unit 1302 can also perform a delay adjustment process (S404) to match the position of the ear that perceives the reflected sound, thus delaying the arrival time of the direct sound. Specifically, the selection unit 1302 can also delay the arrival time of the direct sound by adding (a(sinθ+θ) / c) ms (where a is the radius of the head and c is the speed of sound) to the arrival time of the direct sound.

[0449] [Third variation of the rendering unit's actions] The method for setting a threshold corresponding to the direction of arrival will be explained.

[0450] Figure 23 is a flowchart illustrating yet another example of the selection process. Descriptions of the processes common to the example in Figure 14 are omitted. In the example of Figure 23, the selection unit 1302 uses a threshold corresponding to the direction of arrival to select the reflected sound.

[0451] Specifically, the selection unit 1302 calculates the direct sound arrival path (pd), the reflected sound arrival path (pr), and the orientation information D1 of the avatar, using the orientation of the avatar as a reference, the direct sound arrival direction (θ) and the reflected sound arrival direction (γ) (the direction of the reflected sound (γ)). That is, the selection unit 1302 detects the direct sound arrival direction (θ) and the reflected sound arrival direction (γ) (S231). The orientation of the avatar corresponds to the orientation of the listener. The orientation information D1 of the avatar may also be included in the input signal.

[0452] The selection unit 1302 uses three indices, including the direct sound arrival direction (θ) and the reflected sound arrival direction (γ), as well as the time difference (T), to determine the threshold used in the selection process (S232) based on the three-dimensional arrangement shown in FIG15.

[0453] As an example, this section explains how to set the threshold used in the selection process when an avatar, a sound source object, and an obstacle object are configured as shown in Figure 22.

[0454] The position information of the avatar, the sound source object, and the obstacle object, as well as the orientation information D1 of the avatar, are obtained from the input signal. Using this position information and orientation information D1, the direction of the direct sound (θ) and the direction of the sound image of the reflected sound (γ) are calculated when the orientation of the avatar is set to 0 degrees. In the case of Figure 22, the direction of the direct sound (θ) is about 20 degrees, and the direction of the sound image of the reflected sound (γ) is about 265 degrees (-95 degrees).

[0455] Next, referring to the threshold data stored in a three-dimensional arrangement as shown in FIG15, the threshold is determined from the arrangement region corresponding to the values ​​of the two directions (θ) and (γ) and the value of the time difference (T) calculated by the analysis unit 1301. In the case that there is no index corresponding to the calculated values ​​of (θ), (γ), and (T), the threshold corresponding to the closest index can also be determined.

[0456] Alternatively, the threshold can be determined by interpolation, extrapolation, or other processing based on one or more thresholds corresponding to one or more indices close to the calculated values ​​of (θ), (γ), and (T). For example, the threshold corresponding to (20 degrees, 265 degrees, T) can be determined based on four thresholds corresponding to the four indices (0 degrees, 225 degrees, T), (0 degrees, 270 degrees, T), (45 degrees, 225 degrees, T), and (45 degrees, 270 degrees, T).

[0457] The selection process based on the difference between the angle (θ) of the direction of arrival of the direct sound and the angle (γ) of the direction of arrival of the reflected sound is explained.

[0458] For example, threshold data can be pre-created and set, as shown in Figure 16, with the angle difference (Φ) and time difference (T) between the direction of arrival of the direct tone (θ) and the direction of arrival of the reflected tone (γ) arranged as a two-dimensional index. In this case, the angle difference (Φ) and time difference (T) are referenced in the selection process. Alternatively, the angle difference (Φ) between the angle of arrival of the direct tone (θ) and the angle of arrival of the reflected tone (γ) can be calculated in the selection process, and the calculated angle difference (Φ) can be used to determine the threshold.

[0459] Alternatively, a threshold data can be set to be used as an index for arranging the combination of the angle difference (Φ), the direction of arrival of the direct sound (θ), and the time difference (T), or the combination of the angle difference (Φ), the direction of arrival of the reflected sound (γ), and the time difference (T).

[0460] Alternatively, the threshold data can be set by arranging the values ​​of (θ), (γ), and (T) as a three-dimensional index, as shown in Figure 15.

[0461] [Fourth variation of the operation of the rendering unit] The processing performed by the analysis unit 1301, selection unit 1302 and reproduction unit 1303 described above can also be performed as pipeline processing as described in Patent Document 3.

[0462] Figure 24 is a block diagram showing a configuration example for pipeline processing in the rendering unit 1300.

[0463] The rendering unit 1300 in FIG24 includes a reverberation processing unit 1311, an initial reflection processing unit 1312, a distance attenuation processing unit 1313, a selection unit 1314, a generation unit 1315, and a binaural processing unit 1316. These multiple components may be constituted by the multiple components of the rendering unit 1300 shown in FIG7, or by at least a portion of the multiple components of the sound signal processing apparatus 1001 shown in FIG5.

[0464] Pipeline processing refers to dividing the processing used to impart sound effects into multiple processes and executing these processes sequentially. Each process may perform signal processing on the audio signal or generate parameters used in signal processing.

[0465] The rendering unit 1300 can also perform reverberation processing, initial reflection processing, distance attenuation processing, and binaural processing as a pipeline process. However, these are just examples; pipeline processing may include other processing methods, or it may exclude some of them. For example, pipeline processing may also include diffraction processing and occlusion processing. Furthermore, reverberation processing, for example, can be omitted if it is not needed.

[0466] Furthermore, each process can be represented as a stage. Additionally, the results of each process, and the generated sound signals such as reflected sounds, can be represented as rendering items. The multiple stages in pipeline processing and their order are not limited to the example shown in Figure 24.

[0467] Here, the parameters used in the selection process (arrival path, arrival time, and volume ratio related to direct and reflected tones) can also be calculated in one of the multiple stages used to generate the render item. That is, the parameters used in the selection of reflected tones are calculated in a part of the pipeline processing used to generate the render item. Alternatively, not all stages may be performed by the rendering unit 1300. For example, some stages may be omitted, or they may be performed outside of the rendering unit 1300.

[0468] The reverberation processing, initial reflection processing, distance attenuation processing, selection processing, generation processing, and binaural processing that may be included as stages in pipeline processing are described. Within each stage, metadata contained in the input signal can also be parsed to calculate the parameters used in generating the reflected tones.

[0469] In reverberation processing, the reverberation processing unit 1311 generates parameters used in the generation of a sound signal representing reverberant sound. Reverberant sound refers to the sound that arrives at the listener as reverberation after the direct tone. As an example, reverberant sound is the sound that arrives at the listener after a relatively late stage (e.g., from the arrival of the direct tone to about one hundred and several tens of ms) following the arrival of the initial reflected sound, as described later. It undergoes more reflections (e.g., dozens of times) than the initial reflected sound.

[0470] The reverberation processing unit 1311 refers to the sound signal and spatial information contained in the input signal and calculates the reverberation sound using a pre-prepared function that is used to generate the reverberation sound.

[0471] The reverberation processing unit 1311 can also apply known reverberation generation methods to the sound signal contained in the input signal to generate reverberation. An example of a known reverberation generation method is the Schroeder method, but known reverberation generation methods are not limited to the Schroeder method. Furthermore, the reverberation processing unit 1311 uses the shape and acoustic characteristics of the sound reproduction space represented by spatial information in the application of known reverberation generation methods. Therefore, the reverberation processing unit 1311 can calculate the parameters used to generate the reverberation.

[0472] In the initial reflection processing, the initial reflection processing unit 1312 calculates parameters for generating the initial reflection tone based on spatial information. The initial reflection tone is the reflection tone that reaches the listener after more than one reflection in a relatively early stage (e.g., about tens of milliseconds from the arrival of the direct tone) after the direct tone reaches the listener from the sound source object.

[0473] The initial reflection processing unit 1312 calculates, for example, the path of the reflected sound from the sound source object to the listener via the reflection object, by referring to the sound signal and metadata. For example, the shape of the three-dimensional sound field (space), the size of the three-dimensional sound field, the position of the reflecting object such as the structure, and the reflectivity of the reflecting object can also be used in the path calculation.

[0474] Furthermore, the initial reflection processing unit 1312 may also calculate the path of the direct tone. This path information may also be used as a parameter for the initial reflection processing unit 1312 to generate the initial reflected tone, or as a parameter for the selection unit 1314 to select the reflected tone.

[0475] In distance attenuation processing, the distance attenuation processing unit 1313 calculates the volume of the direct tone and the reflected tone reaching the listener based on the path lengths of the direct tone and the reflected tone. The volume of the direct tone and the reflected tone reaching the listener is attenuated proportionally to the distance of the path to the listener (inversely proportional to the distance) relative to the volume of the sound source. Therefore, the distance attenuation processing unit 1313 can calculate the volume of the direct tone by dividing the volume of the sound source by the path length of the direct tone, and can calculate the volume of the reflected tone by dividing the volume of the sound source by the path length of the reflected tone.

[0476] In the selection process, the selection unit 1314 selects the object to be reflected based on parameters calculated prior to the selection process. A selection method of this disclosure may also be used in the selection of the object to be reflected.

[0477] Selection processing can be performed on all reflected sounds, or, as described above, on only reflected sounds with high evaluation values. That is, reflected sounds with low evaluation values ​​are automatically disqualified without any selection processing. For example, reflected sounds with very low volume can also be considered as having low evaluation values ​​and are therefore disqualified.

[0478] Furthermore, for example, selection processing can be applied to all reflected sounds. Also, the evaluation value of the selected reflected sounds in the selection process can be determined, and reflected sounds with low evaluation values ​​can be re-selected as not selected.

[0479] The selection and evaluation processes can be executed independently or in combination. When the selection and evaluation processes are executed in combination, one of the two processes can be executed first.

[0480] In the generation process, the generation unit 1315 generates direct tones and reflected tones. For example, the generation unit 1315 generates a direct tone based on the sound signal contained in the input signal, according to the arrival time and volume of the direct tone. Furthermore, regarding the reflected tone selected in the selection process, the generation unit 1315 generates a reflected tone based on the sound signal contained in the input signal, according to the arrival time and volume of the reflected tone.

[0481] In binaural processing, the binaural processing unit 1316 performs signal processing to make the sound signal of the direct tone perceived as sound arriving at the listener from the direction of the sound source object. Furthermore, the binaural processing unit 1316 performs signal processing to make the reflected tone selected by the selection unit 1314 perceived as sound arriving at the listener from the reflecting object.

[0482] For example, the binaural processing unit 1316 performs HRIR DB processing based on the position and orientation of the listener in the sound space, so that the sound reaches the listener from the position of the sound source object or the position of the obstacle object.

[0483] Additionally, HRIR (Head-Related Impulse Responses) describes the response characteristics when a single impulse is generated. Specifically, HRIR is the response characteristic obtained by transforming the head-related transfer function from its frequency domain representation to its time domain representation using a Fourier transform. This head-related transfer function represents the changes in sound produced by surrounding objects, including the auricle, head, and shoulders, as a transfer function. The HRIR DB is a database containing such information.

[0484] Furthermore, the position and orientation of the listener in the sound space can be, for example, the position and orientation of a virtual listener in a virtual sound space. Alternatively, the position and orientation of the virtual listener in the virtual sound space can change in accordance with the movement of the listener's head. Furthermore, the position and orientation of the virtual listener in the virtual sound space can also be determined based on information obtained from sensor 1405.

[0485] The program, spatial information, HRIR DB, threshold data or other parameters used in the above processing are obtained from the memory 1404 of the sound signal processing device 1001 or from outside the sound signal processing device 1001.

[0486] Furthermore, pipeline processing may also include other processing. Additionally, the rendering unit 1300 may include a processing unit (not shown) for performing other processing included in pipeline processing. For example, the rendering unit 1300 may also include a diffraction processing unit and an occlusion processing unit.

[0487] The diffraction processing unit performs processing to generate a sound signal representing a sound containing diffracted tones, which are caused by an obstacle object in a three-dimensional sound field (space) located between the listener and the sound source object. A diffracted tone is a sound that reaches the listener from the sound source object by bypassing an obstacle object when such an obstacle object exists between the sound source object and the listener.

[0488] The diffraction processing unit, for example, refers to the sound signal and metadata to calculate the path of the diffracted sound from the sound source object, bypassing the obstacle object, to reach the listener, and generates the diffracted sound based on the path. In the path calculation, the positions of the sound source object, the listener, and the obstacle object in the three-dimensional sound field (space), as well as the shape and size of the obstacle object, can also be used.

[0489] When a sound source object exists on the opposite side of an obstacle object, the occlusion processing unit generates a sound signal of the sound leaking from the sound source object through the obstacle object based on spatial information and information such as the material of the obstacle object.

[0490] [Example of a sound source object] In the above, the position information assigned to the sound source object represents a "point" in the virtual space as the position of the sound source object. That is, in the above, the sound source is defined as a "point sound source".

[0491] On the other hand, a sound source in virtual space can also be defined as an object with length, size, and shape, that is, a non-point sound source that extends spatially. In this case, the distance between the listener and the sound source and the direction of sound arrival are uncertain. Therefore, reflected sounds caused by such sound sources do not need to be analyzed by the analysis unit 1301, or are limited to being selected by the selection unit 1302 regardless of the analysis result. As a result, sound quality degradation that may occur due to failure to select reflected sounds can be avoided.

[0492] Alternatively, a representative point, such as the object's center of gravity, can be determined, and the processing disclosed herein can be applied assuming that sound originates from that point. In this case, the threshold can also be adjusted based on the spatial extension information of the sound source.

[0493] [Examples of direct and reflected sounds] For example, a direct sound is a sound that is not reflected by a reflecting object, while a reflected sound is a sound that is reflected by a reflecting object. A direct sound can also be a sound that reaches the listener from the sound source without being reflected by a reflecting object, and a reflected sound can also be a sound that reaches the listener from the sound source through a reflecting object.

[0494] Furthermore, direct tone and reflected tone are not limited to the sound that reaches the listener; they can also be the sound before it reaches the listener. For example, direct tone can also be the sound output from the sound source, or in other words, the sound of the sound source.

[0495] Figure 25 is a diagram illustrating sound transmission and diffraction. As shown in Figure 25, sometimes the direct sound does not reach the listener because of an obstacle object between the sound source and the listener. In this case, the sound emitted from the sound source, transmitted through the obstacle object, and reaching the listener can be considered as the direct sound. Furthermore, the sound emitted from the sound source, diffracted through the obstacle object, and reaching the listener can be considered as the reflected sound.

[0496] Furthermore, the two sounds compared in the selection process are not limited to the direct and reflected sounds of a sound emitted from a single sound source. For example, sound selection can also be performed by comparing two reflected sounds of a sound emitted from a single sound source. In this case, the direct sound in this disclosure can be replaced with the sound that arrives at the listener first, and the reflected sound in this disclosure can be replaced with the sound that arrives at the listener later.

[0497] [Example of Bitstream Construction] A bitstream may contain, for example, audio signals and metadata. The audio signal is audio data representing sound, and includes information related to the frequency and intensity of the sound. Furthermore, the metadata contains spatial information related to the space of the sound field, i.e., the sound space.

[0498] For example, spatial information is information relating to the space in which a listener is located when receiving sound based on a sound signal. Specifically, spatial information is information related to a predetermined location (location position) used to position a sound image at a specific location in sound space (e.g., a three-dimensional sound field), that is, information used to enable the listener to perceive sound arriving from a direction corresponding to that predetermined location. Spatial information may include, for example, information about the sound source and location information indicating the listener's position.

[0499] Sound source object information refers to the information about the sound source object that generates sound based on the sound signal. That is, sound source object information is information related to the object (sound source object) that reproduces the sound signal, and it is information related to a virtual sound source object configured in a virtual sound space. Here, the virtual sound space can also correspond to the real space where the object generating the sound is configured, and the sound source object in the virtual sound space can also correspond to the object generating sound in the real space.

[0500] Sound source object information can also represent the location of a sound source object configured in the sound space, the orientation of the sound source object, the directionality of the sound emitted by the sound source object, whether the sound source object is a living being, and whether the sound source object is a moving object. For example, a sound signal can be associated with one or more sound source objects represented by the sound source object information.

[0501] Bitstreams, for example, have a data structure consisting of metadata (control information) and sound signals.

[0502] The audio signal and metadata can be contained in a single bitstream or in multiple separate bitstreams. Furthermore, the audio signal and metadata can be contained in a single file or in multiple separate files.

[0503] Bitstreams can exist either per audio source or per playback time. When bitstreams exist per playback time, multiple bitstreams can be processed in parallel simultaneously.

[0504] Metadata can be assigned to each bitstream individually, or it can be assigned to multiple bitstreams together as information to control them. In this case, multiple bitstreams can also share the metadata. Alternatively, metadata can be assigned at each playback time.

[0505] In the presence of multiple bitstreams or multiple files, information indicating associated bitstreams or associated files may be included in more than one bitstream or more than one file. Alternatively, information indicating associated bitstreams or associated files may be included in each of the individual bitstreams or each of the individual files.

[0506] Here, associated bitstreams or associated files refer, for example, to bitstreams or files that may be used simultaneously during audio processing. Additionally, it may also include bitstreams or files that together contain information representing associated bitstreams or associated files.

[0507] Here, the information representing the associated bitstream or file can be, for example, an identifier representing the associated bitstream or file. Alternatively, the information representing the associated bitstream or file can be, for example, the filename, URL (Uniform Resource Locator), or URI (Uniform Resource Identifier).

[0508] In this case, the acquisition unit can also determine and acquire the associated bitstream or associated file based on information representing the associated bitstream or associated file. Alternatively, information representing the associated bitstream or associated file can be included in the bitstream or file, and also in other bitstreams or other files.

[0509] Here, the file containing information representing the associated bitstream or associated file can also be a control file such as a declaration file for content distribution.

[0510] In addition, all or part of the metadata can be obtained from outside the audio signal bitstream. For example, metadata for controlling the audio and metadata for controlling the video can be obtained from outside the bitstream, or metadata for both can be obtained from outside the bitstream.

[0511] Furthermore, metadata for controlling the image may also be included in the bitstream acquired by the stereo sound reproduction system 1000. In this case, the stereo sound reproduction system 1000 may also output the metadata for controlling the image to a display device that displays the image or a stereo image reproduction device that reproduces the stereo image.

[0512] [Examples of information contained in metadata] Metadata can also be information used in the description of a scene represented by a sound space. Here, a scene is a term that refers to the collection of all elements of a sound space, including three-dimensional images and sound events, modeled by a sound signal reproduction system using metadata.

[0513] That is, metadata includes not only information used to control audio processing, but also information used to control video processing. Metadata can contain only one of the information used to control audio processing or the information used to control video processing, or it can contain both.

[0514] The stereo sound reproduction system 1000 processes sound signals using metadata contained in the bitstream and interactive listener location information acquired through appending, to generate virtual sound effects. These sound effects can include initial reflection processing, obstacle removal, diffraction processing, masking, and reverberation processing, as well as other sound processing using metadata. For example, sound effects such as distance attenuation, localization, or Doppler effects can be added.

[0515] In addition, information can be attached to the metadata to toggle the on / off of all or some of the additional sound effects, or priority information for multiple processing of sound effects.

[0516] In addition, as an example, metadata includes information related to the sound space, including sound source objects and obstacle objects, and information related to the positioning location used to locate the sound image in a specified position within the sound space (i.e., to make the listener perceive the sound coming from a specified direction).

[0517] Here, an obstacle object is an object that may affect the listener's perception of sound by blocking or reflecting it before the sound emitted by the sound source reaches the listener. Besides stationary objects, obstacle objects can also include moving bodies such as animals or machines. Animals can also be people.

[0518] Furthermore, when multiple sound source objects exist in the sound space, for any given sound source object, the other sound source objects may become obstacle objects. That is, objects that do not emit sound, such as building materials or inanimate objects (i.e., non-sound-emitting objects), as well as sound source objects that emit sound, can all become obstacle objects.

[0519] The metadata contains all or part of the information representing the shape of the sound space, the shape and location of obstacle objects in the sound space, the shape and location of sound source objects in the sound space, and the location and orientation of the listener in the sound space.

[0520] The sound space can be either enclosed or open. Furthermore, the metadata can also include information about the reflectivity of obstacles within the sound space that can reflect sound. For example, the floor, walls, or ceiling that form the boundary of the sound space can also be considered obstacles.

[0521] Reflectivity is the energy ratio of reflected sound to incident sound, and it can be set for each frequency band of the sound. Of course, reflectivity can also be set uniformly regardless of the frequency band of the sound. In addition, when the sound space is an open space, parameters such as attenuation rate, diffraction tone, and initial reflection tone can be set uniformly, for example.

[0522] Metadata can also include information beyond reflectivity as parameters relating to obstacle or sound source objects. For example, metadata can also include information about the material of an object as parameters relating to both the sound source and the non-sound-producing object. Specifically, metadata can also include information such as diffusivity, transmissivity, and sound absorption.

[0523] Information related to a sound source object may also include volume, radiation characteristics (directivity), reproduction conditions, the number and type of sound sources in an object, and information representing the sound source region within the object. Reproduction conditions may, for example, specify whether the sound is a continuously flowing sound or an event-triggered sound. The sound source region within an object can be set based on the relative position of the listener and the object, or it can be set using the object as a reference.

[0524] For example, when the sound source area is set according to the relative position of the listener and the object, from the listener's perspective, the listener can perceive sound A from the right side of the object and sound B from the left side of the object.

[0525] Furthermore, when using an object as a reference to define a sound source region, it is possible to fix which area of ​​the object emits which sound. For example, when the listener views the object from the front, the listener can perceive high frequencies from the right side of the object and low frequencies from the left side. And when the listener views the object from the back, the listener can perceive low frequencies from the right side of the object and high frequencies from the left side.

[0526] Spatial metadata can also include the time up to the initial reflection, reverberation time, and the ratio of direct to diffuse sound. When the ratio of direct to diffuse sound is zero, the listener can perceive only the direct sound.

[0527] [Brief Summary] This implementation method is briefly summarized here.

[0528] When analyzing the relationship between direct and reflected tones, with the direct tone set as the preceding tone and the reflected tone as the following tone, in the case of a relationship that produces a priority effect, that is, when the reflected tone is below the echo detection limit, the reflected tone will not be perceived. Therefore, even if the reflected tone is deleted, the auditory impact on the listener is minimal.

[0529] Figure 26 is a diagram showing an example of the positional relationship between the listener and the obstacle object in this embodiment. Figure 27 is a diagram showing another example of the positional relationship between the listener and the obstacle object in this embodiment. Furthermore, the positional relationship shown in Figure 26 is the same as that shown in Figure 9, and the positional relationship shown in Figure 27 is the same as that shown in Figure 10. Additionally, Figure 28 is an example of the echo detection limit threshold in this embodiment. Furthermore, the echo detection limit threshold shown in Figure 28 is an example of the threshold data shown in Figure 12C, etc.

[0530] For example, when comparing the positional relationships in Figure 26 and Figure 27, the volume of the reflected sound heard by the listener in the positional relationship of Figure 26 is lower than that in the positional relationship of Figure 27. This is because the path length of the reflected sound in the positional relationship of Figure 26 is longer than that in the positional relationship of Figure 27.

[0531] Therefore, when judging solely by the volume of the reflected sound, the reflected sound shown in Figure 26 has a smaller auditory impact than the reflected sound shown in Figure 27. However, if we compare the arrival time of the reflected sound shown in Figure 26 at the listening position with the arrival time of the reflected sound shown in Figure 27, the case shown in Figure 26 is later.

[0532] Therefore, judging from the perspective of the echo detection limit, as shown in Figure 28, the reflected sound shown in Figure 27 is below the echo detection limit, so it will not be perceived as a reflected sound by the listener, while the reflected sound shown in Figure 26 is above the echo detection limit, so it will be perceived as a reflected sound by the listener.

[0533] In this embodiment, this situation is used to determine the auditory importance of the reflected sound, so that unimportant reflected sounds are not reproduced, thereby reducing the amount of computation related to the processing of reflected sounds.

[0534] The above is a brief summary of this implementation method.

[0535] Therefore, the following research can be conducted.

[0536] As described above, when the direct tone is set as the preceding tone and the reflected tone as the following tone, in the case of a relationship that produces a priority effect, by moving the sound image position of the reflected tone to the position of the direct tone and combining (merging) the sound signals representing the reflected tone and the direct tone into a single sound signal, it is possible to reduce the processing related to the reflected tone. Furthermore, even when using the merged single sound signal, the listener rarely experiences dissonance.

[0537] Here, we consider the arrival time difference between the direct tone and the reflected tone. Compared to the positional relationship shown in Figure 26, the arrival time difference is shorter in the case shown in Figure 27. For example, we investigate the case where no priority effect occurs in the case shown in Figure 26, but a priority effect occurs in the case shown in Figure 27. In this case, as shown in Figure 27, the listener will perceive the sound image position of the reflected tone as if it has moved to the sound image position of the direct tone.

[0538] Taking advantage of this, under the condition that the priority effect is met, direct tones and reflected tones can be integrated to generate a single sound signal (combined sound signal) representing a synthesized tone (merged tone), thereby reducing the computational load associated with processing reflected tones. In this case, the volume of the synthesized tone is set to the sum of the volumes of the direct tone and the reflected tone.

[0539] Figure 29 is a diagram illustrating an example of the shift in the sound image position of the reflected sound in the positional relationship shown in Figure 27. Under the conditions shown in Figure 27, where a priority effect occurs, as shown in Figure 29, processing is performed to shift the sound image position of the reflected sound towards the sound image position of the direct sound, increasing the volume of the direct sound by an amount corresponding to the volume of the reflected sound. As a result, processing related to the reflected sound can be reduced.

[0540] However, in processes that generate merged audio signals only when a priority effect has occurred, the following problems sometimes arise. These are illustrated using Figures 30-32.

[0541] Figure 30 is a diagram illustrating an example of direct and indirect tones arriving at the listener from the same direction. In Figure 30, the indirect tone is a reflected tone caused by reflection from a wall.

[0542] Figure 31 illustrates an example of direct and indirect tones reaching the listener from one sound source (sound source 101) and another sound source (sound source 102), respectively. One sound source (sound source 101) and the other sound source (sound source 102) emit different kinds of sounds; for example, one sound source (sound source 101) emits the sound of a dog barking, and the other sound source (sound source 102) emits the sound of a vehicle driving. As shown in Figure 31, the indirect tone (reflected tone) from one sound source (sound source 101) and the direct tone from the other sound source (sound source 102) reach the listener from the same direction.

[0543] Figure 32 is a diagram illustrating an example of transmitted and diffracted tones arriving at a listener as a sound emanating from a sound source. In Figure 32, the transmitted tone is the preceding tone, and the diffracted tone, as an indirect tone, is the following tone.

[0544] In the examples shown in Figures 30-32, both direct and indirect tones arrive at the listener from the same or approximately the same direction. In this case, it should be possible to integrate the direct and indirect tones, that is, to combine the sound signals representing the indirect tones and the sound signals representing the direct tones into a single sound signal. This is because even with this processing, it rarely causes dissonance to the listener, and the processing of reflected tones can be reduced.

[0545] However, conditions regarding priority effects alone are sometimes insufficient to determine whether a merging process should be performed.

[0546] As a case where it is impossible to determine, an example shown in Figure 30 could be given where the arrival time difference between the direct tone and the indirect tone (reflected tone) is too short or too long to meet the condition for a priority effect. Furthermore, as a case where it is impossible to determine, an example shown in Figure 31 could be given where the indirect tone and the direct tone are dissimilar, i.e., the indirect tone does not originate from the direct tone. Additionally, as a case where it is impossible to determine, an example shown in Figure 32 could be given where the volume of the preceding tone is lower than the volume of the following tone.

[0547] In these cases, since no priority effect is generated, the aforementioned processing, which generates a merged audio signal only when a priority effect is generated, cannot produce a merged audio signal, i.e., it cannot reduce the processing of reflected sounds. In other words, such audio signal processing methods result in the inability to reduce computational load and computational complexity.

[0548] Therefore, the following section will further explain in detail a sound signal processing method that can reduce computational load and computational complexity in the sound space.

[0549] (Embodiment 2) Hereinafter, Embodiment 2 will be described. The description will focus on the differences from Embodiment 1, and the description of the commonalities will be omitted or simplified.

[0550] [Configuration of the rendering unit] First, the configuration of the rendering unit 2300 in this embodiment will be described. FIG33 is a block diagram showing an example of the configuration of the rendering unit 2300 in this embodiment.

[0551] The rendering unit 2300 includes a resolution unit 2301, a selection unit 2302, and a reproduction unit 2303. Furthermore, the audio signal processing apparatus of this embodiment is an example of a decoding apparatus, which includes a decoder, and the decoder includes the rendering unit 2300. That is, it can be said that the audio signal processing apparatus of this embodiment includes a resolution unit 2301, a selection unit 2302, and a reproduction unit 2303. The rendering unit 2300 performs additional audio processing on the audio data contained in the input signal and outputs it.

[0552] Similar to Implementation 1, the input signal consists of, for example, spatial information, sensor information, and sound data. The spatial information also includes physical information such as the reflection coefficient, transmission coefficient, and diffraction coefficient of the non-sound-emitting object (obstacle object).

[0553] Furthermore, in this embodiment, reflected sounds are mainly used as an example of indirect sounds for explanation, but the same treatment is applied even when indirect sounds are used instead of reflected sounds. Additionally, as an example, indirect sounds are reflected sounds or diffracted sounds, etc.

[0554] The analysis unit 2301 can perform all or part of the processing performed by the analysis unit 1301 in Embodiment 1. Furthermore, like the analysis unit 1301 in Embodiment 1, the analysis unit 2301 analyzes the sound signal contained in the input signal and the spatial information received from the spatial information management units 1201 and 1211. Therefore, the analysis unit 2301 calculates the information required for generating direct and reflected sounds in the reproduction unit 2303, as well as the information required for selecting whether to generate reflected sounds. The method by which the analysis unit 2301 calculates this information is as described in Embodiment 1.

[0555] Furthermore, the analysis unit 2301 performs analysis processing of the input signal as in S101 of Embodiment 1, as in the analysis unit 1301 of FIG8. That is, the analysis unit 2301 analyzes the input signal input to the sound signal processing device of this embodiment and detects direct sounds and reflected sounds that may be generated in the sound space.

[0556] In this embodiment, for example, an indirect tone (more specifically, a reflected tone) is an example of the first sound, and a direct tone is an example of the second sound. The first sound and the second sound are different from each other.

[0557] When both direct and reflected sounds are detected, the analysis unit 2301 generates a sound signal representing the reflected sound and a sound signal representing the direct sound based on spatial information and sound data.

[0558] More specifically, the analysis unit 2301 generates sound signals representing reflected sound and sound signals representing direct sound based on the location information of the sound source object, the location information of the non-sound-producing object (obstacle object), the location information and physical information of the listener, and sound data contained in the spatial information.

[0559] That is, the analysis unit 2301 generates a sound signal generated in the virtual space based on spatial information and sound data. Here, a sound signal representing a first sound (first sound signal) and a sound signal representing a second sound (second sound signal) are generated. In this embodiment, the first sound signal is equivalent to a sound signal representing a reflected sound, and the second sound signal is equivalent to a sound signal representing a direct sound.

[0560] Furthermore, the sound signal generated by the analysis unit 2301 is given attribute information for determining the attributes of the sound signal; that is, the analysis unit 2301 generates a sound signal containing such attribute information. The analysis unit 2301 also generates attribute information. The sound signal containing attribute information is generated for each sound produced in the virtual space.

[0561] As described above, the first sound (reflected sound (indirect sound)) or the second sound (direct sound) is represented by a sound signal, but is not limited thereto. It is also possible that the properties of the sound signal contain information indicating whether the sound represented by the sound signal is the first sound (reflected sound (indirect sound)) or the second sound (direct sound).

[0562] Sometimes, a sound signal whose attributes contain information representing reflected sound (indirect sound) is recorded as a sound signal representing reflected sound (indirect sound), and a sound signal whose attributes contain information representing direct sound is recorded as a sound signal representing direct sound.

[0563] The first sound signal is a sound signal containing first attribute information. The first attribute information is information used to determine the attributes of the first sound signal. This attribute may also include information representing the first sound. The second sound signal is a sound signal containing second attribute information. The second attribute information is information used to determine the attributes of the second sound signal. This attribute may also include information representing the second sound.

[0564] Furthermore, the attribute information can also include information necessary for radiating the sound signal into the sound space, such as gain information, gain characteristics for each bandwidth, positional information, and directivity information. That is, this necessary information can also be retained in the attribute information. Additionally, the attribute information can be associated with the sound signal as metadata. For example, the gain characteristics of the sound signal for each bandwidth contained in the attribute information can be determined based on spatial information contained in the input information. Information representing frequency characteristics indicating auditory sensitivity can be determined based on spatial information contained in the input information, and in particular, can be determined as information associated with the listener's persona.

[0565] A sound that reaches the listener's head directly from a sound source is a direct sound. A sound that reaches the listener's head after being output from a sound source and reflected or diffracted by a non-sound-producing object is an indirect sound (reflected sound or diffracted sound).

[0566] In this embodiment, the analysis unit 2301 generates a sound signal representing a reflected tone (indirect tone) and a sound signal representing a direct tone related to the reflected tone (indirect tone).

[0567] Furthermore, the direct tone of an indirect tone refers to a direct tone originating from the same sound source as the indirect tone. The indirect tone of a direct tone refers to an indirect tone originating from the same sound source as the direct tone. More specifically, a reflected tone refers to the sound produced when the direct tone of the reflected tone is reflected by a reflecting object.

[0568] The sound signal representing a reflected tone (indirect tone) contains information about the sound signal representing the direct tone of the reflected tone (indirect tone).

[0569] Furthermore, in this embodiment, the analysis unit 2301 also generates a sound signal representing a reflected sound (indirect sound) and a sound signal representing a direct sound originating from a sound source different from the reflected sound (indirect sound). That is, the analysis unit 2301 generates a sound signal representing a reflected sound (indirect sound) originating from one sound source, and also generates a sound signal representing a direct sound originating from another sound source.

[0570] The analysis unit 2301 can store the generated sound signals in its own memory. Furthermore, the analysis unit 2301 can generate multiple sound signals and store them in the memory.

[0571] Furthermore, similar to Embodiment 1, the analysis unit 2301 can calculate values ​​related to the path to the listening position, the time taken to reach the position, and the volume at the time of arrival for both direct and reflected sounds. Similarly, the analysis unit 2301 can calculate information representing the relationship between the direct and reflected sounds, such as values ​​related to the time difference between the arrival of the direct and reflected sounds (the time difference between the direct and reflected sounds).

[0572] Furthermore, examples of arrival volume include the volume of reflected sound (lr) and the volume of direct sound (ld). The volume of direct sound arrival (ld) refers to the volume of the direct sound reaching the listener's location in virtual space, i.e., the listening position; in other words, it is the volume of the direct sound at the listening position. The volume of reflected sound arrival (lr) refers to the volume of a reflected sound, an example of an indirect sound, reaching the listening position; in other words, it is the volume of the indirect sound (reflected sound volume) at the listening position.

[0573] In this embodiment, the sound signal representing the reflected sound and the sound signal representing the direct sound may each include information indicating the volume of the sound represented by the sound signal at the listening position. That is, in this embodiment, the sound signal representing the reflected sound includes information indicating the volume (lr) when the reflected sound arrives as the indirect sound volume (reflected sound volume). Similarly, the sound signal representing the direct sound includes information indicating the volume (ld) when the direct sound arrives as the direct sound volume.

[0574] More specifically, the first attribute information includes first volume information indicating the volume (lr) of the reflected sound as the volume of the first sound, and the second attribute information includes second volume information indicating the volume (ld) of the direct sound as the volume of the second sound.

[0575] Furthermore, a sound signal representing a reflected tone may also include information representing the volume of the indirect tone (reflected tone volume) and the volume of the direct tone relating to that reflected tone. Similarly, a sound signal representing a direct tone may include information representing the volume of the direct tone and the volume of the indirect tone (reflected tone) relating to that direct tone.

[0576] Furthermore, the analysis unit 2301 generates first position information and second position information based on the position information of the sound source object, the position information of the non-sound-producing object (obstacle object), the position information and physical information of the listener, and the sound data contained in the spatial information.

[0577] The first position information indicates the position of the first sound, which is the information used to locate the sound image of the first sound. The second position information indicates the position of the second sound, which is the information used to locate the sound image of the second sound.

[0578] In this embodiment, the first attribute information includes the first location information, and the second attribute information includes the second location information.

[0579] The selection unit 2302 can perform all or part of the processing performed by the selection unit 1302 in Embodiment 1. Furthermore, the selection unit 2302 determines whether the reproduction unit 2303 outputs (reproduces) an output signal based on the sound signal generated by the analysis unit 2301. That is, the selection unit 2302 first determines whether to merge multiple sound signals generated by the analysis unit 2301 (e.g., the first sound signal and the second sound signal), and if it determines to merge multiple sound signals, generates a merged sound signal. The generated merged sound signal is output by the reproduction unit 2303.

[0580] The selection unit 2302 has an acquisition unit 2302a, a decision unit 2302b, and a merging unit 2302c.

[0581] The acquisition unit 2302a acquires an audio signal containing attribute information generated by the analysis unit 2301 and stored in the memory of the analysis unit 2301. For example, the acquisition unit 2302a acquires a first audio signal and a second audio signal. Additionally, the acquisition unit 2302a acquires values ​​related to the time difference between direct and reflected sounds calculated by the analysis unit 2301.

[0582] The decision unit 2302b calculates the first direction of arrival of the first sound (reflected sound) at the listener's location, i.e., the listening position. The decision unit 2302b calculates the second direction of arrival of the second sound (direct sound) at the listening position. The decision unit 2302b calculates the direction of arrival of the reflected sound as the first direction of arrival and the direction of arrival of the direct sound as the second direction of arrival using the method described in step S231 of FIG. 23 in Embodiment 1.

[0583] Furthermore, the decision unit 2302b determines whether to merge the acquired first sound signal and the acquired second sound signal based on the indicators corresponding to the calculated first and second arrival directions.

[0584] When the decision unit 2302b determines that the first audio signal and the second audio signal should be merged, the merging unit 2302c generates a merged audio signal after merging the first audio signal and the second audio signal. In this case, the merging unit 2302c outputs the generated merged audio signal to the reproduction unit 2303.

[0585] Furthermore, if it is determined that the first and second audio signals will not be merged, the merging unit 2302c does not generate a merged audio signal. In this case, the merging unit 2302c outputs the first and second audio signals to the reproduction unit 2303.

[0586] Furthermore, the combined sound signal represents the signal of the combined sound resulting from combining the first sound and the second sound. The volume of the combined sound is the volume obtained by adding the volume of the first sound to the volume of the second sound. In addition, when the first sound is a reflected sound representing a dog's bark and the second sound is a direct sound representing a vehicle's driving sound, the combined sound is the sound of combining the dog's bark sound and the vehicle's driving sound.

[0587] The reproduction unit 2303 can perform all or part of the processing performed by the reproduction unit 1303 in Embodiment 1. In addition, the reproduction unit 2303 acquires the merged audio signal, the first audio signal and the second audio signal output from the selection unit 2302 (more specifically, the merging unit 2302c), and outputs output signals based on the acquired merged audio signal, the first audio signal and the second audio signal respectively.

[0588] The reproduction unit 2303 generates and outputs an output signal based on each of the acquired combined audio signal, the first audio signal, and the second audio signal by performing binaural filtering and other processing on each of them. The binaural filtering is implemented, for example, by applying a head correlation transfer function (e.g., convolution) to each of the acquired combined audio signal, the first audio signal, and the second audio signal. In this embodiment, for example, the output signal based on the combined audio signal is generated by applying a head correlation transfer function to the combined audio signal.

[0589] Furthermore, the reproduction unit 2303 can also generate and output output signals based on each of the merged sound signal, the first sound signal, and the second sound signal output from the selection unit 2302 by performing both binaural filtering and diffusion filtering on the merged sound signal, the first sound signal, and the second sound signal, respectively. The diffusion filtering process, for example, is a process that enhances the realism of the merged sound by diffusing the merged sound represented by the acquired merged sound signal into the merged sound signal. Furthermore, the diffusion filtering process uses a filter that realistically simulates the audible strength of the diffusion of the merged sound represented by the acquired merged sound signal (i.e., simulates the audible strength of the sound diffusion perceived by the listener). In the diffusion filtering process, a finite-pulse filter and / or an infinite-pulse filter are used.

[0590] Hereinafter, an example of the operation of the sound signal processing method performed by the sound signal processing apparatus (more specifically, the rendering unit 2300) of this embodiment will be described.

[0591] [Example 1 of the operation of the rendering unit] Figure 34 is a flowchart showing Example 1 of the operation of the sound signal processing apparatus of this embodiment. Figure 34 shows the processing mainly performed by the rendering unit 2300 provided by the sound signal processing apparatus of this embodiment. In addition, the description of the common points with Figure 8 of Embodiment 1 is omitted or simplified here.

[0592] In this example of action, the first sound is an indirect sound (more specifically, a reflected sound), and the second sound is a direct sound.

[0593] First, the analysis unit 2301 performs analysis processing (S101a) to analyze the input signal. For example, the analysis unit 2301 generates a plurality of first sound signals and a plurality of second sound signals, and stores the generated plurality of first sound signals and a plurality of second sound signals in the memory of the analysis unit 2301.

[0594] The selection unit 2302 decides whether to merge the first audio signal and the second audio signal. If the decision is to merge, a merged audio signal is generated and output (S102a).

[0595] The reproduction unit 2303 acquires the merged audio signal output from the selection unit 2302 and outputs an output signal based on the acquired merged audio signal (S103a).

[0596] Furthermore, the processing performed by the selection unit 2302 and the reproduction unit 2303 will be explained in more detail using FIG35.

[0597] Figure 35 is a flowchart illustrating an example of the processing performed by the selection unit 2302 and the reproduction unit 2303 in Operation Example 1 of this embodiment. Furthermore, descriptions of similarities with Figure 14 of Embodiment 1 are omitted or simplified here.

[0598] First, the selection unit 2302 selects the reflected sound and direct sound detected by the analysis unit 2301 (S501). That is, the acquisition unit 2302a of the selection unit 2302 selects one of a plurality of first sound signals generated by the analysis unit 2301 and stored in the memory, and acquires the selected first sound signal. In addition, the acquisition unit 2302a selects one of a plurality of second sound signals generated by the analysis unit 2301 and stored in the memory, and acquires the selected second sound signal.

[0599] Furthermore, the decision unit 2302b of the selection unit 2302 calculates the first arrival direction (reflected sound arrival direction) and the second arrival direction (direct sound arrival direction). More specifically, the decision unit 2302b calculates the reflected sound arrival direction with respect to the acquired first sound signal and calculates the direct sound arrival direction with respect to the acquired second sound signal.

[0600] The decision unit 2302b determines whether to merge the acquired first sound signal and the acquired second sound signal based on the indicators corresponding to the calculated first and second arrival directions. As an example, the decision unit 2302b performs the following processing.

[0601] The decision unit 2302b calculates the arrival angle δ between the reflected sound and the direct sound (S502). That is, the decision unit 2302b calculates the angle δ formed by the first arrival direction and the second arrival direction. This arrival angle δ is an example of the above-mentioned index.

[0602] Here, we will explain the indicator (arrival angle δ).

[0603] As explained in Embodiment 1, the orientation information of the sound source and the orientation information of the listener can be represented by azimuth and pitch, respectively. Therefore, the arrival angle δ, as an example of an indicator, can also be represented by azimuth and pitch. The indicator in this embodiment is an indicator composed of two orthogonal axes: the azimuth axis and the pitch axis.

[0604] The decision unit 2302b compares the calculated arrival angle δ with the discrimination limit T of the arrival angle δ. The discrimination limit T of the arrival angle δ refers to the smallest angle at which the listener can perceive the movement of the sound image when the sound image moves. The discrimination limit T can use a known value (e.g., Figure 2.26 of Non-Patent Document 2).

[0605] The discrimination limit T can be preset to a predetermined value before performing this action example. This discrimination limit T (predetermined value) can be stored in the memory of the parsing unit 2301 for example, and the acquisition unit 2302a acquires the discrimination limit T stored in the memory.

[0606] Furthermore, the decision unit 2302b determines whether to merge the first sound signal and the acquired second sound signal based on the indicators corresponding to the calculated first and second arrival directions. That is, the decision unit 2302b determines whether the arrival angle δ based on the first and second arrival directions is less than the discrimination limit T (S503). Thus, the decision unit 2302b decides whether to merge the first and second sound signals.

[0607] When the arrival angle δ is less than the discrimination limit T ("Yes" in S503), the decision unit 2302b decides to merge the first sound signal and the second sound signal. Furthermore, the merging unit 2302c synthesizes (merges) the reflected tone (first sound) and the direct tone (second sound) (S504). That is, the merging unit 2302c generates a merged sound signal after merging the first sound signal and the second sound signal. The merging unit 2302c outputs the generated merged sound signal to the reproduction unit 2303. In addition, in this case, the merging unit 2302c can perform a process that invalidates the first sound signal and the second sound signal before merging. The invalidated first sound signal and the second sound signal will not be output by the reproduction unit 2303.

[0608] Furthermore, the reproduction unit 2303 reproduces the synthesized sound (merged sound) (S505). That is, the reproduction unit 2303 acquires the merged sound signal output from the selection unit 2302 and outputs an output signal based on the acquired merged sound signal. More specifically, the reproduction unit 2303 outputs an output signal generated by applying a head correlation transfer function to the merged sound signal.

[0609] When the arrival angle δ is greater than or equal to the discrimination limit T ("No" in S503), the decision unit 2302b decides not to merge the first sound signal and the second sound signal. The merging unit 2302c outputs the first sound signal and the second sound signal to the reproduction unit 2303.

[0610] The reproduction unit 2303 reproduces both reflected sound and direct sound (S506). That is, the reproduction unit 2303 acquires the first sound signal and the second sound signal output from the selection unit 2302, and outputs an output signal based on the acquired first sound signal and an output signal based on the acquired second sound signal. More specifically, the reproduction unit 2303 outputs an output signal generated by applying a head correlation transfer function to the first sound signal, and an output signal generated by applying a head correlation transfer function to the second sound signal.

[0611] Here, the set of the first sound (reflected sound) represented by one of the multiple first sound signals and multiple second sound signals generated and stored in the memory by the analysis unit 2301, and the set of the second sound (direct sound) represented by one of the first sound signals, is taken as a first combination. The selection unit 2302 determines whether the processing of steps S501 to S506 has been performed on all the first combinations (S507). That is, the selection unit 2302 determines whether the processing of steps S501 to S506 has been performed on the set of each of the first combinations, that is, the first sound signal and the second sound signal representing the first sound and the second sound.

[0612] If the answer in step S507 is "No", then the process in step S501 is repeated. If the answer in step S507 is "Yes", then the action ends.

[0613] [Rendering Unit Action Example 2] Action Example 2 performs the same process as Action Example 1, except that it performs the process shown in Figure 36 instead of the process shown in Figure 35.

[0614] Figure 36 is a flowchart illustrating an example of the processing performed by the selection unit 2302 and the reproduction unit 2303 in Operation Example 2 of this embodiment. Furthermore, the description of commonalities with Operation Example 1 of Embodiment 2 is omitted or simplified here.

[0615] In Action Example 2, the steps S101a shown in FIG34 are performed in the same manner as in Action Example 1.

[0616] As shown in FIG36, the selection unit 2302 specifies the direct tone detected by the analysis unit 2301 (S501a). That is, the acquisition unit 2302a of the selection unit 2302 specifies one of the multiple second sound signals generated by the analysis unit 2301 and stored in the memory, and acquires the specified second sound signal.

[0617] The decision unit 2302b of the selection unit 2302 calculates the second direction of arrival (direct tone direction of arrival) (S508).

[0618] The selection unit 2302 selects the reflected sound detected by the analysis unit 2301 (S501aa). That is, the acquisition unit 2302a of the selection unit 2302 selects one of the multiple first sound signals generated by the analysis unit 2301 and stored in the memory, and acquires the selected first sound signal.

[0619] The decision unit 2302b of the selection unit 2302 calculates the first direction of arrival (direction of arrival of reflected sound) (S509).

[0620] The decision unit 2302b calculates the arrival angle δ between the reflected sound and the direct sound (S502). That is, the decision unit 2302b calculates the angle δ formed by the first arrival direction and the second arrival direction.

[0621] The decision unit 2302b determines the discrimination limit T of the arrival angle δ relative to the direction of arrival of the direct tone (S509a). Here, the decision unit 2302b determines the discrimination limit T corresponding to the arrival angle δ relative to the direction of arrival of the direct tone (the second direction of arrival).

[0622] Here, the case where the direct tone arrives from above, behind, below, to the right, or to the left of the listener is investigated. The discrimination limit T of the arrival angle δ in this case can be determined based on its direction. Here, a database of the discrimination limits T of the arrival angle δ pre-set for each direct tone arrival direction is stored in the memory of the analysis unit 2301, and the determination unit 2302b determines the discrimination limit T corresponding to the calculated second arrival direction.

[0623] In this database, the direction of arrival of a direct tone is defined as having a one-to-one correspondence with a discrimination limit T. For example, the decision unit 2302b refers to this database, selects a discrimination limit T that corresponds one-to-one with the direction of arrival of the direct tone (the second direction of arrival), extracts the selected value, and determines the extracted value as the value of the discrimination limit T.

[0624] A database of the discrimination limit T for each direct tone arrival angle δ, which is predetermined according to the direction of arrival, can also be developed based on actual measurements through experiments conducted by test subjects.

[0625] On the other hand, if the device or service equipped with this technology has a database of head-related transfer functions (so-called HRTF (Head Related Impulse Response) SOFA (Spatially Oriented Format for Acoustics)), the discrimination limit T of the arrival angle δ can be determined for each arrival direction by calculation processing of the SOFA conforming to the HRTF. The method will be described below.

[0626] Figure 37 is a diagram showing the SOFA of the HRTF in this embodiment. Figure 38 is a diagram showing a cone used to illustrate the relationship between the position of the HRTF and the listening position in this embodiment.

[0627] In Figure 37, the listener is located at the center of the celestial sphere, on which multiple points are arranged. Each of these multiple points is assigned (arranged) an HRTF.

[0628] Figure 38 shows a cone with the listener's position (listening position) as the vertex, the direction of the direct sound as the direction of the perpendicular line, and the angle between the perpendicular line and the generatrix as γ.

[0629] When applying such a cone to the SOFA of HRTF, the largest γ, where no more than K HRTFs are defined within the cone, is set as the discrimination limit T of the arrival angle δ in the direction of arrival of the direct tone. Ideally, the discrimination limit T should be set to K=1. However, in reality, by setting N to a sufficiently large value (K=10, etc.) to represent the discrimination limit T as a large value, the chances of synthesizing the direct tone and the reflected tone can be increased, thus reducing the computational load. Such a database can be used in step S509a.

[0630] Figure 36 will be used again for illustration. The decision unit 2302b compares the calculated arrival angle δ with the discrimination limit T of the determined arrival angle δ.

[0631] The decision unit 2302b determines whether the arrival angle δ based on the first arrival direction and the second arrival direction is less than the discrimination limit T (S503). Therefore, the decision unit 2302b decides whether to merge the first sound signal and the second sound signal.

[0632] When the arrival angle δ is less than the discrimination limit T ("Yes" in S503), the decision unit 2302b decides to merge the first sound signal and the second sound signal. The merging unit 2302c synthesizes (merges) the reflected sound (first sound) and the direct sound (second sound) (S504).

[0633] The reproduction section 2303 reproduces the synthesized sound (combined sound) (S505).

[0634] When the arrival angle δ is greater than or equal to the discrimination limit T ("No" in S503), the decision unit 2302b decides not to merge the first sound signal and the second sound signal. The merging unit 2302c outputs the first sound signal and the second sound signal to the reproduction unit 2303.

[0635] The reproduction unit 2303 reproduces both reflected sounds and direct sounds (S506).

[0636] The selection unit 2302 determines whether steps S501aa to S506 have been processed on all of the plurality of first sound signals generated by the analysis unit 2301 and stored in the memory. That is, the selection unit 2302 determines whether steps S501aa to S506 have been processed on all reflected sounds represented by the plurality of first sound signals generated by the analysis unit 2301 and stored in the memory (S507a).

[0637] If the answer in step S507a is "No", then the process in step S501aa is repeated. If the answer in step S507a is "Yes", then the following process is performed further.

[0638] The selection unit 2302 determines whether steps S501a to S507a have been performed on all of the plurality of second sound signals created by the analysis unit 2301 and stored in the memory. That is, the selection unit 2302 determines whether steps S501a to S507a (S507aa) have been performed on all the direct tones represented by the plurality of second sound signals created by the analysis unit 2301 and stored in the memory.

[0639] If the answer in step S507aa is "No", then the process in step S501a is repeated. If the answer in step S507aa is "Yes", then the operation ends.

[0640] [Rendering Unit Action Example 3] Action Example 3 performs the same process as Action Example 1, except that it performs the process shown in FIG39 instead of the process shown in FIG35.

[0641] Figure 39 is a flowchart illustrating an example of the processing performed by the selection unit 2302 and the reproduction unit 2303 in Operation Example 3 of this embodiment. Furthermore, descriptions of commonalities with Operation Example 1 of Embodiment 2 are omitted or simplified here.

[0642] Furthermore, in Action Example 1, the first sound is an indirect sound (more specifically, a reflected sound), and the second sound is a direct sound. However, in Action Example 3, this is not the case. That is, in Action Example 2, both the first and second sounds can be direct sounds, both the first and second sounds can be indirect sounds, the first sound can be a direct sound and the second sound can be an indirect sound, or the first sound can be an indirect sound and the second sound can be a direct sound.

[0643] In Action Example 3, the steps S101a shown in FIG34 are performed in the same manner as in Action Example 1.

[0644] Furthermore, as shown in FIG39, the selection unit 2302 specifies the two sounds detected by the analysis unit 2301 (S501b). That is, the acquisition unit 2302a of the selection unit 2302 specifies two sound signals from the plurality of first sound signals and the plurality of second sound signals generated and stored in the memory by the analysis unit 2301, and acquires the specified two sound signals.

[0645] Furthermore, the two specified sounds can be two first sounds, two second sounds, or one first sound and one second sound. That is, the two specified sound signals can be two first sound signals, two second sound signals, or one first sound signal and one second sound signal.

[0646] The decision unit 2302b of the selection unit 2302 calculates the direction of arrival of each of the two sounds.

[0647] The decision unit 2302b calculates the arrival angle δ of the two sounds (S502b). That is, the decision unit 2302b calculates the angle δ formed by the arrival direction of one of the two sounds and the arrival direction of the other sound.

[0648] For example, when the two specified sounds are two first sounds, the determination unit 2302b calculates the angle δ formed by the first arrival direction corresponding to one first sound and the first arrival direction corresponding to the other first sound. Similarly, when the two specified sounds are two second sounds, the determination unit 2302b calculates the angle δ formed by the second arrival direction corresponding to one second sound and the second arrival direction corresponding to the other second sound. Furthermore, when the two specified sounds are one first sound and one second sound, the determination unit 2302b calculates the angle δ formed by the first arrival direction corresponding to one first sound and the second arrival direction corresponding to one second sound.

[0649] The decision unit 2302b compares the calculated arrival angle δ with the discrimination limit T of the arrival angle δ. In this example, the discrimination limit T is determined as follows.

[0650] First, before performing this action example, the memory of the analysis unit 2301 stores a database showing the values ​​of the discrimination limit T corresponding to the direction of arrival of each sound. The decision unit 2302b selects one of the two specified sounds. The decision unit 2302b extracts the value corresponding to the direction of arrival of the selected sound by referring to the database (S510), and determines the extracted value as the value of the discrimination limit T.

[0651] In this database, the direction of arrival of a sound is defined as having a one-to-one correspondence with a discrimination limit T. For example, the decision unit 2302b refers to this database, selects a discrimination limit T that corresponds one-to-one with the first direction of arrival, extracts the selected value, and determines the extracted value as the discrimination limit T.

[0652] Furthermore, not limited to this, in this action example, the discrimination limit T can also be preset to a specified value before performing this action example.

[0653] The decision unit 2302b determines whether to merge the two sound signals based on the indices corresponding to the calculated arrival directions of the two sounds. That is, the decision unit 2302b determines whether the arrival angle δ based on the two arrival directions is less than the discrimination limit T (S503b). Thus, the decision unit 2302b decides whether to merge the two sound signals.

[0654] When the arrival angle δ is less than the discrimination limit T ("Yes" in S503b), the decision unit 2302b decides to merge the two sound signals. Then, the merging unit 2302c merges the two sounds (S504b). That is, the merging unit 2302c generates a merged sound signal that combines the two sound signals. The merging unit 2302c outputs the generated merged sound signal to the reproduction unit 2303. Furthermore, in this case, the merging unit 2302c can perform a process that invalidates the two sound signals before merging. The two invalidated sound signals are not output by the reproduction unit 2303.

[0655] When the arrival angle δ is greater than or equal to the discrimination limit T ("No" in S503b), the decision unit 2302b decides not to merge the two sound signals. The merging unit 2302c outputs the two sound signals to the reproduction unit 2303. The two sounds represented by these two sound signals are equivalent to the unmerged sounds.

[0656] Here, a set of two sounds represented by two of the multiple first sound signals generated and stored in the memory by the analysis unit 2301 and the multiple second sound signals is defined as a second combination. The selection unit 2302 determines whether steps S501b to S504b have been processed (S507b) for all second combinations. That is, the selection unit 2302 determines whether steps S501b to S504b have been processed for each set of all second combinations, i.e., the two sound signals representing the two sounds.

[0657] If the answer in step S507b is "No", then the process in step S501b is repeated. If the answer in step S507b is "Yes", then the following process is performed.

[0658] The reproduction unit 2303 applies a head-related transfer function based on the direction of arrival of both unmerged and merged sounds (merged tones) (S505b). More specifically, the reproduction unit 2303 outputs an output signal generated by applying a head-related transfer function based on the direction of arrival of the sound represented by one of the two sound signals determined by the determination unit 2302b to be unmerged. Furthermore, the reproduction unit 2303 outputs an output signal generated by applying a head-related transfer function based on the direction of arrival of the sound represented by the other of the two sound signals determined by the determination unit 2302b to be unmerged.

[0659] That is, for example, if the two audio signals determined not to be merged are a first audio signal and a second audio signal, the reproduction unit 2303 performs the following processing. The reproduction unit 2303 outputs an output signal generated by applying a head correlation transfer function based on the first arrival direction to the first audio signal, and an output signal generated by applying a head correlation transfer function based on the second arrival direction to the second audio signal.

[0660] Furthermore, the reproduction unit 2303 outputs an output signal generated by applying a head-related transfer function based on the direction of arrival of the merged sound signal generated by the merging unit 2302c, wherein the direction of arrival of the merged sound signal is the direction in which the merged sound represented by the merged sound signal reaches the listening position.

[0661] Furthermore, the reproduction unit 2303 may, for example, use a database (e.g., SOFA of HRIR) that illustrates head-related transfer functions defined for each direction of sound arrival to select the head-related transfer function to apply. In this database, a direction of sound arrival is defined (determined) to correspond one-to-one with a head-related transfer function.

[0662] For example, the reproduction unit 2303 can refer to this database, select a head-related transfer function that corresponds one-to-one with the first direction of arrival, apply the selected head-related transfer function to the first sound signal, thereby generating and outputting an output signal. Furthermore, the same processing is performed for the second direction of arrival and the direction of arrival of the combined tone.

[0663] Here, the direction of arrival of the combined tone used in step S505b will be explained. Furthermore, in the explanation of the direction of arrival of the combined tone, the two sounds specified in step S501b will be described as a first sound and a second sound.

[0664] Since a first sound and a second sound were specified in step S501b, the acquisition unit 2302a acquires the first sound signal and the second sound signal.

[0665] As described above, the first attribute information contained in the acquired first sound signal includes first position information indicating the position of the first sound and first volume information indicating the volume of the first sound. Furthermore, the second attribute information contained in the acquired second sound signal includes second position information indicating the position of the second sound and second volume information indicating the volume of the second sound.

[0666] The merging unit 2302c determines the direction of arrival of the merged tone representing the merged sound signal at the listening position based on the first position information, the first volume information, and the second position information and the second volume information. A specific example will be explained using Figure 40.

[0667] Figure 40 is a diagram illustrating the direction of arrival of the combined sound used in Operation Example 3 of this embodiment.

[0668] The first and second position information are information representing the positions of the first and second sounds in polar coordinates. For example, as shown in Figure 40, the first position information indicates that the polar coordinates of the position of the first sound have a deflection angle of 135°, and the second position information indicates that the polar coordinates of the position of the second sound have a deflection angle of 45°. Furthermore, when the first and second position information are represented in polar coordinates, the deflection angle with respect to the sound corresponds to the direction in which the sound arrives at the listener.

[0669] Furthermore, as shown in Figure 40, the first volume information indicates that the volume of the first sound (the volume at the listening position) is 1.0, and the second volume information indicates that the volume of the second sound (the volume at the listening position) is 3.0.

[0670] The merging section 2302c first determines the position of the merged notes as follows.

[0671] The merging unit 2302c considers the volume of the first sound represented by the first volume information as a weight corresponding to the position of the first sound represented by the first position information. The merging unit 2302c considers the volume of the second sound represented by the second volume information as a weight corresponding to the position of the second sound represented by the second position information. In this case, the merging unit 2302c determines that the position of the merged tone is the position of the center of gravity of the first and second sounds.

[0672] That is, the merging unit 2302c determines the position of the merged tone (more specifically, the position where the merged tone is generated) by the position of the center of gravity of the first and second sounds. In Figure 40, the angle (direction of arrival) of the center of gravity position is calculated to be 67.5°. In addition, the volume of the merged tone is the volume of the first sound and the volume of the second sound added together, and is therefore calculated to be 4.0. Thus, the merging unit 2302c determines the position of the merged tone.

[0673] Furthermore, the merging unit 2302c uses the determined position of the merged tone, the position information of the avatar (listener) contained in the input signal, and the orientation information D1 of the avatar (listener) to determine the direction of arrival of the merged tone. That is, the merging unit 2302c determines the direction of arrival of the merged tone from the determined position of the merged tone (the position of the center of gravity of the first sound and the second sound) toward the listening position.

[0674] Thus, in this embodiment, the merging unit 2302c determines the direction of arrival of the merged sound based on the first position information and the first volume information, as well as the second position information and the second volume information.

[0675] In addition, volume was used in the above to determine the position of the center of gravity, but it is not limited to this. Loudness can also be used instead of volume.

[0676] Furthermore, in step S505b, the reproduction unit 2303 outputs an output signal generated by applying a head-related transfer function based on the direction of arrival of the merged sound signal to the merged sound signal.

[0677] Furthermore, in step S510, the decision unit 2302b can select one of the two specified sounds according to the following conditions. For example, if one of the two specified sounds is a direct tone, the decision unit 2302b selects the direct tone. For example, if both specified sounds are direct tones or indirect tones (reflected tones), the decision unit 2302b selects the sound with a large volume, a loud sound, a sound whose direction of arrival is stopped (i.e., a sound from which the sound source is not moving), or a sound whose direction of arrival is close to the listener's front. The decision unit 2302b may also determine the value corresponding to the direction of arrival of the combined sound of the two sounds as the value of the discrimination limit T.

[0678] [Execution Example 4 of the Rendering Department]Execution Example 4 performs the same process as Execution Example 1, except that it performs the process shown in FIG41 instead of the process shown in FIG35.

[0679] Figure 41 is a flowchart illustrating an example of the processing performed by the selection unit 2302 and the reproduction unit 2303 in Operation Example 4 of this embodiment. Furthermore, the description of commonalities with Operation Example 3 of Embodiment 2 is omitted or simplified here.

[0680] In Action Example 4, the steps S101a shown in FIG34 are performed in the same manner as in Action Example 1.

[0681] Furthermore, as shown in FIG41, the selection unit 2302 specifies the two sounds detected by the analysis unit 2301 (S501b).

[0682] The decision unit 2302b of the selection unit 2302 calculates the position where the merged tone is generated when the two specified sounds are merged, and the direction of arrival of the merged tone from the position where it is generated to the listening position (the direction of arrival of the merged tone).

[0683] In addition, the method for calculating the position where the merged sound is generated and the direction of arrival of the merged sound is as described in Action Example 3.

[0684] In addition, the decision unit 2302b calculates the respective directions of arrival of the two sounds specified.

[0685] The decision unit 2302b calculates the arrival angles (arrival angle δ1 and arrival angle δ2) of the combined tone and the two specified sounds when two sounds are combined (S502c). That is, the decision unit 2302b calculates the angle δ1 formed by the arrival direction of the combined tone and the arrival direction of one of the two specified sounds, and calculates the angle δ2 formed by the arrival direction of the combined tone and the arrival direction of the other of the two specified sounds.

[0686] The decision unit 2302b compares the calculated arrival angles δ1 and δ2 with the discrimination limit T of the arrival angles. In this example, the discrimination limit T is determined as follows.

[0687] The memory of the analysis unit 2301 stores a database showing the values ​​of the discrimination limit T corresponding to the direction of arrival of each sound. The decision unit 2302b refers to this database, extracts the value corresponding to the direction of arrival (direction of arrival of the combined sound) from the position where the combined sound is generated to the listening position when two specified sounds are combined (S511), and determines the extracted value as the value of the discrimination limit T.

[0688] The decision unit 2302b determines whether to merge the two sound signals based on indicators corresponding to the calculated arrival directions of the two sounds. That is, the decision unit 2302b determines whether the calculated arrival angle δ1 is less than the discrimination limit T, and whether the calculated arrival angle δ2 is less than the discrimination limit T (S503c). Thus, the decision unit 2302b decides whether to merge the two sound signals.

[0689] When the arrival angle δ1 is less than the discrimination limit T and the arrival angle δ2 is less than the discrimination limit T ("Yes" in S503c), the decision unit 2302b decides to merge the two sound signals. Then, the merging unit 2302c merges the two sounds (S504b).

[0690] If the arrival angle δ1 is greater than or equal to the discrimination limit T or the arrival angle δ2 is greater than or equal to the discrimination limit T ("No" in S503c), the decision unit 2302b decides not to merge the two sound signals.

[0691] The selection unit 2302 determines whether steps S501b to S504b have been processed for all second combinations (S507b).

[0692] If the answer in step S507b is "No", then the process in step S501b is repeated. If the answer in step S507b is "Yes", then the following process is performed.

[0693] The reproduction unit 2303 applies a head-related transfer function (S505b) to both unmerged and merged sounds (merged tones) according to their direction of arrival.

[0694] [Execution Example 5 of the Rendering Department]Execution Example 5 performs the same process as Execution Example 1, except that it performs the process shown in FIG42 instead of the process shown in FIG35.

[0695] Figure 42 is a flowchart illustrating an example of the processing performed by the selection unit 2302 and the reproduction unit 2303 in operation example 5 of this embodiment. Furthermore, the explanation of commonalities with operation examples 1 and 4 of embodiment 2 is omitted or simplified here.

[0696] In Action Example 5, the steps S101a shown in FIG34 are performed in the same manner as in Action Example 1.

[0697] The selection unit 2302 specifies the reflected tone and direct tone detected by the analysis unit 2301 (S501).

[0698] The decision unit 2302b calculates the arrival angle δ of the reflected sound and the direct sound (S502).

[0699] The decision unit 2302b determines whether the arrival angle δ based on the first arrival direction and the second arrival direction is less than the discrimination limit T (S503). The discrimination limit T is the same as the discrimination limit T of the arrival angle δ in Operation Example 1, and can be obtained by the acquisition unit 2302a.

[0700] When the arrival angle δ is less than the discrimination limit T ("Yes" in S503), the merging unit 2302c synthesizes (merges) the reflected tone (first sound) and the direct tone (second sound) (S504). Furthermore, in this case, the merging unit 2302c can perform a process that invalidates the first sound signal and the second sound signal before merging.

[0701] When the arrival angle δ is greater than or equal to the discrimination limit T ("No" in S503), the decision unit 2302b decides not to merge the first sound signal and the second sound signal. The merging unit 2302c outputs the first sound signal and the second sound signal to the reproduction unit 2303.

[0702] Selection unit 2302 determines whether steps S501 to S504 have been processed for all first combinations (S507).

[0703] If the answer in step S507 is "No", then the process in step S501 is repeated. If the answer in step S507 is "Yes", then the following process is performed.

[0704] The reproduction unit 2303 applies a head-related transfer function (S505d) to the direct tone and the non-invalid reflected tone, corresponding to their direction of arrival. More specifically, the reproduction unit 2303 outputs an output signal generated by applying a head-related transfer function corresponding to the direction of arrival of the second sound represented by the second sound signal to the second sound signal representing the direct tone. Furthermore, the reproduction unit 2303 outputs an output signal generated by applying a head-related transfer function corresponding to the direction of arrival of the reflected tone represented by the first sound signal, which has not been invalidated by the merging unit 2302c.

[0705] Similar to Action Example 3, the reproduction unit 2303 can use a database that shows head-related transfer functions corresponding to the direction of sound arrival.

[0706] Furthermore, a reflected sound was used in action example 5, but the same treatment was applied even if an indirect sound was used instead of a reflected sound.

[0707] [Rendering Department Action Example 6] Action Example 6 is the same as Action Example 5 except that it extracts the discrimination limit T as follows.

[0708] Figure 43 is a flowchart illustrating an example of the processing performed by the selection unit 2302 and the reproduction unit 2303 in Operation Example 6 of this embodiment. Furthermore, the explanation of commonalities with Operation Example 5 of Embodiment 2 is omitted or simplified here.

[0709] In Action Example 5, the steps S101a shown in FIG34 are performed in the same manner as in Action Example 1.

[0710] Proceed to steps S501 and S502.

[0711] In this example, similar to example 3, the memory of the analysis unit 2301 stores a database showing the values ​​of the discrimination limit T corresponding to the direction of arrival of each sound. The decision unit 2302b extracts the value corresponding to the direction of arrival of the direct sound by referring to the database (S512), and determines the extracted value as the value of the discrimination limit T.

[0712] Using the discrimination limit T determined as described above, proceed to step S503.

[0713] Proceed to steps S504, S507, and S505d.

[0714] Furthermore, databases are used in steps S512 and S505d respectively. The thresholds shown in the database used in step S512 and the head-related transfer functions shown in the database used in step S505d can be stored in a common address region according to each direction of sound arrival.

[0715] [Execution Example 7 of the Rendering Department]Execution Example 7 performs the same process as Execution Example 1, except that it performs the process shown in FIG44 instead of the process shown in FIG35.

[0716] Figure 44 is a flowchart illustrating an example of the processing performed by the selection unit 2302 and the reproduction unit 2303 in operation example 7 of this embodiment. Furthermore, descriptions of commonalities with operation examples 1, 2, and 6 of embodiment 2 are omitted or simplified here.

[0717] In Action Example 7, the steps S101a shown in FIG34 are performed in the same manner as in Action Example 1.

[0718] The selection unit 2302 specifies the reflected tone (S501aa) detected by the analysis unit 2301. In addition, sometimes the reflected tone specified at this time (the first sound signal) is recorded as the first reflected tone.

[0719] The selection unit 2302 specifies the direct tone detected by the analysis unit 2301 (S501a).

[0720] The decision unit 2302b calculates the arrival angle δ of the reflected sound and the direct sound (S502).

[0721] In this example, similar to example 3, the memory of the analysis unit 2301 stores a database showing the values ​​of the discrimination limit T corresponding to the direction of arrival of each sound. The decision unit 2302b refers to this database, extracts the value corresponding to the direction of arrival of the reflected sound or the value corresponding to the direction of arrival of the direct sound (S513), and determines the extracted value as the value of the discrimination limit T.

[0722] The decision unit 2302b determines whether the arrival angle δ based on the first and second arrival directions is less than the discrimination limit T (S503). In this example, for example, the value corresponding to the arrival direction of the reflected sound is used as the value of the discrimination limit T.

[0723] When the arrival angle δ is less than the discrimination limit T ("Yes" in S503), the decision unit 2302b determines the direct tone (second tone) as the synthesis (merging) object (S514).

[0724] When the arrival angle δ is greater than or equal to the discrimination limit T ("No" in S503), the decision unit 2302b determines that the direct tone (the second sound) is not a synthesis (merging) object. That is, the decision unit 2302b determines that the first sound signal and the second sound signal will not be merged.

[0725] The selection unit 2302 determines whether steps S501a to S514 have been processed on all of the plurality of second sound signals created by the analysis unit 2301 and stored in the memory. That is, the selection unit 2302 determines whether steps S501a to S514 have been processed on all the direct tones represented by the plurality of second sound signals created by the analysis unit 2301 and stored in the memory (S515).

[0726] If the answer in step S515 is "No", then proceed to step S501a again. If the answer in step S515 is "Yes", then proceed to the following steps.

[0727] The merging unit 2302c combines the direct tone determined as the merging target in step S514 with the reflected tone (first reflected tone) specified in step S501aa (S516). That is, the merging unit 2302c generates a merged sound signal by merging the second sound signal representing the direct tone determined as the merging target and the first sound signal representing the first reflected tone. The merging unit 2302c outputs the generated merged sound signal to the reproduction unit 2303. Furthermore, in this case, the merging unit 2302c can perform a process that invalidates the first sound signal before merging.

[0728] Furthermore, when there are multiple direct tones determined to be merge targets, the merging unit 2302c may perform the following processing, for example.

[0729] The merging unit 2302c can also merge each of the multiple direct tones with the first reflected tone to generate the same number of merged sound signals as the multiple direct tones. Additionally, the merging unit 2302c can also merge the direct tone closest to the first reflected tone with the first reflected tone. The merging unit 2302c can also merge the direct tone with the loudest volume with the first reflected tone. The merging unit 2302c can also merge the direct tone whose sound production location is the first reflected tone with the first reflected tone. The merging unit 2302c can also merge the direct tone whose sound production location is most easily seen by the listener with the first reflected tone.

[0730] The selection unit 2302 determines whether steps S501aa to S516 have been processed on all of the plurality of first sound signals generated by the analysis unit 2301 and stored in the memory. That is, the selection unit 2302 determines whether steps S501aa to S516 have been processed on all the reflected sounds represented by the plurality of first sound signals generated by the analysis unit 2301 and stored in the memory (S517).

[0731] If "No" is received in step S517, the process of step S501aa is performed again. Furthermore, if the process of step S501aa is performed again, the reflected tone is specified again, and the specified reflected tone (the first sound signal) becomes the second reflected tone.

[0732] Subsequently, after further performing step S516, the merging unit 2302c merges the direct tone that was determined to be the merging target in S514 with the reflected tone (the second reflected tone) that was specified again in step S501aa. The processing is the same for the third and subsequent steps.

[0733] If "yes" is selected in step S517, proceed to step S505d.

[0734] Furthermore, a reflected sound was used in action example 7, but the same treatment was applied even if an indirect sound was used instead of a reflected sound.

[0735] In addition, in action examples 3, 4, 6 and 7, the discrimination limit T can also be changed according to the following.

[0736] For example, the discrimination limit T can also be changed according to the amount of movement of the listener, that is, the amount of rotation of the listener's head.

[0737] Even with high-speed head rotation, the perception of the listener is almost unaffected compared to when the listener is stationary, for example, even if the first and second sounds are combined. Therefore, when the listener is highly mobile, the discrimination limit T is set (determined) to be larger than when the listener is stationary, thereby reducing the processing load without impairing the listener's listening experience. The listener's motion can be obtained by sensors integrated into the audio signal processing device, or by the audio signal processing device from external sensors.

[0738] Additionally, for example, the discrimination limit T can also be changed depending on whether the image is displayed or not.

[0739] When visual cues are presented to the listener simultaneously with the audio output, even if the perceived direction of the sound is not precise, the visual perception of the direction of arrival can still follow. Therefore, with visual cues, the discrimination limit T is set (determined) to be larger than without cues, thereby reducing the processing load without compromising the listener's listening experience.

[0740] The following is a summary of this implementation method.

[0741] The sound signal processing method of this embodiment is a sound signal processing method executed by a sound signal processing device. The sound signal processing method includes an acquisition step, a decision step, a merging step, and a reproduction step.

[0742] In the acquisition step, a first sound signal and a second sound signal are obtained. The first sound signal represents a first sound and includes first attribute information for determining the attributes of the first sound signal. The second sound signal represents a second sound and includes second attribute information for determining the attributes of the second sound signal. In the decision step, based on indicators corresponding to the first direction of arrival of the first sound at the listener's location (i.e., the listening location) and the second direction of arrival of the second sound at the listening location, a decision is made on whether to merge the acquired first sound signal and the acquired second sound signal.

[0743] In the merging step, if it is determined that the first and second audio signals will be merged, a merged audio signal is generated by combining the first and second audio signals. In the reproduction step, an output signal based on the generated merged audio signal is output.

[0744] Therefore, based on the indices corresponding to the first arrival direction of the first sound and the second arrival direction of the second sound, it is determined whether to merge the first and second sound signals. If it is determined that the first and second sound signals should be merged, an output signal based on the merged sound signal is output in the reproduction step. If it is assumed that the first and second sound signals should not be merged, an output signal based on the first sound signal and an output signal based on the second sound signal are output in the reproduction step. Compared to the case where only one output signal based on the first sound signal and one output signal based on the second sound signal are output, the number of output signals is reduced when outputting an output signal based on the merged sound signal, thus reducing the computational load. In other words, a sound signal processing method that reduces computational load can be implemented.

[0745] Furthermore, as an example, the indicator is the angle of arrival δ. When the angle of arrival δ is less than the discrimination limit T, the decision step determines to merge the first sound signal and the second sound signal. The case where the angle of arrival δ is less than the discrimination limit T corresponds to the case where the first sound and the second sound arrive at the listener from the same direction or approximately the same direction. In this case, the process of integrating the first sound (reflected sound) and the second sound (direct sound), that is, merging the first sound signal and the second sound signal to generate a merged sound signal, is performed. Even with this process, the listener rarely experiences any dissonance. In addition, as mentioned above, since the number of output signals is reduced, it is possible to realize a sound signal processing method that can more appropriately reduce the computational load.

[0746] In the sound signal processing method of this embodiment, the index is an index composed of two orthogonal axes.

[0747] Therefore, a simpler index can be used, that is, compared with the use of complex indexes, it does not require a large amount of computation and a large computational load, thus enabling a sound signal processing method that can reduce the amount of computation and the computational load.

[0748] In the sound signal processing method of this embodiment, the first attribute information contained in the acquired first sound signal includes first position information indicating the position of the first sound and first volume information indicating the volume of the first sound. The second attribute information contained in the acquired second sound signal includes second position information indicating the position of the second sound and second volume information indicating the volume of the second sound. In the merging step, based on the first position information and the first volume information, and the second position information and the second volume information, the direction of arrival of the merged sound representing the merged sound signal at the listening position is determined.

[0749] Therefore, in the merging step, the direction of arrival of the merged tone can be determined based on the first position information, the first volume information, the second position information, and the second volume information.

[0750] In the sound signal processing method of this embodiment, in the merging step, the volume of the first sound represented by the first volume information is regarded as the weight corresponding to the position of the first sound represented by the first position information. Furthermore, in the merging step, when the volume of the second sound represented by the second volume information is regarded as the weight corresponding to the position of the second sound represented by the second position information, the following processing is performed. In the merging step, the position of the merged sound is determined to be the position of the center of gravity of the first sound and the second sound, and the direction from the determined center of gravity position toward the listening position is determined as the direction of arrival of the merged sound.

[0751] Therefore, the direction of arrival of the merged tone can be determined based on the position of the center of gravity of the first and second sounds, thus making the direction of arrival of the merged tone easy to determine. In other words, determining the direction of arrival of the merged tone does not require a large amount of computation or a large computational load, thereby enabling a sound signal processing method that reduces the amount of computation and computational load.

[0752] The sound signal processing method of this embodiment is a sound signal processing method executed by a sound signal processing device. The sound signal processing method includes an acquisition step, a merging step, and a reproduction step.

[0753] In the acquisition step, a first sound signal and a second sound signal are acquired. The first sound signal represents a first sound and includes first attribute information for determining the attributes of the first sound signal. The second sound signal represents a second sound and includes second attribute information for determining the attributes of the second sound signal. In the merging step, if it is determined that the acquired first sound signal and the acquired second sound signal should be merged, a merged sound signal is generated by merging the first sound signal and the second sound signal. In the reproduction step, an output signal based on the generated merged sound signal is output.

[0754] The first attribute information contained in the acquired first sound signal includes first position information indicating the location of the first sound and first volume information indicating the volume of the first sound. The second attribute information contained in the acquired second sound signal includes second position information indicating the location of the second sound and second volume information indicating the volume of the second sound. In the merging step, based on the first position information and the first volume information, and the second position information and the second volume information, the direction of arrival of the merged sound represented by the merged sound signal at the listener's location, i.e., the listening location, is determined.

[0755] Therefore, if it is decided to merge the first and second audio signals, the reproduction step outputs an output signal based on the merged audio signal. If it is assumed that the first and second audio signals are not merged, the reproduction step outputs an output signal based on the first audio signal and an output signal based on the second audio signal. Compared to outputting both an output signal based on the first and second audio signals, outputting an output signal based on the merged audio signal reduces the number of output signals, thus reducing computational complexity and workload. In other words, an audio signal processing method that reduces computational complexity and workload can be implemented.

[0756] In the sound signal processing method of this embodiment, when it is determined that the first sound signal and the second sound signal should be merged, in the reproduction step, an output signal is output by applying a head-related transfer function based on the determined direction of arrival of the merged sound signal to the generated merged sound signal.

[0757] This enables a sound signal processing method that allows listeners to hear a more immersive, combined sound.

[0758] In the audio signal processing method of this embodiment, if it is determined that the first audio signal and the second audio signal will not be merged, in the reproduction step, an output signal generated by applying a head-related transfer function based on the first arrival direction to the first audio signal is output. Furthermore, in the reproduction step, an output signal generated by applying a head-related transfer function based on the second arrival direction to the second audio signal is output.

[0759] This enables a sound signal processing method that allows listeners to hear a more immersive first and second sound.

[0760] The computer program in this embodiment is a computer program used to enable a computer to execute the above-described sound signal processing method.

[0761] Therefore, the computer can execute the above-mentioned sound signal processing method according to the computer program.

[0762] The audio signal processing apparatus of this embodiment includes an acquisition unit 2302a, a decision unit 2302b, a merging unit 2302c, and a reproduction unit 2303.

[0763] The acquisition unit 2302a acquires a first sound signal and a second sound signal. The first sound signal represents a first sound and includes first attribute information for determining the attributes of the first sound signal. The second sound signal represents a second sound and includes second attribute information for determining the attributes of the second sound signal. The decision unit 2302b determines whether to merge the acquired first sound signal and the acquired second sound signal based on indicators corresponding to a first direction of arrival of the first sound at the listener's location (i.e., the listening location) and a second direction of arrival of the second sound at the listening location.

[0764] If it is determined that the first and second audio signals should be merged, the merging unit 2302c generates a merged audio signal after merging the first and second audio signals. The reproduction unit 2303 outputs an output signal based on the generated merged audio signal.

[0765] Therefore, based on indicators corresponding to the first arrival direction of the first sound and the second arrival direction of the second sound, it is determined whether to merge the first sound signal and the second sound signal. If it is determined that the first sound signal and the second sound signal should be merged, the reproduction unit 2303 outputs an output signal based on the merged sound signal. If it is assumed that the first sound signal and the second sound signal should not be merged, the reproduction unit 2303 outputs an output signal based on the first sound signal and an output signal based on the second sound signal. Compared to the case where both an output signal based on the first sound signal and an output signal based on the second sound signal are output, the number of output signals is reduced when an output signal based on the merged sound signal is output, thus reducing the computational load and computational complexity. In other words, a sound signal processing apparatus capable of reducing computational load and computational complexity can be realized.

[0766] Furthermore, as an example, the indicator is the angle of arrival δ. When the angle of arrival δ is less than the discrimination limit T, the decision unit 2302b decides to merge the first sound signal and the second sound signal. The case where the angle of arrival δ is less than the discrimination limit T corresponds to the case where the first sound and the second sound arrive at the listener from the same direction or approximately the same direction. In this case, processing is performed to integrate the first sound (reflected sound) and the second sound (direct sound), that is, to merge the first sound signal and the second sound signal to generate a merged sound signal. Even with this processing, the listener rarely experiences any dissonance. In addition, as described above, since the number of output signals is reduced, it is possible to realize a sound signal processing device that can more appropriately reduce the computational load.

[0767] (Embodiment 3) Hereinafter, Embodiment 3 will be described. The description will focus on the differences from Embodiments 1 and 2, and the description of the commonalities will be omitted or simplified.

[0768] [Configuration of the rendering unit] First, the configuration of the rendering unit 3300 in this embodiment will be described. FIG45 is a block diagram showing an example of the configuration of the rendering unit 3300 in this embodiment.

[0769] The rendering unit 3300 includes a resolution unit 3301, a selection unit 3302, and a reproduction unit 3303. Furthermore, the audio signal processing apparatus of this embodiment is an example of a decoding apparatus, which includes a decoder, and the decoder includes the rendering unit 3300. That is, it can be said that the audio signal processing apparatus of this embodiment includes a resolution unit 3301, a selection unit 3302, and a reproduction unit 3303. The rendering unit 3300 performs additional audio processing on the audio data contained in the input signal and outputs it.

[0770] In embodiment 2, a first sound signal and a second sound signal are used, that is, a first sound (e.g., a reflected sound) and a second sound (e.g., a direct sound). However, this embodiment is not limited to this.

[0771] In this embodiment, instead of the first and second sound signals, a sound signal representing a predetermined tone and containing attribute information for determining the attributes of the sound signal is used. That is, the sound signal contains attribute information, which is information used to determine the attributes of the sound signal. This attribute may include information representing the predetermined tone. The predetermined tone is either an indirect tone (reflected tone) or a direct tone as described in embodiments 1 and 2.

[0772] In addition, for simplicity, sometimes a sound signal whose attribute represents information about a specified tone is recorded as a sound signal representing a specified tone.

[0773] The following describes the components of the rendering unit 3300.

[0774] The analysis unit 3301 performs analytical processing on the input signal. Specifically, it is described below.

[0775] The analysis unit 3301 can perform all or part of the processing performed by the analysis unit 1301 of Embodiment 1 and the analysis unit 2301 of Embodiment 2.

[0776] Furthermore, the analysis unit 3301 performs analysis processing of the input signal as in S101 of Embodiment 1, as shown in FIG8. That is, the analysis unit 3301 analyzes the input signal input to the sound signal processing device of this embodiment and detects the specified tone that may be generated in the sound space.

[0777] Upon detecting such a prescribed tone, the analysis unit 3301 generates a sound signal representing the prescribed tone based on spatial information and sound data. The method by which the analysis unit 3301 generates the sound signal is the same as the method by which the analysis unit 2301 generates the first sound signal and the second sound signal in Embodiment 2.

[0778] The analysis unit 3301 can store the generated sound signals in its own memory. In this embodiment, the analysis unit 3301 detects M predetermined tones that may be generated in the sound space, generates M sound signals, and stores them in the memory. M is an integer greater than or equal to 2.

[0779] Furthermore, the M sound signals can all be direct tones or all be reflected tones. Alternatively, some of the M sound signals can be direct tones, while the rest can be reflected tones.

[0780] In addition, similar to embodiments 1 and 2, the analysis unit 2301 can calculate values ​​related to the path to the listening position, the time taken to reach the position, and the volume at the time of arrival for each of the M specified tones.

[0781] For example, if a specified tone is a direct tone, the volume of that specified tone when it arrives (the volume of the specified tone) is the volume of the direct tone when it arrives (ld). For example, if another specified tone is a reflected tone, the volume of that other specified tone when it arrives (the volume of the specified tone) is the volume of the reflected tone when it arrives (lr).

[0782] The attribute information contained in a sound signal includes specified tone volume information representing the volume of a specified tone represented by that sound signal (more specifically, the volume at which the specified tone arrives).

[0783] In addition, the analysis unit 3301 generates specified sound position information based on the position information of the sound source object, the position information of the non-sound-producing object (obstacle object), the position information and physical information of the listener, and the sound data contained in the spatial information.

[0784] The position information of a specified tone indicates the position of the specified tone, which is the information used to locate the sound image of the specified tone.

[0785] In this embodiment, the attribute information contained in a sound signal includes specified tone position information indicating the position of a specified tone represented by the sound signal.

[0786] The selection unit 3302 can perform all or part of the processing performed by the selection unit 1302 in Embodiment 1 and the selection unit 2302 in Embodiment 2. Furthermore, the selection unit 3302 determines whether the reproduction unit 2303 outputs (reproduces) an output signal based on the M sound signals generated by the analysis unit 2301. That is, the selection unit 3302 first determines whether to merge N sound signals from the M sound signals generated by the analysis unit 2301, and gener...

Claims

1. A sound signal processing method, executed by a sound signal processing device, wherein, include: The process includes: an acquisition step, acquiring a first sound signal and a second sound signal, wherein the first sound signal represents a first sound and includes first attribute information for determining the attributes of the first sound signal, and the second sound signal represents a second sound and includes second attribute information for determining the attributes of the second sound signal; a decision step, based on indices corresponding to a first arrival direction and a second arrival direction, deciding whether to merge the acquired first sound signal and the acquired second sound signal, wherein the first arrival direction is the direction in which the first sound reaches the listener's location (i.e., the listening position), and the second arrival direction is the direction in which the second sound reaches the listening position; and a merging step, if it is decided to merge the first sound signal and the second sound signal, generating a merged sound signal obtained by merging the first sound signal and the second sound signal. And a reproduction step, outputting an output signal based on the generated merged sound signal.

2. The sound signal processing method according to claim 1, wherein, The indicator is composed of two orthogonal axes.

3. The sound signal processing method according to claim 1 or 2, wherein, The first attribute information contained in the first sound signal obtained includes first position information indicating the location of the first sound and first volume information indicating the volume of the first sound. The second attribute information contained in the second sound signal obtained includes second position information indicating the location of the second sound and second volume information indicating the volume of the second sound. In the merging step, based on the first position information and the first volume information, and the second position information and the second volume information, the direction of arrival of the merged sound represented by the merged sound signal at the listening position is determined.

4. The sound signal processing method according to claim 3, wherein, In the merging step, when the volume of the first sound represented by the first volume information is regarded as the weight corresponding to the position of the first sound represented by the first position information, and the volume of the second sound represented by the second volume information is regarded as the weight corresponding to the position of the second sound represented by the second position information, the position of the merged sound is determined to be the position of the center of gravity of the first sound and the second sound, and the direction from the determined position of the center of gravity toward the listening position is determined as the direction of arrival of the merged sound.

5. A sound signal processing method, executed by a sound signal processing device, wherein, include: The process includes: an acquisition step, acquiring a first sound signal and a second sound signal, wherein the first sound signal represents a first sound and includes first attribute information for determining the attributes of the first sound signal, and the second sound signal represents a second sound and includes second attribute information for determining the attributes of the second sound signal; a merging step, whereby, if it is determined that the acquired first sound signal and the acquired second sound signal are to be merged, generating a merged sound signal obtained by merging the first sound signal and the second sound signal; and a reproduction step, outputting an output signal based on the generated merged sound signal. The first attribute information contained in the first sound signal includes first position information indicating the location of the first sound and first volume information indicating the volume of the first sound. The second attribute information contained in the obtained second sound signal includes second position information indicating the location of the second sound and second volume information indicating the volume of the second sound. In the merging step, based on the first position information and the first volume information, and the second position information and the second volume information, the direction of arrival of the merged sound represented by the merged sound signal at the listener's location, i.e., the listening location, is determined.

6. The sound signal processing method according to any one of claims 3 to 5, wherein, If it is determined that the first sound signal and the second sound signal should be merged, in the reproduction step, the output signal is generated by applying a head-related transfer function based on the determined direction of arrival of the merged sound to the generated merged sound signal.

7. The sound signal processing method according to any one of claims 1 to 4, wherein, If it is decided not to merge the first sound signal and the second sound signal, in the reproduction step, the output signal generated by applying the head-related transfer function based on the first arrival direction to the first sound signal and the output signal generated by applying the head-related transfer function based on the second arrival direction to the second sound signal are output.

8. A sound signal processing method, executed by a sound signal processing device, wherein, include: The process includes: an acquisition step, acquiring M sound signals, each sound signal representing a specified tone and containing attribute information for determining the attributes of the sound signals, where M is an integer greater than or equal to 2; a decision step, based on the direction of arrival of the M specified tones at the listener's location, determining whether to merge N sound signals from the acquired M sound signals, where N is an integer greater than or equal to 1 and less than or equal to M; a merging step, if the decision is to merge N sound signals, generating a merged sound signal obtained by merging the N sound signals; and a reproduction step, outputting an output signal based on the generated merged sound signal.

9. The sound signal processing method according to claim 8, wherein, The attribute information contained in each of the M sound signals obtained includes specified tone position information representing the position of the specified tone and specified tone volume information representing the volume of the specified tone. In the merging step, based on the N specified tone position information and the N specified tone volume information, the direction of arrival of the merged tone represented by the merged sound signal to the listening position is determined.

10. The sound signal processing method according to claim 9, wherein, In the merging step, when the volume of the specified tone represented by the specified tone volume information is regarded as the weight corresponding to the position of the specified tone represented by the specified tone position information, the position of the merged tone is determined to be the position of the center of gravity of N specified tones, and the direction from the determined position of the center of gravity toward the listening position is determined as the direction of arrival of the merged tone.

11. A sound signal processing method, executed by a sound signal processing device, wherein, include: The acquisition step involves acquiring a first sound signal and a second sound signal, wherein the first sound signal represents a first sound and includes first attribute information for determining the attributes of the first sound signal, and the second sound signal represents a second sound and includes second attribute information for determining the attributes of the second sound signal; the decision step involves determining whether to merge the acquired first sound signal and the acquired second sound signal based on a first direction of arrival of the first sound to the listener's location (i.e., the listening location) and a second direction of arrival of the second sound to the listening location; the merging step involves generating a merged sound signal obtained by merging the first sound signal and the second sound signal if it is determined that the first sound signal and the second sound signal should be merged. The process includes a reproduction step, outputting an output signal based on the generated combined sound signal; in the decision step, among multiple arrival direction information representing the direction of arrival of sound reaching the listening position and a head-related transfer function based on the arrival direction, determining one arrival direction information corresponding to the first arrival direction and one arrival direction information corresponding to the second arrival direction; in the merging step, if the determined arrival direction information corresponding to the first arrival direction and the determined arrival direction information corresponding to the second arrival direction are the same, generating the combined sound signal; and in the reproduction step, outputting an output signal based on the generated combined sound signal. The output signal generated by applying the head-related transfer function represented by the arrival direction information corresponding to the first arrival direction to the sound signal, and the output signal generated by applying the head-related transfer function represented by the arrival direction information corresponding to the first arrival direction to the sound signal, and the output signal generated by applying the head-related transfer function represented by the arrival direction information corresponding to the second arrival direction to the sound signal, when the determined arrival direction information corresponding to the first arrival direction and the determined arrival direction information corresponding to the second arrival direction are different.

12. A computer program for causing a computer to execute the sound signal processing method according to any one of claims 1 to 11.

13. A sound signal processing device, wherein, The device comprises: an acquisition unit that acquires a first sound signal and a second sound signal, wherein the first sound signal represents a first sound and includes first attribute information for determining attributes of the first sound signal, and the second sound signal represents a second sound and includes second attribute information for determining attributes of the second sound signal; a determination unit that determines whether to merge the acquired first sound signal and the acquired second sound signal based on indicators corresponding to a first arrival direction and a second arrival direction, wherein the first arrival direction is the direction in which the first sound arrives at the listener's location, i.e., the listening position, and the second arrival direction is the direction in which the second sound arrives at the listening position; a merging unit that, if it is determined that the first sound signal and the second sound signal should be merged, generates a merged sound signal obtained by merging the first sound signal and the second sound signal; and a reproduction unit that outputs an output signal based on the generated merged sound signal.

Citation Information

Patent Citations

  • Signal processor

    JP2019022049A

  • Apparatus and method for rendering a sound scene using pipeline stages

    WO2021180938A1