Audio signal processing method, computer program, and audio signal processing device
By calculating factors such as time difference and reflective properties to determine the reflected sound signal, the problem of excessive computational load in virtual or real space is solved, and the computational load is appropriately reduced, extending battery life and optimizing the use of computing resources is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
- Filing Date
- 2024-10-04
- Publication Date
- 2026-04-24
AI Technical Summary
Existing technologies struggle to adequately reduce the computational load and complexity of sound signal processing in virtual or real spaces, especially in environments with multiple sound sources and complex conditions, leading to shortened battery life and wasted computing resources.
By calculating the time difference and factors such as the nature and volume of the reflector, the system determines whether to select and output the reflected sound signal, reducing unnecessary computational processing.
It effectively reduces the computational load and computational burden of sound signal processing in virtual or real space, extends battery life, and optimizes the use of computing resources.
Smart Images

Figure CN121925868A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to methods for processing sound signals, etc. Background Technology
[0002] In recent years, goods and services utilizing ER (Extended Reality), including VR (Virtual Reality), AR (Augmented Reality), and MR (Mixed Reality), have become increasingly popular. Consequently, the importance of sound signal processing technologies—which assign acoustic effects corresponding to the environment of a virtual sound source to the sound emitted in virtual or real space, thereby providing listeners with immersive audio—has increased.
[0003] Additionally, the listener can also be a listener or a user. Furthermore, technologies related to the sound signal processing method disclosed herein are shown in Patent Document 1, Patent Document 2, Patent Document 3, and Non-Patent Document 1.
[0004] Existing technical documents Patent documents Patent Document 1: Japanese Patent No. 6288100 Patent Document 2: Japanese Patent Application Publication No. 2019-22049 Patent Document 3: International Publication No. 2021 / 180938 Non-patent literature Non-Patent Literature 1: BCJ Moore, *An Outline of Auditory Psychology*, Chengxin Bookstore, April 20, 1994, Chapter 6: Spatial Perception, p. 225 Summary of the Invention
[0005] The problem that the invention aims to solve However, in the technology shown in Patent Document 1, it is sometimes difficult to properly reduce the amount of computation and the computational load.
[0006] Therefore, the purpose of this disclosure is to provide a sound signal processing method that can appropriately reduce computational load and computational complexity.
[0007] Methods used to solve problems One aspect of this disclosure is a sound signal processing method executed by a sound signal processing device, comprising: a parsing step, calculating a time difference, the time difference being the time difference between the moment when a sound emitted from a sound source as a direct tone arrives at the listener's location (i.e., the sound source position) from the moment the sound is reflected by a reflector and arrives at the listener's location (i.e., the listener's position) as a reflected tone; an acquisition step, acquiring the calculated time difference and a sound signal representing the reflected tone; a determination step, based on the acquired time difference, the nature of the object present in the sound propagation, a first volume, and a second volume, determining whether to select the reflected tone represented by the acquired sound signal, wherein the first volume is a volume corresponding to the direction from the sound source position to the listener's location, and the second volume is a volume corresponding to the direction from the sound source position to the reflector's location (i.e., the reflector position); and a reproduction step, outputting an output signal based on a sound signal representing the selected reflected tone.
[0008] In addition, a computer program of one embodiment of this disclosure enables a computer to perform the aforementioned sound signal processing method.
[0009] Furthermore, one aspect of the sound signal processing apparatus disclosed herein includes: a resolution unit that calculates a time difference, which is the time difference between the moment when a sound emitted from a sound source as a direct sound travels from the location of the sound source (i.e., the sound source position) to the moment when the sound is reflected by a reflector and becomes a reflected sound, which then travels to the listening position; an acquisition unit that acquires the calculated time difference and a sound signal representing the reflected sound; a determination unit that, based on the acquired time difference, the nature of an object present in the propagation of the sound, a first volume, and a second volume, determines whether to select the reflected sound represented by the acquired sound signal, wherein the first volume is a volume corresponding to the direction from the sound source position to the listening position, and the second volume is a volume corresponding to the direction from the sound source position to the location of the reflector (i.e., the reflector position); and a reproduction unit that outputs an output signal based on a sound signal representing the reflected sound determined to be selected.
[0010] Furthermore, these inclusive or specific solutions can also be implemented by non-transitory recording media such as systems, devices, methods, integrated circuits, computer programs, or computer-readable CD-ROMs, or by any combination of systems, devices, methods, integrated circuits, computer programs, and recording media.
[0011] Invention Effects According to one of the solutions disclosed herein, the sound signal processing method, etc., can appropriately reduce the amount of computation and computational load. Attached Figure Description
[0012] Figure 1 This is a diagram illustrating an example of direct and reflected tones generated in sound space.
[0013] Figure 2 This is a diagram illustrating an example of the stereo sound reproduction system of Embodiment 1.
[0014] Figure 3A This is a block diagram illustrating an example of the configuration of the encoding device in Embodiment 1.
[0015] Figure 3B This is a block diagram illustrating an example of the configuration of the decoding device in Embodiment 1.
[0016] Figure 3C This is a block diagram illustrating another configuration example of the encoding device according to Embodiment 1.
[0017] Figure 3D This is a block diagram illustrating another configuration example of the decoding device according to Embodiment 1.
[0018] Figure 4A This is a block diagram illustrating a configuration example of the decoder in Implementation Method 1.
[0019] Figure 4B This is a block diagram illustrating another configuration example of the decoder in Implementation 1.
[0020] Figure 5 This is a diagram illustrating an example of the physical configuration of the sound signal processing device in Embodiment 1.
[0021] Figure 6 This is a diagram illustrating an example of the physical configuration of the encoding device according to Embodiment 1.
[0022] Figure 7 This is a block diagram showing an example of the configuration of the rendering unit in Embodiment 1.
[0023] Figure 8 This is a flowchart illustrating an example of the operation of the sound signal processing device in Embodiment 1.
[0024] Figure 9 It is a diagram that shows the relative positions of the listener and the obstacle.
[0025] Figure 10 It is a diagram that shows the relative positional relationship between the listener and the obstacle object.
[0026] Figure 11 It is a graph showing the relationship between the time difference and threshold of direct and reflected tones.
[0027] Figure 12A This is a diagram that represents part of an example of how threshold data is set.
[0028] Figure 12B This is a diagram that represents part of an example of how threshold data is set.
[0029] Figure 12C This is a diagram that represents part of an example of how threshold data is set.
[0030] Figure 13 This is a diagram illustrating an example of how a threshold is set.
[0031] Figure 14 This is a flowchart representing an example of a selection process.
[0032] Figure 15 It is a graph showing the relationship between the direction of the direct tone, the direction of the reflected tone, the time difference, and the threshold.
[0033] Figure 16 It is a graph showing the relationship between angle difference, time difference, and threshold.
[0034] Figure 17 This is a block diagram representing another example of the rendering unit.
[0035] Figure 18 This is a flowchart representing another example of the selection process.
[0036] Figure 19 This is another flowchart representing a selection process.
[0037] Figure 20 This is a flowchart illustrating a first variation of the operation of the sound signal processing apparatus according to Embodiment 1.
[0038] Figure 21 This is a flowchart illustrating a second variation of the operation of the sound signal processing device in Embodiment 1.
[0039] Figure 22 This is a diagram showing a configuration example of avatars, sound source objects, and obstacle objects.
[0040] Figure 23 This is another flowchart representing a selection process.
[0041] Figure 24 This is a block diagram illustrating a configuration example for pipeline processing in the rendering unit.
[0042] Figure 25 It is a diagram showing the transmission and diffraction of sound.
[0043] Figure 26 This is a diagram illustrating an example of the positional relationship between the listener and the obstacle object in Implementation 1.
[0044] Figure 27This is a diagram illustrating another example of the positional relationship between the listener and the obstacle object in Implementation 1.
[0045] Figure 28 This is an example of the echo detection limit threshold in Implementation Method 1.
[0046] Figure 29 This is a block diagram showing an example of the configuration of the rendering unit in Embodiment 2.
[0047] Figure 30 This is a diagram used to illustrate the first sound, the second sound, the direct sound, and the reflected sound in Embodiment 2.
[0048] Figure 31 This is a diagram showing the directionality of sound in Implementation Method 2.
[0049] Figure 32A This is a diagram used to illustrate the relationship between the directivity of the sound in Embodiment 2 and the first sound, the second sound, the direct sound, and the reflected sound.
[0050] Figure 32B This is a diagram illustrating an example of the method for calculating the second volume in Embodiment 2.
[0051] Figure 33 This is a flowchart illustrating an example of the operation of the sound signal processing device in Embodiment 2.
[0052] Figure 34 This is a flowchart illustrating the first example of the parsing and selection processes in Implementation Method 2.
[0053] Figure 35 This is a diagram illustrating the process of step S314 in embodiment 2.
[0054] Figure 36 This is a diagram illustrating the process of step S322 in embodiment 2.
[0055] Figure 37 This is a diagram illustrating the process of step S332 in embodiment 2.
[0056] Figure 38 This is a flowchart illustrating the second example of the parsing and selection processes in Implementation Method 2.
[0057] Figure 39 This is a flowchart illustrating the third example of the analysis and selection process in Implementation Method 2.
[0058] Figure 40 This is a flowchart illustrating the fourth example of the parsing and selection processes in Implementation Method 2.
[0059] Figure 41This is a flowchart illustrating the fifth example of the parsing and selection processes in Implementation Method 2.
[0060] Figure 42 This is a flowchart illustrating the sixth example of the parsing and selection processes in Implementation Method 2.
[0061] Figure 43 This is a flowchart illustrating the seventh example of the parsing and selection processes in Implementation Method 2.
[0062] Figure 44 This is a block diagram illustrating the configuration of another example of the rendering unit. Detailed Implementation
[0063] (The insights that form the basis of this disclosure) Previous studies have explored sound signal processing techniques that imbue virtual sound sources with acoustic effects based on the environment of the space in virtual or real spaces and provide immersive audio to listeners.
[0064] Such a sound signal processing technique is disclosed in Patent Document 1. More specifically, Patent Document 1 discloses a technique for detecting the importance of an audio signal (sound signal) and not outputting audio signals with low detected importance. In this way, by not outputting audio signals with low importance, it is expected that the computational load and computational complexity will be appropriately reduced in this sound signal processing technique.
[0065] Additionally, reflected sound sometimes becomes important in sound space (virtual or real space).
[0066] Figure 1 This diagram illustrates an example of direct and reflected sounds generated in a sound space. In sound processing that uses sound to represent the characteristics of a virtual space, in order to represent the breadth of the space and the material of the walls, as well as to accurately determine the location of the sound source (sound image localization), it is effective to reproduce not only direct sounds but also reflected sounds.
[0067] For example, in such Figure 1 When listening to sound within a rectangular room, a sound source produces six primary reflections corresponding to the six walls. The reproduction of these reflections provides clues for a proper understanding of the space and the sound image. Furthermore, for each reflection, secondary reflections are produced on surfaces other than the one that produced the reflection. These reflections also serve as perceptually valid clues.
[0068] However, even considering only the case of secondary reflection, a sound source will produce 1 direct tone and 36 (6+6×5) reflected tones, resulting in 37 vocal lines. Processing these vocal lines requires a considerable amount of computation.
[0069] Furthermore, in recent years, applications related to the concept of the Metaverse, such as virtual meetings, virtual shopping, or virtual concerts, have inevitably involved multiple sound sources, thus requiring a much larger amount of computation.
[0070] Furthermore, listeners in virtual spaces use headphones or VR goggles to receive audio. To provide stereo sound to such listeners, binaural processing is performed on each sound ray, assigning sound pressure levels and phase differences between the two ears to reproduce the direction of arrival and the sense of distance. Therefore, the computational load is extremely high if all the reflected sounds are to be reproduced.
[0071] On the other hand, small rechargeable batteries are sometimes used as batteries for VR goggles worn by listeners experiencing virtual space, due to their convenience. To extend battery life, the computational load required for the processing described above should ideally be low. Therefore, it is desirable to reduce the number of sound lines generated on a scale of several hundred without compromising sound localization and spatial accuracy.
[0072] Furthermore, in systems that reproduce sound, there are sometimes 6 degrees of freedom (DoF) allowed for the listener's position (i.e., the listening position as the listener's location) and orientation. In such cases, the positional relationship between the listener, the sound source, and the object reflecting the sound cannot be determined if it is not during reproduction (rendering). Therefore, the reflected sound also cannot be determined if it is not during reproduction. Consequently, it is difficult to predetermine the reflected sound of the object being processed.
[0073] Therefore, the appropriate selection and output (reproduction) of one or more reflected sounds from multiple reflected sounds generated in the sound space during reproduction is beneficial for the appropriate reduction of computational load.
[0074] Furthermore, controlling whether to select a sound corresponds to determining whether to select a sound, and more specifically, to determining whether to select and output (reproduce) a sound. Additionally, selecting a sound can mean selecting it as the target sound for processing, or selecting it as a non-target sound.
[0075] Furthermore, while Patent Document 1 detects the importance of an audio signal, more specifically the importance of the direct tone represented by that audio signal, it does not investigate the importance of reflected tones. Therefore, in cases such as Figure 1 When indirect sounds such as reflected sounds are generated, the computational load and computational complexity will increase, making it sometimes difficult to appropriately reduce the computational load and computational complexity.
[0076] Therefore, there is a need for sound signal processing methods that can appropriately reduce computational load in the sound space.
[0077] Therefore, the sound signal processing method of the first embodiment of this disclosure is a sound signal processing method executed by a sound signal processing device, comprising: a parsing step, calculating a time difference, the time difference being the time difference between the moment when a sound emitted from a sound source as a direct tone arrives at the listener's location (i.e., the sound source position) from the moment the sound is reflected by a reflector and arrives at the listener's location (i.e., the listener's position) as a reflected tone; an acquisition step, acquiring the calculated time difference and a sound signal representing the reflected tone; a determination step, based on the acquired time difference, the nature of the object present in the sound propagation, a first volume, and a second volume, determining whether to select the reflected tone represented by the acquired sound signal, wherein the first volume is a volume corresponding to the direction from the sound source position to the listener's location, and the second volume is a volume corresponding to the direction from the sound source position to the reflector's location (i.e., the reflector's position); and a reproduction step, outputting an output signal based on a sound signal representing the selected reflected tone.
[0078] Therefore, based on the time difference, the property, the first volume, and the second volume, it is selected whether to output an output signal based on the sound signal representing the reflected sound. That is, it is appropriately selected whether to output an output signal based on the sound signal. By not outputting an output signal, the computational complexity and computational load can be reduced. In other words, a sound signal processing method that can appropriately reduce the computational complexity and computational load can be implemented.
[0079] The second aspect of the sound signal processing method disclosed herein is that, in the first aspect of the sound signal processing method, the direction from the sound source position to the position of the reflector, i.e., the position of the reflector, is calculated based on the direction from the position of the mirror image of the sound source formed via the reflector to the listening position.
[0080] Therefore, it is possible to calculate the direction from the sound source location to the reflector location based on the direction from the position of the mirror image of the sound source to the listening position.
[0081] The third embodiment of the sound signal processing method disclosed herein is as follows: In the first embodiment of the sound signal processing method, the property is the directivity of the sound. In the determination step, based on the obtained time difference, the first volume calculated according to the directivity, and the second volume calculated according to the directivity, it is determined whether to select the reflected sound.
[0082] Therefore, based on this time difference and the first and second volume levels calculated according to the directivity, it is determined whether to output an output signal. That is, it is possible to more appropriately select whether to output an output signal based on this sound signal. In other words, it is possible to implement a sound signal processing method that can more appropriately reduce the computational load.
[0083] The fourth embodiment of the sound signal processing method disclosed herein is as follows: In the sound signal processing method of the third embodiment, in the determination step, it is determined whether to select the reflected sound based on the obtained time difference and the volume ratio of the first volume calculated according to the directivity to the second volume calculated according to the directivity.
[0084] Therefore, based on this time difference and the volume ratio of the first volume to the second volume calculated according to directivity, it is selected whether to output an output signal. That is, it is possible to more appropriately select whether to output an output signal based on this sound signal. In other words, it is possible to realize a sound signal processing method that can more appropriately reduce the amount of computation and computational load.
[0085] The fifth aspect of the sound signal processing method disclosed herein is that, in the first aspect of the sound signal processing method, the property is the reflection coefficient of the reflector.
[0086] Therefore, based on the time difference, the reflection coefficient of the reflector, the first volume, and the second volume, it is determined whether to output an output signal based on the sound signal representing the reflected sound. That is, it is possible to more appropriately select whether to output an output signal based on the sound signal. In other words, it is possible to realize a sound signal processing method that can more appropriately reduce the computational load.
[0087] The sixth embodiment of the sound signal processing method disclosed herein is as follows: In the first embodiment of the sound signal processing method, the properties are the directivity of the sound and the reflection coefficient of the reflector. In the determination step, it is determined whether to select the reflected sound based on the obtained time difference, the reflection coefficient, the first volume calculated based on the directivity, and the second volume calculated based on the directivity.
[0088] Therefore, based on the time difference, the reflection coefficient, and the volume ratio of the first volume to the second volume calculated according to the directivity, it is determined whether to output an output signal. That is, it is possible to more appropriately select whether to output an output signal based on the sound signal. In other words, it is possible to realize a sound signal processing method that can more appropriately reduce the computational load.
[0089] The seventh embodiment of the sound signal processing method disclosed herein is as follows: In the sixth embodiment, when the first volume is A, the second volume is B, and the reflection coefficient is Ref, in the determination step, based on the obtained time difference and the following formula, it is determined whether to select the reflected sound. (B×Ref) / A.
[0090] Therefore, based on this time difference and the above formula, it is possible to choose whether to output an output signal. That is, to more appropriately choose whether to output an output signal based on this sound signal. In other words, it is possible to realize a sound signal processing method that can more appropriately reduce the amount of computation and computational load.
[0091] The eighth embodiment of the sound signal processing method disclosed herein is as follows: In the sixth embodiment of the sound signal processing method, when the volume ratio of the first volume to the second volume is Dir and the reflection coefficient is Ref, in the determination step, based on the obtained time difference and the following formula, it is determined whether to select the reflected sound. Dir×Ref.
[0092] Therefore, based on this time difference and the above formula, it is possible to choose whether to output an output signal. That is, to more appropriately choose whether to output an output signal based on this sound signal. In other words, it is possible to realize a sound signal processing method that can more appropriately reduce the amount of computation and computational load.
[0093] The sound signal processing method of the ninth embodiment of this disclosure is, in any of the sound signal processing methods of the first to eighth embodiments, in the determination step, further determining whether to select the reflected sound based on the position of the object present in the propagation of the sound.
[0094] Therefore, the decision to output a signal is made based on the positions of the sound source, reflector, and listener—objects present in the propagation of sound. In other words, it allows for a more appropriate selection of whether to output a signal based on the sound signal. This enables a sound signal processing method that can more appropriately reduce computational complexity and workload.
[0095] The sound signal processing method of the tenth embodiment of this disclosure is, in the sound signal processing method of the ninth embodiment, in the determination step, determining whether to select the reflected sound based on the path length ratio of the direct sound path length to the reflected sound path length, wherein the direct sound path length is the path length of the direct sound to the listening position calculated based on the position of the object, and the reflected sound path length is the path length of the reflected sound to the listening position calculated based on the position of the object.
[0096] Therefore, the decision to output a signal is also based on the path length ratio. That is, it allows for a more appropriate selection of whether to output a signal based on the audio signal. In other words, it enables an audio signal processing method that can more appropriately reduce computational complexity and workload.
[0097] The computer program of the 11th embodiment of this disclosure is a computer program for causing a computer to execute the sound signal processing method of any one of the 1st to 10th embodiments.
[0098] Therefore, the computer can execute the above-mentioned sound signal processing method according to the computer program.
[0099] The sound signal processing apparatus of the 12th embodiment of this disclosure includes: a parsing unit that calculates a time difference, which is the time difference between the time when a sound emitted from a sound source as a direct sound travels from the position of the sound source (i.e., the sound source position) to the position of the listener (i.e., the listening position) and the time when the sound is reflected by a reflector and becomes a reflected sound to reach the listening position; an acquisition unit that acquires the calculated time difference and a sound signal representing the reflected sound; a determination unit that, based on the acquired time difference, the nature of an object present in the propagation of the sound, a first volume, and a second volume, determines whether to select the reflected sound represented by the acquired sound signal, wherein the first volume is a volume corresponding to the direction from the sound source position to the listening position, and the second volume is a volume corresponding to the direction from the sound source position to the position of the reflector (i.e., the reflector position); and a reproduction unit that outputs an output signal based on a sound signal representing the reflected sound determined to be selected.
[0100] Therefore, based on the time difference, the property, the first volume, and the second volume, it is selected whether to output an output signal based on the sound signal representing the reflected sound. That is, it is appropriately selected whether to output an output signal based on the sound signal. When no output signal is output, the computational load and computational complexity can be reduced. In other words, a sound signal processing device that can appropriately reduce the computational load and computational complexity can be realized.
[0101] (Implementation Method 1) (Example of a stereo sound reproduction system) Figure 2 This diagram illustrates an example of a stereo sound reproduction system 1000. Specifically, Figure 2 This describes a stereo sound reproduction system 1000 as an example of a system capable of applying the sound processing or decoding techniques disclosed herein. Stereo sound is also referred to as immersive audio. The stereo sound reproduction system 1000 includes a sound signal processing device 1001 and a sound prompting device 1002.
[0102] The sound signal processing device 1001 is also manifested as an audio processing device, which performs audio processing on the sound signal emitted by the virtual sound source to generate an audio-processed sound signal for the listener. The sound signal is not limited to speech; any audible sound is acceptable. Audio processing, for example, is signal processing performed on the sound signal to reproduce the effects that the sound undergoes from the sound source to the listener.
[0103] The sound signal processing device 1001 performs sound processing based on spatial information describing the reasons for the aforementioned effects. Spatial information includes, for example, information indicating the location of the sound source, the listener, and surrounding objects; information indicating the shape of the space; and parameters related to sound propagation. The sound signal processing device 1001 may be, for example, a PC (Personal Computer), a smartphone, a tablet computer, or a game console.
[0104] The processed audio signal is presented to the listener by the audio prompt device 1002. The audio prompt device 1002 is connected to the audio signal processing device 1001 via wireless or wired communication. The processed audio signal generated by the audio signal processing device 1001 is transmitted to the audio prompt device 1002 via wireless or wired communication.
[0105] When the sound prompting device 1002 is composed of multiple devices, such as a device for the right ear and a device for the left ear, the multiple devices simultaneously provide prompts through communication between the multiple devices or through communication between each of the multiple devices and the sound signal processing device 1001. The sound prompting device 1002 may be, for example, a headset, earplugs, head-mounted display, or a surround sound system composed of multiple fixed speakers, worn on the listener's head.
[0106] Furthermore, the stereo sound reproduction system 1000 can also be used in combination with image prompting devices or stereoscopic image prompting devices that visually provide ER experiences, including AR / VR. For example, the space processed by spatial information is a virtual space, where the positions of sound sources, listeners, and objects are virtual positions of virtual sound sources, virtual listeners, and virtual objects within the virtual space. This space can also be represented as a sound space. Additionally, spatial information can also be represented as sound spatial information.
[0107] also, Figure 2 This is an example of a system configuration where the sound signal processing device 1001 and the sound prompting device 1002 are different devices, but the stereo sound reproduction system 1000, which can apply the sound processing method (sound signal processing method) or decoding method of this disclosure, is not limited to this. Figure 2The configuration can be as follows. For example, the sound signal processing device 1001 can be included in the sound prompting device 1002, and the sound prompting device 1002 performs both sound processing and sound prompting.
[0108] Alternatively, the sound signal processing device 1001 and the sound prompting device 1002 may share the implementation of the sound processing described in this disclosure. Furthermore, a portion or all of the sound processing described in this disclosure may be implemented via a server connected to the sound signal processing device 1001 or the sound prompting device 1002 through a network.
[0109] Furthermore, the audio signal processing apparatus 1001 can also perform audio processing by decoding a bitstream generated by encoding at least a portion of the audio signal and spatial information data used for audio processing. Therefore, the audio signal processing apparatus 1001 can also be exemplified as a decoding device.
[0110] (Example of an encoding device) Figure 3A This is a block diagram illustrating an example of the configuration of the encoding device 1100. Specifically, Figure 3A The configuration of encoding device 1100, which is an example of the encoding device of this disclosure, is shown.
[0111] Input data 1101 is encoded object data containing spatial information and / or audio signals input to encoder 1102. Details regarding the spatial information will be explained later.
[0112] Encoder 1102 encodes the input data 1101 to generate encoded data 1103. Encoded data 1103 is, for example, a bitstream generated through encoding processing.
[0113] The memory 1104 stores the encoded data 1103. The memory 1104 may be, for example, a hard disk or an SSD (Solid-State Drive), or other types of memory.
[0114] Furthermore, in the above description, the bitstream generated through encoding processing was listed as an example of encoded data 1103 stored in memory 1104, but the encoded data 1103 can also be data other than a bitstream. For example, the encoding device 1100 may also store transformed data generated by converting the bitstream into a specified data format in memory 1104. The transformed data may, for example, be a file or multiplexed stream corresponding to more than one bitstream.
[0115] Here, the file is a file with a file format such as ISOBMFF (ISO Base Media File Format). Furthermore, the encoded data 1103 can also be in the form of multiple packets generated by splitting the aforementioned bitstream or file.
[0116] For example, the bitstream generated by encoder 1102 can be transformed into data different from the bitstream. In this case, encoding device 1100 has a transformation unit (not shown), which can perform the transformation processing either by the transformation unit or by a CPU (Central Processing Unit), which is an example of a processor described later.
[0117] (Example of a decoding device) Figure 3B This is a block diagram illustrating an example of the configuration of the decoding device 1110. Specifically, Figure 3B The configuration of decoding device 1110, which is an example of the decoding device of this disclosure, is shown.
[0118] Memory 1114 stores, for example, the same data as the encoded data 1103 generated by encoding device 1100. The stored data is read from memory 1114 and input as input data 1113 into decoder 1112. Input data 1113 is, for example, a bitstream intended for decoding. Memory 1114 can be, for example, a hard disk or SSD, or other type of storage.
[0119] Alternatively, the decoding device 1110 may not input the data read from the memory 1114 as input data 1113 into the decoder 1112 as is, but instead transform the read data and input the transformed data as input data 1113 into the decoder 1112. The data before transformation may be, for example, multiplexed data containing more than one bitstream. Here, the multiplexed data may also be a file with a file format such as ISOBMFF.
[0120] Furthermore, the data before conversion can also be multiple packets generated by splitting the aforementioned bitstream or file. Alternatively, data different from the bitstream can be read from memory 1114 and converted into a bitstream. In this case, the decoding device 1110 may also include a conversion unit (not shown), and the conversion processing may be performed by the conversion unit, or by a CPU, as an example of a processor described later.
[0121] Decoder 1112 decodes the input data 1113 and generates an audio signal 1111 representing a prompt to the listener.
[0122] (Another example of an encoding device) Figure 3C This is a block diagram illustrating another configuration example of an encoding device. Specifically, Figure 3C This illustrates the configuration of encoding device 1120, which is another example of the encoding device disclosed herein. Figure 3C In China, for the sake of Figure 3A The same constituent elements are assigned to the same constituent elements as the constituent elements. Figure 3A The same labels are used for these constituent elements, and the descriptions are omitted for these constituent elements.
[0123] Encoding device 1100 stores encoded data 1103 in memory 1104. On the other hand, encoding device 1120 differs from encoding device 1100 in that it has a transmitting unit 1121 that transmits encoded data 1103 to the outside.
[0124] The transmitting unit 1121 transmits a transmission signal 1122 generated based on encoded data 1103 or data transformed from encoded data 1103 into other data formats to other devices or servers. The data used in generating the transmission signal 1122 may be, for example, a bit stream, multiplexed data, file, or packet as described in the encoding device 1100.
[0125] (Another example of a decoding device) Figure 3D This is a block diagram illustrating another configuration example of a decoding device. Specifically, Figure 3D This illustrates the configuration of decoding device 1130, another example of a decoding device disclosed herein. Figure 3D In China, for the sake of Figure 3B The same constituent elements are assigned to the same constituent elements as the constituent elements. Figure 3B The same labels are used for these constituent elements, and the descriptions are omitted for these constituent elements.
[0126] Decoding device 1110 reads input data 1113 from memory 1114. On the other hand, decoding device 1130 differs from decoding device 1110 in that it has a receiving unit 1131 that receives input data 1113 from the outside.
[0127] The receiving unit 1131 receives the received signal 1132 to obtain received data, and outputs the input data 1113 input to the decoder 1112. The received data can be the same as the input data 1113 input to the decoder 1112, or it can be data in a different format than the input data 1113.
[0128] If the format of the received data differs from the format of the input data 1113, the receiving unit 1131 may convert the received data into the input data 1113. Alternatively, the receiving data may be converted into the input data 1113 by a conversion unit (not shown) or CPU of the decoding device 1130. The received data may be, for example, a bit stream, multiplexed data, a file, or a packet as described in the encoding device 1120.
[0129] (Example of a decoder) Figure 4A This is a block diagram illustrating an example of the configuration of decoder 1200. Specifically, Figure 4A Indicates as Figure 3B or Figure 3D The decoder 1200 is an example of the decoder 1112 in the example.
[0130] Input data 1113 is the encoded bitstream, which contains encoded audio data as the encoded audio signal and metadata used in audio processing.
[0131] The Spatial Information Management Unit 1201 acquires the metadata contained in the input data 1113 and parses the metadata. The metadata contains information describing the elements that act on sound and are configured in the sound space. The Spatial Information Management Unit 1201 manages the spatial information used in sound processing obtained by parsing the metadata and provides the spatial information to the Rendering Unit 1203.
[0132] Furthermore, in this disclosure, the information used in sound processing is represented as spatial information, but other representations may also be used. For example, the information used in sound processing may be represented as sound spatial information or scene information. In addition, when the information used in sound processing changes over time, the spatial information input to the rendering unit 1203 may also be represented as information such as spatial state, sound spatial state, or scene state.
[0133] Furthermore, spatial information can be managed on a per-sound-space or per-scene basis. For example, when multiple distinct rooms are represented as virtual spaces, these rooms can be managed as separate scenes. Additionally, even within the same space, spatial information can be managed as different scenes depending on the presented conditions.
[0134] Therefore, multiple spatial information can be managed for multiple sound spaces or multiple scenes. In the management of multiple spatial information, each spatial information can be assigned an identifier to distinguish between the multiple spatial information.
[0135] Spatial information data can also be included in the bitstream, which is one example of input data 1113. Alternatively, the bitstream may contain identifiers for spatial information, and the spatial information data may be obtained from an information source outside the bitstream. Specifically, when the bitstream only contains identifiers for spatial information, the identifiers can be used during rendering to obtain spatial information data stored in the device's memory or an external server as input data 1113.
[0136] Furthermore, the information managed by the Spatial Information Management Department 1201 is not limited to the information contained in the bitstream. For example, the input data 1113 may include data on the characteristics and structure of the representation space obtained from software or servers providing VR or AR, as data not included in the bitstream.
[0137] Furthermore, the input data 1113 may also include data representing the characteristics and location of the listener or object. Additionally, the input data 1113 may include information about the listener's location obtained by sensors possessed by the terminal, including the decoding devices (1110, 1130), and may also include information representing the terminal's location inferred based on the information obtained from the sensors.
[0138] That is, the spatial information management unit 1201 can also communicate with external systems or servers to obtain spatial information and the listener's location (i.e., listening location). The spatial information management unit 1201 can obtain clock synchronization information from external systems and perform clock synchronization processing with the rendering unit 1203.
[0139] Furthermore, the space described above can be a virtually formed space, i.e., VR space, or a real space or a virtual space corresponding to a real space, i.e., AR space or MR space. Additionally, the virtual space can also be represented as a sound field or acoustic space. Furthermore, the information indicating position described above can be information such as coordinate values representing the position within the space, information representing the relative position with respect to a defined reference position, or information representing the motion or acceleration of the position within the space.
[0140] The audio data decoder 1202 decodes the encoded audio data contained in the input data 1113 to obtain the audio signal.
[0141] The encoded audio data acquired by the stereo sound reproduction system 1000 is, for example, a bitstream encoded in a format specified by MPEG-H 3D Audio (ISO / IEC 23008-3). Furthermore, MPEG-H 3D Audio is merely one example of an encoding method that can be used when generating encoded audio data contained within a bitstream. The encoded audio data can also be a bitstream encoded in other encoding methods.
[0142] For example, the encoding method can also be an irreversible codec such as MP3 (MPEG-1 Audio Layer-3), AAC (Advanced Audio Coding), WMA (Windows Media Audio), AC3 (Audio Codec-3), or Vorbis. Alternatively, the encoding method can be a reversible codec such as ALAC (Apple Lossless Audio Codec) or FLAC (Free Lossless Audio Codec).
[0143] Alternatively, any encoding method other than those described above can be used. For example, PCM (pulse code modulation) data can be used as encoded audio data. In this case, for example, if the number of quantization bits of the PCM data is N, the decoding process can be a process of converting the N-bit binary number into a number form (e.g., floating-point form) that the rendering unit 1203 can process.
[0144] The rendering unit 1203 acquires the sound signal and spatial information, uses the spatial information to perform sound processing on the sound signal, and outputs the sound processed sound signal (sound signal 1111).
[0145] Before rendering begins, the Spatial Information Management Unit 1201 reads the metadata of the input signal, detects the rendering items such as objects and sounds defined by the spatial information, and sends them to the Rendering Unit 1203. After rendering begins, the Spatial Information Management Unit 1201 monitors the changes in spatial information and the listener's location over time, updates and manages the spatial information, and sends the updated spatial information to the Rendering Unit 1203.
[0146] The rendering unit 1203 generates and outputs a sound signal with added audio processing based on the sound signal contained in the input data 1113 and the spatial information received from the spatial information management unit 1201.
[0147] Spatial information update processing and sound signal output processing with added audio processing can also be performed by the same thread. Spatial information management unit 1201 and rendering unit 1203 can assign processing to their respective independent threads. When spatial information update processing and sound signal output processing with added audio processing are performed in different threads, the start frequency of each thread can be set separately, or the processing can be performed in parallel.
[0148] When the spatial information management unit 1201 and the rendering unit 1203 perform processing in different independent threads, computing resources can be allocated preferentially to the rendering unit 1203. As a result, sound output processing that does not allow for even small delays, such as generating noise like a popping sound with a delay of 1 sample (0.02 msec), can be performed safely.
[0149] At this time, the allocation of computing resources to the spatial information management unit 1201 is restricted. However, compared to the output processing of audio signals, the updating of spatial information is a low-frequency process (e.g., updating the listener's facial orientation), and therefore does not need to be instantaneous like the output processing of audio signals. Therefore, even if the allocation of computing resources is restricted, it will not have a significant impact on the sound quality.
[0150] Spatial information updates can be performed periodically at preset times or intervals, or they can be performed when preset conditions are met. Furthermore, spatial information updates can be performed manually by the listener or the administrator of the sound space, or they can be triggered by changes in external systems.
[0151] For example, the listener can operate the controller to update the spatial information when their avatar's standing position is instantly distorted, or when it moves forward or backward. Alternatively, the spatial information can be updated when the virtual space administrator implements a performance that suddenly changes the environment of the venue. In these cases, the thread used to update the spatial information managed by the Spatial Information Management Department 1201 can be started either periodically or as a single interrupt.
[0152] The information update thread, which performs spatial information update processing, handles tasks such as updating the position or orientation of the listener's avatar within the virtual space based on the position or orientation of the VR goggles worn by the listener, and updating the position of moving objects within the virtual space. This is provided within a relatively low-frequency processing thread, typically around tens of Hz. Processing reflecting the properties of direct tones can also be performed within this low-frequency processing thread. This is because the frequency of changes in the properties of direct tones is lower than the frequency of audio processing frames used for audio output. This approach reduces the computational load on the processing and avoids the risk of impulsive noise that would result from updating information at unnecessarily high frequencies.
[0153] Figure 4B This is a block diagram representing another example of a decoder. Specifically, Figure 4B Indicates as Figure 3B or Figure 3D Another example of decoder 1112 is the construction of decoder 1210.
[0154] Figure 4B The point is that input data 1113 does not contain encoded audio data but rather unencoded audio signals. Figure 4A Different. Input data 1113 includes a bitstream containing metadata and an audio signal.
[0155] Spatial Information Management Department 1211 due to its relationship with Figure 4A The same applies to the Spatial Information Management Department 1201, so the explanation is omitted.
[0156] Rendering Department 1213 due to its relationship with Figure 4A The rendering unit 1203 is the same, so the description is omitted.
[0157] Alternatively, decoders 1112, 1200, and 1210 can also be represented as audio processing units that perform audio processing. Furthermore, decoders 1110 and 1130 can also be audio signal processing devices 1001, or can be represented as audio processing devices.
[0158] (Physical structure of a sound signal processing device) Figure 5 This diagram illustrates an example of the physical configuration of the sound signal processing device 1001. Additionally, Figure 5 The sound signal processing device 1001 can also be Figure 3B Decoding device 1110 or Figure 3D Decoding device 1130. Figure 3B or Figure 3D The multiple constituent elements shown can also be obtained through Figure 5 The various components shown are installed. Furthermore, a portion of the components described herein can also be incorporated into the sound prompting device 1002.
[0159] Figure 5 The sound signal processing device 1001 includes a processor 1402, a memory 1404, a communication interface 1403, a sensor 1405, and a speaker 1401.
[0160] The processor 1402 is, for example, a CPU, a DSP (Digital Signal Processor), or a GPU (Graphics Processing Unit). The audio processing or decoding processing of this disclosure can also be implemented by the CPU, DSP, or GPU executing a program stored in memory 1404. Furthermore, the processor 1402 is, for example, a circuit that performs information processing. The processor 1402 can also be a dedicated circuit that performs signal processing of sound signals, including the audio processing of this disclosure.
[0161] The memory 1404 may be composed of, for example, RAM (Random Access Memory) or ROM (Read Only Memory). The memory 1404 may also include magnetic recording media such as a hard disk or semiconductor memory such as an SSD. Furthermore, the memory 1404 may be internal memory built into the CPU or GPU. In addition, the memory 1404 may store spatial information managed by the spatial information management units 1201 and 1211. Furthermore, it may also store threshold data, which will be described later.
[0162] The communication IF1403 is, for example, a communication module corresponding to a communication method such as Bluetooth (registered trademark) or WIGIG (registered trademark). The audio signal processing device 1001 communicates with other communication devices, for example, via the communication IF1403, to obtain the bitstream of the decoded object. The obtained bitstream is stored, for example, in the memory 1404.
[0163] The communication IF1403, for example, consists of signal processing circuitry and an antenna corresponding to the communication method. The communication method is not limited to Bluetooth (registered trademark) and WIGIG (registered trademark), but can also be LTE (Long Term Evolution), NR (New Radio), or Wi-Fi (registered trademark), etc.
[0164] Furthermore, the communication method is not limited to the wireless communication methods mentioned above. The communication method can also be a wired communication method such as Ethernet (registered trademark), USB (Universal Serial Bus), or HDMI (registered trademark) (High-Definition Multimedia Interface).
[0165] Sensor 1405 performs sensing to infer the listener's position and orientation. Specifically, sensor 1405 infers the listener's position and / or orientation based on the detection results of one or more of the following: position, orientation, movement, velocity, angular velocity, and acceleration of a part or the whole of the body, and generates position / or orientation information representing the listener's position and / or orientation.
[0166] Alternatively, the sensor 1405 may be an external device of the sound signal processing device 1001. A part of the body may also be the listener's head, etc. The position / or orientation information may be information indicating the listener's position and / or orientation in real space, or it may be information indicating the displacement of the listener's position and / or orientation based on a predetermined point in time. Furthermore, the position / or orientation information may also be information indicating the relative position and / or orientation to the stereo sound reproduction system 1000 or the external device equipped with the sensor 1405.
[0167] Sensor 1405 can be, for example, a camera or a ranging device such as LiDAR (Light Detection and Ranging). Sensor 1405 can also capture images of the listener's head movements and detect these movements by processing the captured images. Furthermore, a device that uses wireless communication in any frequency band, such as millimeter waves, to perform position estimation can also be used as sensor 1405.
[0168] Furthermore, the sound signal processing device 1001 can also obtain location information from an external device equipped with sensor 1405 via communication IF 1403. In this case, the sound signal processing device 1001 may also not include sensor 1405. Here, the external device is, for example, a... Figure 2 The sound prompt device 1002 described herein may be a stereoscopic image reproduction device worn on the head of the listener. In this case, the sensor 1405 is configured by combining various sensors such as a gyroscope sensor and an accelerometer sensor.
[0169] For example, as the speed of the listener's head movement, sensor 1405 can detect the angular velocity of rotation about at least one of three mutually orthogonal axes in the sound space, and can also detect the acceleration of displacement about at least one of the three axes.
[0170] For example, as a measure of head movement in the listener, sensor 1405 can detect rotational motion about at least one of three mutually orthogonal axes in the sound space, and displacement about at least one of these three axes. Specifically, sensor 1405 detects 6DoF position (x, y, z) and angle (yaw, pitch, roll) as the listener's position. Sensor 1405 is constructed by combining various sensors used for motion detection, such as gyroscopes and accelerometers.
[0171] Alternatively, sensor 1405 can be implemented using a camera or a GPS (Global Positioning System) receiver, which detects the listener's location. Location information obtained by inferring location using LiDAR or similar devices as sensor 1405 can also be used. For example, in the case where the stereo sound reproduction system 1000 is implemented in a smartphone, sensor 1405 can be built into the smartphone.
[0172] Furthermore, sensor 1405 may also include a temperature sensor such as a thermocouple that detects the temperature of the sound signal processing device 1001. Additionally, sensor 1405 may also include a battery included in the sound signal processing device 1001, or a sensor that detects the remaining amount of the battery connected to the sound signal processing device 1001.
[0173] The loudspeaker 1401 includes, for example, a drive mechanism and an amplifier, such as a diaphragm, a magnet, or a voice coil, which transmits the processed sound signal as a sound prompt to the listener. The loudspeaker 1401 activates the drive mechanism based on the sound signal amplified by the amplifier (more specifically, a waveform signal representing the waveform of the sound), causing the diaphragm to vibrate. The diaphragm, vibrating in response to the sound signal, generates sound waves, which propagate through the air and reach the listener's ear, allowing the listener to perceive the sound.
[0174] Additionally, an example is given here of a sound signal processing device 1001 that includes a speaker 1401 and prompts the sound signal after sound processing via the speaker 1401, but the sound signal prompting mechanism is not limited to the above configuration.
[0175] For example, the processed sound signal can also be output to an external sound prompting device 1002 connected via a communication module. Communication via the communication module can be either wired or wireless. Furthermore, as another example, the sound signal processing device 1001 has a terminal for outputting an analog sound signal, to which a cable for an earphone or similar device is connected, and a prompting sound signal is received from the earphone or similar device.
[0176] In the above-described cases, the sound prompting device 1002 may also be a headset, earphone, head-mounted display, neck speaker, or wearable speaker worn on the head or part of the listener's body. Alternatively, the sound prompting device 1002 may also be a surround sound system consisting of multiple fixed speakers. Furthermore, the sound prompting device 1002 can also reproduce sound signals.
[0177] (Physical structure of the encoding device) Figure 6 This is a diagram illustrating an example of the physical configuration of the encoding device 1500. Figure 6 The encoding device 1500 can also be Figure 3A Encoding device 1100 or Figure 3C The encoding device 1120 can also Figure 3A or Figure 3C The multiple constituent elements shown are composed of Figure 6 The multiple components shown are installed.
[0178] Figure 6 The encoding device 1500 includes a processor 1501, a memory 1503, and a communication IF 1502.
[0179] The processor 1501 is, for example, a CPU, DSP, or GPU. The encoding processing of this disclosure can also be implemented by the CPU, DSP, or GPU executing a program stored in memory 1503. Furthermore, the processor 1501 is, for example, a circuit that performs information processing. The processor 1501 can also be a dedicated circuit that performs signal processing on an audio signal, including the encoding processing of this disclosure.
[0180] The memory 1503 may be composed of, for example, RAM or ROM. The memory 1503 may also include magnetic recording media, such as a hard disk, or semiconductor memory, such as an SSD. Furthermore, the memory 1503 may also be internal memory embedded in a CPU or GPU.
[0181] The communication IF1502 is, for example, a communication module corresponding to communication methods such as Bluetooth (registered trademark) or WIGIG (registered trademark). The encoding device 1500 communicates with other communication devices, for example, via the communication IF1502, and transmits the encoded bit stream.
[0182] The communication IF1502, for example, consists of a signal processing circuit and an antenna corresponding to the communication method. The communication method is not limited to Bluetooth (registered trademark) and WIGIG (registered trademark), but can also be LTE, NR, or Wi-Fi (registered trademark), etc. Furthermore, the communication method is not limited to wireless communication. The communication method can also be wired communication methods such as Ethernet (registered trademark), USB, or HDMI (registered trademark).
[0183] The communication module, for example, consists of a signal processing circuit and an antenna corresponding to the communication method. In the examples above, Bluetooth (registered trademark) or WIGIG (registered trademark) are cited as communication methods, but it can also correspond to communication methods such as LTE (Long Term Evolution), NR (New Radio), or Wi-Fi (registered trademark). Furthermore, the communication IF may not be a wireless communication method as described above, but rather a wired communication method such as Ethernet (registered trademark), USB (Universal Serial Bus), or HDMI (registered trademark) (High-Definition Multimedia Interface).
[0184] [The Structure of the Rendering Department] Figure 7 This is a block diagram illustrating an example of the configuration of the rendering unit 1300. Specifically, Figure 7 Indicates and Figure 4A and Figure 4B An example of the detailed configuration of the rendering unit 1300 corresponding to the rendering units 1203 and 1213.
[0185] The rendering unit 1300 consists of a resolution unit 1301, a selection unit 1302, and a reproduction unit 1303. It performs additional audio processing on the audio data contained in the input signal and outputs it.
[0186] The input signal may consist of spatial information, sensor information, and sound data. The input signal may also contain a bitstream of sound data and metadata (control information), in which case spatial information may also be included in the metadata.
[0187] Spatial information is information related to the sound space (three-dimensional sound field) formed by the stereo sound reproduction system 1000. It consists of information related to the objects contained in the sound space and information related to the listener. Among the objects, there are sound source objects that emit sound and become sound sources, and non-sound-emitting objects that do not emit sound. Sound source objects can also be simply represented as sound sources.
[0188] Non-sound-producing objects can act as obstacles reflecting the sound emitted by a sound source object, but there are also cases where a sound source object acts as an obstacle reflecting the sound emitted by other sound source objects. Obstacle objects can also be represented as reflecting objects.
[0189] As information shared by both the sound source object and the non-sound-producing object, it includes location information, shape information, and the attenuation rate of the volume when the object reflects the sound.
[0190] Position information is represented by coordinate values along three axes in Euclidean space, such as the X, Y, and Z axes, but it doesn't necessarily have to be three-dimensional. For example, position information can also be two-dimensional, represented by coordinate values along only the X and Y axes. The position information of an object is determined by the representative position of its shape, represented by a mesh or voxels.
[0191] Shape information can also include information related to the material of the surface.
[0192] The attenuation rate can be represented by a real number above 0 and below 1, or by a negative decibel value. Since volume is not amplified by reflection in real space, the attenuation rate is set to a negative decibel value. However, for example, to present the sense of horror in an unreal space, an attenuation rate above 1, i.e., a positive decibel value, can be deliberately set.
[0193] Furthermore, the attenuation rate can be set to a different value for each of the multiple frequency bands, or it can be set independently for each frequency band. Additionally, when setting the attenuation rate for each type of material on the object's surface, the corresponding attenuation rate value can be used based on information related to the surface material.
[0194] Furthermore, spatial information can also include information indicating whether an object is a living organism or whether it is a moving object. If the object is a moving object, the position indicated by the location information can also change over time. In this case, information about the changed position or the amount of change is transmitted to the rendering unit 1300.
[0195] Information related to the sound source object includes not only the information shared by the sound source object and the non-sound-producing object, but also the sound data and the information required to project the sound data into the sound space. Sound data represents information related to the frequency and intensity of the sound, and is data that reflects the sound perceived by the listener.
[0196] The audio data is typically a PCM signal, but it can also be data compressed using encoding methods such as MP3. In this case, the signal needs to be decoded at least before it reaches the reproduction unit 1303, so the rendering unit 1300 may also include a decoding unit (not shown). Alternatively, the signal can also be decoded by the audio data decoder 1202.
[0197] For a sound source object, one sound data point or multiple sound data points can be set. Additionally, recognition information can be assigned to each sound data point, and information related to the sound source object can also include this recognition information.
[0198] Information required to project sound data into the sound space may include, for example, reference volume information used as a reference in the reproduction of sound data, information representing the properties (also called characteristics) of the sound data, information related to the location of the sound source object, and information related to the orientation of the sound source object (i.e., information related to the directivity of the sound emitted by the sound source object).
[0199] Reference volume information can be, for example, the effective value of the amplitude of the sound data at the sound source location when the sound data is radiated into the sound space, or it can be represented as a decibel (dB) value in floating point.
[0200] For example, when the reference volume is 0dB, it can also mean that the volume of the signal level represented by the sound data is not increased or decreased, but the sound is radiated into the sound space at the original volume from the position indicated by the information related to the location of the sound source object. Alternatively, when the reference volume is -6dB, it can also mean that the volume of the signal level represented by the sound data is set to approximately half, and the sound is radiated into the sound space from the position indicated by the information related to the location of the sound source object.
[0201] The reference volume information can be assigned to each sound data point individually, or it can be assigned to multiple sound data points uniformly.
[0202] Information representing the properties of sound data can be, for example, information related to the volume of the sound source, and information representing the time-series variation of the sound source's volume.
[0203] For example, in a virtual conference room where the sound space is a speaker and the sound source is the speaker, the volume changes intermittently over a short period of time. That is, the spoken and silent parts alternate. Conversely, in a concert hall where the sound space is a concert hall and the sound source is a performer, the volume is maintained for a certain duration. Furthermore, in a battlefield where the sound space is an explosive device and the sound source is an explosive device, the volume of the explosion sound only increases momentarily, then remains silent or low.
[0204] In this way, the volume information of the sound source includes not only the magnitude of the sound, but also information about changes in the magnitude of the sound. This information can also be used to represent the properties of the sound data.
[0205] Transition information can also be represented by time-series data showing frequency characteristics. Transition information can also be represented by data showing the duration of the sound interval. Transition information can also be represented by time-series data showing the duration of both the sound and silent intervals. Transition information can also be represented by multiple sets of time-series data listing the duration during which the amplitude of the sound signal can be considered stationary (approximately constant) and the amplitude values of the signal during that period.
[0206] Transition information can also be represented by data that allows the frequency characteristics of a sound signal to be considered as a stationary duration. Transition information can also be represented by listing multiple sets of data, in time series, of the duration during which the frequency characteristics of the sound signal can be considered stationary and the frequency characteristics during that period. Transition information can also be represented, for example, by data showing the approximate shape of a spectrogram.
[0207] Furthermore, the volume used as the reference for the aforementioned frequency characteristics can also be the aforementioned reference volume. Information about the reference volume and information representing the properties of the sound data can be used for calculating the volume of direct or reflected sounds perceived by the listener, or for selecting whether or not to allow the listener to perceive them. Other examples and methods of using information representing the properties of the sound data will be described later.
[0208] Furthermore, the reflected sound in this embodiment is an example of an indirect sound. An indirect sound can be a reflected sound, a diffracted sound, or the like. In this embodiment, a reflected sound, as an example of an indirect sound, is used for explanation, but the same treatment is performed even if an indirect sound is used instead of a reflected sound.
[0209] Information related to the orientation of a sound source object (orientation information) is typically represented by yaw, pitch, and roll. Alternatively, the roll rotation can be omitted, and the orientation information of the sound source object can be represented by azimuth (yaw) and pitch (pitch). The orientation information of the sound source object can also change over time, and any changes are transmitted to the rendering unit 1300.
[0210] Information relevant to the listener relates to their position and orientation within the sound space. Position information is represented by the XYZ axes of Euclidean space, but it doesn't necessarily have to be three-dimensional; it can also be two-dimensional. Orientation information is typically represented by yaw, pitch, and roll. Alternatively, the roll can be omitted, and the listener's orientation information can be represented by azimuth (yaw) and pitch (pitch).
[0211] The listener's location and orientation information can also change over time, and if changes occur, they will be transmitted to the rendering unit 1300.
[0212] The sensor information includes information such as the amount of rotation or displacement detected by the sensor 1405 worn by the listener, as well as the listener's position and orientation. The sensor information is transmitted to the rendering unit 1300, which updates the listener's position and orientation information based on the sensor information. The sensor information may also include, for example, position information obtained by the portable terminal using GPS, a camera, or LiDAR for self-position estimation.
[0213] Alternatively, instead of sensor 1405, information obtained from an external source via a communication module can be used as sensor information for detection. Information indicating the temperature of the sound signal processing device 1001 and the remaining battery level can also be obtained from sensor 1405. Furthermore, the computing resources (CPU capacity, memory resources, or PC performance, etc.) of the sound signal processing device 1001 or the sound prompting device 1002 can be obtained in real time.
[0214] The analysis unit 1301 analyzes the sound signal contained in the input signal and the spatial information received from the spatial information management units 1201 and 1211, and calculates the information required for generating direct sound and reflected sound in the reproduction unit 1303, as well as the information required for selecting whether to generate reflected sound.
[0215] The information required for the generation of direct and reflected tones includes values related to the path taken to reach the listening position, the time taken to reach the position, and the volume upon arrival for each direct and reflected tone. These values represent, for example, the path taken to reach the listening position, the time taken to reach the position, and the volume upon arrival.
[0216] The information required for selecting the output reflected tone is information representing the relationship between the direct tone and the reflected tone, such as values related to the time difference between the direct tone and the reflected tone, and values related to the volume ratio of the direct tone and the reflected tone at the listening position. For example, the values related to the time difference between the direct tone and the reflected tone, and the values related to the volume ratio of the direct tone and the reflected tone at the listening position, are respectively values representing the time difference between the direct tone and the reflected tone, and values representing the volume ratio of the direct tone and the reflected tone at the listening position.
[0217] Furthermore, when volume is expressed in decibels on a logarithmic axis (in the case of representing volume in the decibel range), the volume ratio of two signals is naturally represented by the difference in decibel values. Specifically, the volume ratio of two signals can be the difference between the amplitude values of each signal expressed in the decibel range. This value can also be calculated based on energy values or power values, etc. Moreover, this difference can be referred to as the gain difference or simply the gain difference in the decibel range.
[0218] That is, the volume ratio in this disclosure is essentially the ratio of signal amplitudes, so it can also be expressed as Soundvolume ratio, Volume ratio, Amplitude ratio, Soundlevel ratio, Sound intensity ratio, or Gain ratio, etc. Furthermore, when the unit of volume is decibels, the volume ratio in this disclosure can of course also be referred to as volume difference.
[0219] In this disclosure, "volume ratio" typically refers to the gain difference between two sounds expressed in decibels. In examples of embodiments, the threshold data is also typically defined by the gain difference expressed in decibels. However, the volume ratio is not limited to the gain difference in decibels. When using a volume ratio expressed outside the decibel range, the threshold data defined in the decibel range can be converted to the units of the calculated volume ratio for use. Alternatively, the threshold data defined in each unit can be stored in memory in advance.
[0220] That is, even if a ratio such as energy value or power value is used instead of volume ratio, it is obvious that the algorithm in this disclosure can be applied to the solution of the problem of this disclosure.
[0221] The time difference between direct and reflected sound refers, for example, the time difference between the moment the direct sound arrives at the listening position from the sound source location and the moment the sound emitted from the sound source is reflected by a reflective object and arrives at the listening position. Furthermore, for simplicity, the time difference between the moment the direct sound arrives at the listening position and the moment the reflected sound arrives at the listening position is sometimes recorded as the time difference between direct and reflected sound, or simply as the time difference. Additionally, the time difference between direct and reflected sound can be, for example, the time difference between the arrival time of the direct sound and the arrival time of the reflected sound. The time difference between direct and reflected sound can also be the time difference between the moments when the direct sound and the reflected sound arrive at the listening position, or the difference in the time required for the direct sound and the reflected sound to arrive at the listening position. The calculation methods for these values will be described later.
[0222] The selection unit 1302 uses the information calculated by the analysis unit 1301 and the threshold data to select whether the reproduction unit 1303 will generate the reflected sound. In other words, the selection unit 1302 determines whether to select the reflected sound as the target reflected sound for generation. In other words, the selection unit 1302 selects which of the multiple reflected sounds generated by the reproduction unit 1303 will be the reflected sound.
[0223] Threshold data, for example, can be represented in a graph where the horizontal axis represents the time difference between the direct and reflected sounds, and the vertical axis represents the volume ratio of the direct and reflected sounds. This represents the boundary (threshold) at which the reflected sound is perceived or not. Threshold data can also be represented by an approximation using the time difference between the direct and reflected sounds as a variable, or by an arrangement of values indexed by the time difference between the direct and reflected sounds and corresponding thresholds.
[0224] Selection unit 1302 selects to generate a reflected sound if, for example, the ratio of the volume of the direct sound at arrival to the volume of the reflected sound at arrival is greater than a threshold set based on reference threshold data, among the values of the time difference between the arrival time of the direct sound and the arrival time of the reflected sound. Furthermore, the arrival volume refers to the volume of the sound when it reaches the listening position.
[0225] The time difference between the arrival time of the direct tone and the arrival time of the reflected tone is, in other words, the difference in time required for the direct tone and the reflected tone to reach the listening position respectively. Alternatively, the time difference between the end of the direct tone's emission and the arrival time of the reflected tone at the listening position can also be used as the time difference between the direct tone and the reflected tone. In this case, different threshold data can be used than the threshold data set based on the time difference between the arrival times of the direct tone and the reflected tone, or a common threshold data can be used.
[0226] The threshold data can be obtained from the memory 1404 of the audio signal processing device 1001, or from an external storage device via the communication module. The method for storing the threshold data and the method for setting the threshold will be described later.
[0227] The reproduction unit 1303 synthesizes the direct sound signal with the reflected sound signal selected and generated by the selection unit 1302.
[0228] Specifically, the reproduction unit 1303 processes the input sound signal to generate a direct tone based on the information about the arrival time and volume of the direct tone calculated by the analysis unit 1301. Furthermore, the reproduction unit 1303 processes the input sound signal to generate a reflected tone based on the information about the arrival time and volume of the reflected tone selected by the selection unit 1302. Then, the reproduction unit 1303 combines the generated direct tone and reflected tone and outputs them.
[0229] The reproduction unit 1303 may also include a reproduction unit with a dispersion filter. A dispersion filter is a filter that causes the phase of sound to spread (generate deviation), and can simulate the diffusion of sound. For the dispersion filter, the signal amplification rate may be preset, or the reproduction unit 1303 may be equipped with a unit that sets the signal amplification rate.
[0230] Figure 7 The rendering unit 1300 shown may include multiple resolution units 1301 and selection units 1302. The processing performed by the resolution units 1301, selection units 1302, and reproduction unit 1303 may, for example, be performed as part of the pipeline processing described in Patent Document 3. Pipeline processing refers to dividing the processing for imparting sound effects into multiple processes and executing these processes sequentially. In each of the multiple processes, for example, signal processing for a sound signal or the generation of parameters for signal processing may be performed.
[0231] [Example of actions in the rendering department] Figure 8 This is a flowchart illustrating an example of the operation of the sound signal processing device 1001. Figure 8 The text indicates that the processing is mainly performed by the rendering unit 1300 of the sound signal processing device 1001.
[0232] In the analysis and processing of input signals ( Figure 8 In step S101, the analysis unit 1301 analyzes the input signal input to the sound signal processing device 1001 and detects direct tones and reflected tones that can be generated in the sound space. The reflected tones detected here are candidate reflected tones selected by the selection unit 1302 as the final reflected tones to be generated by the reproduction unit 1303. Furthermore, the analysis unit 1301 analyzes the input signal and calculates the information needed for generating direct tones and reflected tones, as well as the information needed for selecting the target reflected tones.
[0233] First, the characteristics of direct and reflected tones are calculated. Specifically, the arrival time and volume of the direct and reflected tones when they reach the listener are calculated. If multiple objects exist in the sound space as reflection targets, the characteristics of the reflected tones are calculated for each object separately.
[0234] The direct arrival time (td) is calculated based on the direct arrival path (pd). The direct arrival path (pd) is the path connecting the location information S(xs, ys, zs) of the sound source object with the location information A1(xa, ya, za) of the listener. The direct arrival time (td) is the value obtained by dividing the length of the path connecting the location information S(xs, ys, zs) and the location information A1(xa, ya, za) by the speed of sound (approximately 340 m / s).
[0235] For example, the path length (X) is calculated using (xs-xa)^2 + (ys-ya)^2 + (zs-za)^2)^0.5. Volume decreases inversely with distance. Therefore, given the volume as N in the location information S(xs, ys, zs) of the sound source object and the unit distance as U, the volume (ld) when the direct sound arrives is calculated using ld = N. Find U / X.
[0236] The volume N at the sound source location can also be the reference volume described earlier.
[0237] The arrival time (tr) of the reflected sound is calculated based on the arrival path (pr). The arrival path (pr) is the path that connects the position of the sound image of the reflected sound with the position information A1 (xa, ya, za).
[0238] Furthermore, the location of the sound image of reflected sound can be derived using methods such as the "mirror method" or the "ray tracing method," or any other method. The mirror method assumes that the reflected wave on the wall of a room has a mirror image at a position symmetrical to the sound source relative to the wall, and simulates the sound image by assuming that sound waves are emitted from that mirror image location. The ray tracing method simulates the image (sound image) observed at a certain point by tracing waves that propagate in straight lines, such as light rays or sound rays.
[0239] Figure 9 It is a diagram that shows the relative positions of the listener and the obstacle. Figure 10 This is a diagram showing the relatively close positional relationship between the listener and the obstacle object. That is, Figure 9 and Figure 10 These examples illustrate sound images formed at positions symmetrical to the sound source location, separated by a wall. By determining the position of the sound image of the reflected sound on the x, y, and z axes based on this relationship, the arrival time of the reflected sound can be calculated in the same way as the arrival time of the direct sound.
[0240] The arrival time (tr) of the reflected sound is obtained by dividing the length (Y) of the path connecting the position of the sound image of the reflected sound to the position information A1 (xa, ya, za) by the speed of sound (approximately 340 m / s). Volume decreases inversely with distance. Therefore, when the volume at the sound source is N, the unit distance is U, and the attenuation rate of the volume in the reflection is G, the volume (lr) when the reflected sound arrives decreases by lr = N. G Find U / Y.
[0241] As explained above, the attenuation rate G can be represented by a real number greater than or equal to 0 and less than 1, or by a negative decibel value. In this case, the overall volume attenuation of the signal corresponds to the amount of G. Furthermore, the attenuation rate can also be set for each of the multiple frequency bands. In this case, the analysis unit 1301 applies the specified attenuation rate for each frequency component of the signal. Additionally, to reduce computational complexity, the analysis unit 1301 can use representative values or average values of multiple attenuation rates from multiple frequency bands as the overall attenuation rate, thereby correspondingly attenuating the overall volume of the signal.
[0242] Next, the analysis unit 1301 calculates the volume ratio (L) required in the selection of the reflected sound of the generated object, namely the ratio of the volume (ld) when the direct sound arrives to the volume (lr) when the reflected sound arrives, and the time difference (T) between the direct sound and the reflected sound.
[0243] The ratio of the volume (ld) of the direct tone to the volume (lr) mentioned above, i.e., the volume ratio (L), is, for example, the value obtained by dividing the volume (lr) of the reflected tone by the volume (ld) of the direct tone, given by L = (N G U / Y) / (N) U / X) = G X / Y is calculated. Since the calculated value is a volume ratio, the values of N and U can be any pre-set values.
[0244] The time difference (T) between the direct tone and the reflected tone can also be the time difference between the direct tone and the reflected tone when they arrive at the listening position. For example, the time difference (T) between the direct tone and the reflected tone when they arrive at the listening position can be calculated by T = tr - td.
[0245] Furthermore, the time difference (T) can also be the difference between the arrival times of the direct tone and the reflected tone at the listening position. Alternatively, the time difference (T) can also be the time difference between the end of the direct tone's speech and the arrival time of the reflected tone at the listening position. In other words, the time difference (T) can also be the time difference between the end of the direct tone and the beginning of the reflected tone at the listening position.
[0246] Next, in the selective processing of reflected sounds ( Figure 8 In step S102, the selection unit 1302 selects whether the reproduction unit 1303 generates the reflected tone calculated by the analysis unit 1301. In other words, the selection unit 1302 determines whether to select the reflected tone as the target reflected tone for generation. When multiple reflected tones exist, the selection unit 1302 selects whether to generate each reflected tone. The result of the selection unit 1302's decision on whether to generate each reflected tone can be either selecting more than one target reflected tone from multiple reflected tones, or selecting none of the target reflected tones.
[0247] Furthermore, the selection unit 1302 is not limited to generation processing; it can also select reflected sounds that are the application targets of other processing. For example, the selection unit 1302 can also select reflected sounds that are the application targets of binaural processing. In addition, the selection unit 1302 generally selects only one or more reflected sounds that are the processing targets. However, the selection unit 1302 can also select only one or more reflected sounds that are not the processing targets. Furthermore, processing can be applied to one or more reflected sounds that are not selected.
[0248] For example, the selection of reflected tones is based on the volume ratio (L) and time difference (T) calculated by the analysis unit 1301. By performing selection processing based on the time difference (T) between the direct tone and the reflected tone, compared to the case where selection processing is based solely on the volume difference between the direct tone and the reflected tone, it is possible to more appropriately select reflected tones that have a greater impact on the listener's perception.
[0249] Specifically, the decision to generate a reflected tone is made, for example, by comparing the volume ratio of the direct tone to the reflected tone, corresponding to the time difference between the direct tone and the reflected tone, with a pre-set threshold. The threshold is set with reference to threshold data. The threshold data is an indicator representing the boundary at which the reflected tone of a direct tone is perceived by the listener, defined by the ratio of the volume (Id) when the direct tone arrives to the volume (lr) when the reflected tone arrives.
[0250] Furthermore, a threshold corresponds to a value expressed as a numerical value set in relation to a time difference (T). Threshold data corresponds to the relationship between the time difference (T) and the threshold, and to tabular data or formulas used to determine or calculate the threshold under the time difference (T). The form and type of threshold data are not limited to tabular data or formulas.
[0251] Figure 11 This is a graph showing the relationship between the time difference of the direct tone and the reflected tone and the threshold. For example, you can also refer to... Figure 11 The threshold data shown is for the volume ratio preset according to each value of the time difference between the direct tone and the reflected tone. Alternatively, one can refer to the data from... Figure 11 The threshold data shown is obtained through interpolation or extrapolation.
[0252] Furthermore, the threshold for the volume ratio under the time difference (T) calculated by the analysis unit 1301 is determined based on the threshold data. The selection unit 1302 then decides whether to select the reflected sound as the target reflected sound based on whether the volume ratio (L) of the direct sound to the reflected sound calculated by the analysis unit 1301 is higher than this threshold.
[0253] By using threshold data of volume ratios preset according to each value of the time difference between direct and reflected tones, selection processing that takes into account backmasking or priority effects can be achieved. Detailed explanations of the types, formats, storage methods, and setting methods of the threshold data will be provided later.
[0254] Next, in the generation and processing of direct and reflected tones ( Figure 8 In S103, the reproduction unit 1303 generates a direct sound signal and a reflected sound signal selected by the selection unit 1302 as the target reflected sound, and synthesizes them.
[0255] The direct tone sound signal is generated by applying the direct tone arrival time (td) and the direct tone arrival volume (ld) calculated by the analysis unit 1301 to the sound data of the sound source object contained in the input information. Specifically, the sound data is delayed by the direct tone arrival time (td) and multiplied by the direct tone arrival volume (ld). The sound data delay process is a process of shifting the position of the sound data back and forth on the time axis. For example, the sound data delay process disclosed in Patent Document 2 without degrading the sound quality can also be applied.
[0256] The sound signal of the reflected sound is generated in the same way as the direct sound by applying the sound data of the sound source object to the arrival time (tr) of the reflected sound and the volume (lr) of the reflected sound when it arrives, which is calculated by the analysis unit 1301.
[0257] However, the volume (lr) of the reflected sound during its arrival differs from the volume (ld) of the direct sound. It is affected by the attenuation rate G applied to the volume of the reflected sound. G can be an attenuation rate applied across the entire frequency band. Alternatively, it can be a reflection rate specified for each defined frequency band to reflect the bias in the frequency components generated by reflection. In this case, the application of the volume (lr) of the reflected sound can also be implemented as a frequency equalizer process, multiplying the volume by the attenuation rate for each frequency band.
[0258] In the example above, the path lengths of the direct and reflected tone candidates as they reach the listener are calculated. Then, the arrival time and volume are calculated based on each path length. Finally, the reflected tone candidates are selected based on their time difference and volume ratio.
[0259] Alternatively, as another example, selection can be performed based on the path lengths of the direct and reflected sounds as they reach the listener, omitting the calculations of arrival time and volume, as well as time difference and volume ratio. In this case, a threshold corresponding to the path length difference can be pre-set for the path length ratio. Furthermore, selection can be performed based on whether the calculated path length ratio is above the threshold corresponding to the calculated path length difference. Thus, selection can be performed based on the path length difference corresponding to the time difference while reducing computational load.
[0260] In addition to the path length difference, the value of a parameter representing the speed of sound propagation, or the value of a parameter that affects the speed of sound propagation, can also be used.
[0261] (Select the detailed processing option) The details of the selection process for whether or not to generate reflected sound are explained.
[0262] The selection of the reflected sound is performed by comparing a threshold value for the volume ratio of the direct sound to the reflected sound at a given time difference (T), i.e., a volume ratio threshold, with a volume ratio (L) calculated by the analysis unit 1301. For example, among the volume ratio thresholds preset for each value of the time difference between the direct sound and the reflected sound, the threshold value for the volume ratio at the time difference (T) between the direct sound and the reflected sound calculated by the analysis unit 1301 is referenced. Furthermore, whether the reflected sound is selected as the target reflected sound is determined based on whether the volume ratio (L) calculated by the analysis unit 1301 is higher than the threshold value.
[0263] The time difference (T) can be, for example, the difference in the time when the direct tone and the reflected tone arrive at the listening position, the time difference in the time required for the direct tone and the reflected tone to arrive at the listening position, or the time difference between the time when the direct tone ends and the time when the reflected tone arrives at the listening position. Here, the end time of the direct tone can also be obtained, for example, by adding the duration of the direct tone to the time of its arrival.
[0264] Regarding threshold data, it can also be determined through auditory nerve action or cognitive function of the brain. More specifically, it can be determined based on the minimum time difference between two sounds that the listener can perceive and detect, through the prioritization effect, the time-dependent masking phenomenon, or a combination thereof (described later). Specific values can be derived from known research findings on time-dependent masking effects, prioritization effects, or echo detection limits, or they can be determined through audiovisual experiments applied to this virtual space.
[0265] Figure 12A , Figure 12B and Figure 12C This is a diagram illustrating an example of how threshold data is set. For example... Figure 12A , Figure 12B and Figure 12C As shown in the figure, the threshold data is represented by the boundary (threshold) at which the reflected sound is perceived or not in the curve with the horizontal axis representing the time difference between the direct sound and the reflected sound and the vertical axis representing the volume ratio between the direct sound and the reflected sound.
[0266] Threshold data can also be approximated by using the time difference between the direct tone and the reflected tone as a variable. Furthermore, threshold data can also be used as... Figure 11 The indexes of the time differences between the direct and reflected tones, and the corresponding thresholds, are stored in the region of memory 1404.
[0267] In addition, Figure 12C In Example 4, when the height of the line parallel to the horizontal axis (the minimum audible limit) is used as the threshold, the volume ratio (L) of the direct tone to the reflected tone is not compared to the threshold; rather, the volume of the reflected tone itself is compared to the threshold. This is because the threshold represents the volume boundary of whether a sound can be perceived by a listener, and it is the threshold used to determine sounds with volumes lower than this threshold as unreproducible sounds. That is, the threshold corresponding to the minimum audible limit is not a threshold for the ratio of the volume of the reflected tone to the volume of the direct tone.
[0268] When the minimum audible limit is used as the threshold, the threshold is constant and independent of the time difference (T), so the time difference (T) can be disregarded.
[0269] Additionally, in the parsing process ( Figure 8 In the case where multiple reflected sounds are generated in S101, selection processing can be performed on all reflected sounds, or selection processing can be performed only on reflected sounds with high evaluation values based on the evaluation values derived from each reflected sound using a pre-set evaluation method. Here, the evaluation value of a reflected sound corresponds to its perceived importance. Furthermore, a high evaluation value corresponds to a large evaluation value, and these interpretations can be interchanged.
[0270] The selection unit 1302 may, for example, calculate the evaluation value of the reflected sound using a pre-set evaluation method corresponding to the volume of the sound source, the visuality of the sound source, the localization of the sound source, the visuality of the reflecting object (obstacle object), or the geometric relationship between the direct sound and the reflected sound.
[0271] Specifically, a higher volume of the sound source can result in a higher evaluation value. Furthermore, to ensure visual and acoustic localization are consistent, a higher evaluation value can also be achieved when the sound source object or its reflection (obstacle) is visible to the listener, or when the sound source object has high localization.
[0272] Furthermore, the opening of the angles of arrival of the direct and reflected tones, as well as the difference in their arrival times, significantly impact spatial perception. Therefore, a higher evaluation value can be achieved when the openings of the angles of arrival of the direct and reflected tones are large, or when the difference in their arrival times is significant.
[0273] The volume information of the audio source can also represent the base volume set for each content, the time transition of the volume, or both.
[0274] For example, in a virtual conference room where the virtual space is a virtual meeting room and the direct audio is conversational sound, the volume changes intermittently over a short period of time. That is, the audible and silent parts alternate. Conversely, in a concert hall where the virtual space is a concert hall and the direct audio is a musical performance, the volume is maintained for a certain duration. Finally, in a battlefield where the virtual space is a battlefield and the direct audio is an explosion, the volume increases only momentarily, then remains silent or low.
[0275] In this way, the volume information of the sound source not only includes information about the reference volume that corresponds to the volume setting when the sound is radiated into the virtual space, but also information about the changes in the volume of the sound.
[0276] Transition information can also be represented by time-series data showing frequency characteristics. Transition information can also be represented by data showing the duration of the sound interval. Transition information can also be represented by time-series data showing the duration of both the sound and silent intervals. Transition information can also be represented by multiple sets of time-series data listing the duration during which the amplitude of the sound signal can be considered stationary (approximately constant) and the amplitude values of the signal during that period.
[0277] Information about the transition can also be represented by data that shows the duration of a stationary frequency response of a sound signal. Alternatively, it can be represented by a time series listing multiple sets of data showing the duration of a stationary frequency response of a sound signal and the frequency response during that period.
[0278] Furthermore, there has been a long-standing and widespread practice of using temporal shifts in the frequency characteristics of signals for audio processing in virtual space (Patent Document 1, etc.). Given this prior art, the aforementioned group could also be a group of time lengths with constant frequency characteristics and their frequency characteristics.
[0279] Geometric relationships can also refer to the relationships between the positions of a sound source, a listener, and a reflecting object within a virtual space. Through these relationships, the path lengths of both direct and reflected sounds can be calculated geometrically. Therefore, by utilizing the inverse relationship between volume and distance, the reference volume of the reflected sound relative to the reference volume of the direct sound can be calculated.
[0280] In calculating the reference volume of reflected sound, the reflection coefficient of the reflecting object can also be used. Alternatively, a typical value commonly used can be used as the reflection coefficient. On the other hand, in special cases where the reflecting object is covered by sound-absorbing material, a specially assigned reflection coefficient can be used as the reflection coefficient of the reflecting object.
[0281] Reflected sound can also be evaluated based on its volume. The volume of the reflected sound can be determined based on the geometric relationship between the direct sound and the reflected sound as described above, as well as the index assigned to the reflecting object. The volume can also be compared to a pre-set threshold to evaluate the reflected sound.
[0282] Furthermore, information representing the temporal change in the volume of the sound source can also be reflected in the evaluation. For example, if the information representing the temporal change in the volume of the sound source represents the duration of the sound interval, the evaluation value of the reflected sound can remain unchanged when the time is within the sound interval. On the other hand, if the time is outside the sound interval, even if the reference volume of the reflected sound exceeds a threshold, the evaluation value of the reflected sound can be reduced or set to zero.
[0283] Alternatively, information representing the temporal shift in the volume of a sound source can also be data that lists multiple groups of amplitude values of the signal over a duration during which the amplitude of the sound signal is considered approximately constant, in a time series. In this case, the processing of the reflected sound can also be evaluated by changing the reference volume of the reflected sound in conjunction with the changes in the amplitude values in the data.
[0284] Alternatively, information representing the volume of a direct tone can be obtained using both reference volume information and information about the volume that changes over time. For example, after calculating an evaluation value based on the reference volume information, the information about the changed volume can be used to correct that evaluation value.
[0285] In evaluating reflected sounds, all of the above methods can be performed, or only some of them can be performed. For example, reflected sounds can be evaluated using multiple evaluation methods, or they can be evaluated using only one evaluation method.
[0286] When evaluating reflected sounds using multiple evaluation methods, the decision to select a reflected sound can be based on the combined evaluation value determined by the multiple evaluation methods, or it can be based on the individual evaluation values of the multiple evaluation methods.
[0287] The sound signal processing device 1001 may select a sound if, when deciding whether to select a reflected sound based on each of multiple evaluation methods, all evaluation results based on multiple evaluation methods indicate that sound should be selected. Alternatively, the sound signal processing device 1001 may select a sound if one of the evaluation results based on multiple evaluation methods indicates that sound should be selected.
[0288] Alternatively, priorities can be set for the first to third evaluation methods. Furthermore, the sound signal processing device 1001 can also ultimately determine that sound is not selected if the first evaluation method determines that sound is not selected, regardless of the determination results in the second and third evaluation methods. Additionally, the sound signal processing device 1001 can also ultimately determine that sound is selected if one of the second and third evaluation methods determines that sound is not selected, but the other determines that sound is selected.
[0289] Furthermore, selection processing and evaluation processing can be performed independently, or only one of them can be performed. Alternatively, evaluation processing can be performed only on reflected sounds that were determined to be selected in the selection processing, and the evaluation processing can then determine whether to select the reflected sound again. Or, evaluation processing can be performed only on reflected sounds that were determined not to be selected in the selection processing, and the evaluation processing can then determine whether to select the reflected sound again.
[0290] The selection process described above can be interpreted as selecting reflected sounds based on the properties of the direct sound. For example, in selecting reflected sounds based on the properties of the direct sound, a threshold used in the selection of reflected sounds is set or adjusted according to the properties of the direct sound. Alternatively, an evaluation value used in the selection of reflected sounds can be calculated based on one or more of the following: the volume of the sound source, the visibility of the sound source, the localization of the sound source, the visibility of the reflecting object (obstacle object), and the geometric relationship between the direct sound and the reflected sound.
[0291] Furthermore, the process of selecting reflected tones based on the properties of direct tones is not limited to setting or adjusting thresholds based on the properties of direct tones, or calculating evaluation values used in selecting reflected tones of the processed object; other processes may also be performed. Moreover, when performing processes such as setting or adjusting thresholds based on the properties of direct tones, or calculating evaluation values used in selecting reflected tones of the processed object, a portion of the process may be modified, or new processes may be added.
[0292] In addition, setting thresholds can also include adjusting thresholds and changing thresholds.
[0293] [How to set the threshold] The threshold data used in the selection process can be set by referring to known values of echo detection limits based on priority effects or masking thresholds based on backmasking effects.
[0294] The priority effect refers to the phenomenon where, when two sounds are heard from different locations, the listener perceives which sound source was heard first in time. If two short sounds merge and sound like one, the overall sound's perceived location (locality) is largely determined by the location of the initial sound. The echo detection limit, a phenomenon occurring through the priority effect, is the minimum time difference that allows a listener to perceive the discrepancy between two sounds.
[0295] exist Figure 12C In Example 2, the horizontal axis corresponds to the arrival time of the reflected sound (echo), specifically, the delay time from the arrival time of the direct sound to the arrival time of the reflected sound. The vertical axis corresponds to the volume ratio of the detectable reflected sound to the direct sound, specifically, the threshold for whether the reflected sound arriving with the delay time can be detected.
[0296] Figure 13 This is a diagram illustrating an example of how a threshold is set. Figure 13 The horizontal axis in the figure corresponds to the arrival time of the reflected sound, specifically the time difference (T) between the direct sound and the reflected sound. Figure 13 The vertical axis corresponds to the volume of the reflected sound. Specifically, Figure 13 The vertical axis can correspond to either the volume of the reflected sound (volume ratio) which is set relatively to the volume of the direct tone, or the volume of the reflected sound which is determined absolutely without depending on the volume of the direct tone.
[0297] For example, in such Figure 9 When the listener is relatively far from the obstacle, the reflected sound arrives later, such as... Figure 13 As shown in C, the threshold is set low. As a result, in Figure 9 In such cases, reflected sounds are generated. On the other hand, in cases such as Figure 10 When the listener is relatively close to the obstacle, the arrival time of the reflected sound is... Figure 9 The condition occurs earlier, such as Figure 13 As shown in B, the threshold is set high. As a result, in Figure 10 In this case, no reflected sound is generated.
[0298] In addition, threshold data can also be stored in memory 1404 and retrieved from memory 1404 for use in selection processing.
[0299] Figure 14This is a flowchart illustrating an example of the selection process. First, the selection unit 1302 specifies the reflected tone detected by the analysis unit 1301 (S201). Next, the selection unit 1302 detects the volume ratio (L) of the direct tone and the reflected tone, as well as the time difference (T) between the direct tone and the reflected tone (S202 and S203).
[0300] The time difference (T) can be, for example, the time difference between the arrival time of the direct tone and the reflected tone at the listening position, the time difference between the arrival time of the direct tone and the arrival time of the reflected tone, or the time difference between the end of the direct tone's sound and the arrival time of the reflected tone at the listening position. An example based on the time difference between the arrival time of the direct tone and the arrival time of the reflected tone is given here.
[0301] Specifically, the selection unit 1302 calculates the difference between the length of the direct sound path and the length of the reflected sound path based on the location information of the sound source object and the listener, as well as the location and shape information of the obstacle object. Furthermore, the selection unit 1302 detects the time difference (T) between the time when the direct sound arrives at the listener's position and the time when the reflected sound arrives at the listener's position by dividing this length by the speed of sound.
[0302] The volume reaching the listener decreases proportionally (inversely proportional to distance) to the volume of the sound source. Therefore, the volume of the direct tone is obtained by dividing the volume of the sound source by the length of the direct tone's path. The volume of the reflected tone is obtained by dividing the volume of the sound source by the length of the reflected tone's path and then multiplying by the attenuation rate assigned to the virtual obstacle object. The selection unit 1302 detects the volume ratio by calculating the ratio of their volumes.
[0303] Furthermore, the selection unit 1302 uses threshold data to determine a threshold corresponding to the time difference (T) (S204). Next, the selection unit 1302 determines whether the detected volume ratio (L) is above the threshold (S205).
[0304] When the volume ratio (L) is above the threshold ("Yes" in S205), the selection unit 1302 selects the reflected sound as the reflected sound of the generation target (S206). When the volume ratio (L) is below the threshold ("No" in S205), the selection unit 1302 does not select the reflected sound as the reflected sound of the generation target (S207). That is, in this case, the selection unit 1302 determines the reflected sound as a reflected sound other than the generation target.
[0305] Then, the selection unit 1302 determines whether there is an unspecified reflected sound (S208). If there is an unspecified reflected sound ("Yes" in S208), the selection unit 1302 repeats the above process (S201 to S207). If there is no unspecified reflected sound ("No" in S208), the selection unit 1302 ends the process.
[0306] This selection process can be performed on all reflected sounds generated in the parsing process, or only on reflected sounds with high evaluation values.
[0307] [Details on the threshold storage method] The threshold data for this embodiment is stored in the memory 1404 of the sound signal processing apparatus 1001. The stored threshold data can be of any form and type. When multiple forms and types of thresholds are stored, it is possible to determine which form and type of threshold to use for the selection processing of reflected sounds during the selection process. The method for determining which threshold data to use for the selection process will be described later.
[0308] Furthermore, threshold data of multiple forms and types can be combined and stored. The combined threshold data can also be read from the spatial information management units 1201 and 1211 and set as the threshold used in the selection process. Alternatively, threshold data stored in the memory 1404 can be stored in the spatial information management units 1201 and 1211.
[0309] Threshold data can also be stored, for example, as thresholds at various time differences, to depict... Figure 12C The threshold lines shown in [Example 1] and [Example 2] are examples of this.
[0310] In addition, threshold data can also be used as follows Figure 11 The diagram shows a table data storage system that establishes a correspondence between the threshold and the time difference (T). That is, the threshold data can also be stored as table data indexed by the time difference (T). Of course, Figure 11 The threshold shown is an example; the threshold is not limited to... Figure 11 Examples are provided. Alternatively, instead of storing the threshold itself, one can approximate the threshold with a function that takes the time difference (T) as a variable, and store the coefficients of that function. Furthermore, multiple approximations can be combined and stored.
[0311] For example, the time difference (T) can be set as timeDiff, and the threshold can be set as gainThresh, and the threshold data can be represented by the following formula.
[0312] [Mathematical Expression 1] The threshold is defined solely by the time range in which the priority effect occurs. When the time difference is outside this time range (a value less than 1 ms or greater than 40 ms in the above formula), a judgment based on gainThresh may not be performed, and the judgment may be made solely by the threshold representing the minimum volume reproduced in the virtual space, as described later.
[0313] Through experiments, the inventors have determined that, within the timeframe in which the preference effect occurs, the threshold is preferably approximated by an upwardly convex function. The above formula is an example of an approximation generated based on this experiment.
[0314] The memory 1404 may also store information related to the relational formulas representing the relationship between time difference (T) and threshold. That is, it may also store formulas that take time difference (T) as a variable. The threshold for each time difference (T) may also be approximated by a straight line or curve, and parameters representing the geometric shape of the straight line or curve may be stored. For example, if the geometric shape is a straight line, the starting point and slope of the straight line may also be stored.
[0315] Furthermore, the type and format of threshold data can be set and stored according to each property of the direct tone. Additionally, parameters used to adjust the threshold based on the property of the direct tone and for selection processing can be stored. The process of adjusting the threshold based on the property of the direct tone and for selection processing will be described later as a variation of the threshold setting method.
[0316] As an example of storing multiple threshold data combinations, it can also be like... Figure 12C As shown in [Example 3], the larger of the masking threshold and the echo detection limit threshold is stored for each time difference (T). Alternatively, as... Figure 12C As shown in Example 4, the value of the larger of the minimum volume and the echo detection limit threshold that are reproduced in the virtual space is stored for each time difference (T).
[0317] The combination of multiple types of threshold data is not limited to this. For example, information on the maximum value can also be stored in multiple threshold data for each time difference (T).
[0318] Furthermore, in the above, the information related to the threshold has time items as a one-dimensional index. The information related to the threshold can also have a two-dimensional or three-dimensional index that also includes variables related to the direction of arrival.
[0319] Figure 15 It is a graph showing the relationship between the direction of the direct tone, the direction of the reflected tone, the time difference, and the threshold. For example, as... Figure 15 As shown, thresholds can also be pre-calculated based on the relationship between the direction of the direct tone (θ), the direction of the reflected tone (γ), the time difference (T), and the volume ratio (L).
[0320] The direction of the direct tone (θ) corresponds to the angle relative to the direction in which the direct tone arrives at the listener. The direction of the reflected tone (γ) corresponds to the angle relative to the direction in which the reflected tone arrives at the listener. Here, the direction the listener is facing is set to 0 degrees. The time difference (T) corresponds to the difference between the arrival time of the direct tone and the arrival time of the reflected tone towards the listening position. The volume ratio (L) corresponds to the volume ratio of the volume of the direct tone arrival to the volume of the reflected tone arrival.
[0321] certainly, Figure 15 The threshold shown is an example; the threshold is not limited to... Figure 15 Examples. Furthermore, in Figure 15 The example primarily illustrates the threshold when the angle (θ) of the direction of arrival of the direct tone is 0 degrees. However, the threshold for cases where the direction of arrival of the direct tone (θ) is other than 0 degrees is also stored in memory 1404.
[0322] Furthermore, in the above, the threshold is stored as an arrangement where the direction of the direct tone (θ) (more specifically, the angle (θ) of the direction of arrival of the direct tone) and the direction of the reflected tone (γ) (more specifically, the angle (γ) of the direction of arrival of the reflected tone) are treated as independent variables or indices. However, the angle (θ) of the direction of arrival of the direct tone and the angle (γ) of the direction of arrival of the reflected tone may also not be used as independent variables.
[0323] For example, the angle difference between the direction of arrival of the direct tone (θ) and the direction of arrival of the reflected tone (γ) can also be used. This angle difference corresponds to the angle formed by the direction of arrival of the direct tone and the direction of arrival of the reflected tone, and can also be expressed as the arrival angles of the direct tone and the reflected tone.
[0324] Figure 16 This is a graph representing the relationship between angular difference, time difference, and threshold. For example, it can also be like... Figure 16 The example shown stores a pre-calculated threshold, using the angle difference (Φ) between the angle of arrival of the direct sound (θ) and the angle of arrival of the reflected sound (γ) as a variable. Of course, Figure 16 The threshold shown is an example; the threshold is not limited to... Figure 16 Examples.
[0325] exist Figure 16 In the example, the number of variables used in deriving the threshold can be reduced. Therefore, the number of thresholds stored in memory 1404 can be reduced. Consequently, the amount of data stored in memory 1404 can be reduced.
[0326] Furthermore, when using the angle difference (Φ) between the angle (θ) of the direction of arrival of the direct sound and the angle (γ) of the direction of arrival of the reflected sound, the threshold data can also be stored in a two-dimensional arrangement. Additionally, in the selection process, a three-dimensional arrangement can be used to calculate the difference between the angle (θ) of the direction of arrival of the direct sound and the angle (γ) of the direction of arrival of the reflected sound.
[0327] The method of selecting reflected sounds using a threshold corresponding to the direction of arrival will be described later.
[0328] [First variation of the threshold setting method] exist Figure 12A , Figure 12B and Figure 12C In the example, multiple forms and types of thresholds can also be stored in the spatial information management units 1201 and 1211. Furthermore, it can be determined which form and type of threshold among the multiple forms and types will be used for the selection processing of reflected sounds. Specifically, it can also be done as follows: Figure 12C As shown in Example 3, the highest threshold is used when the time difference (T) corresponds to the arrival time of the reflected sound.
[0329] Alternatively, as shown in Example 4, a masking threshold, an echo detection limit threshold, and a threshold representing the minimum volume reproduced in the virtual space can be stored. Furthermore, the highest threshold can be used for the time difference (T) corresponding to the arrival time of the reflected sound.
[0330] [Second variation of the threshold setting method] As another example of a method for setting a threshold, a method for setting a threshold based on the properties of a direct tone will be explained.
[0331] Figure 17 It means Figure 7 A block diagram of another configuration example of the rendering unit 1300 shown. Figure 17 The rendering unit 1300 and Figure 7 The rendering unit 1300 differs from the one in that it includes a threshold adjustment unit 1304. The description other than the threshold adjustment unit 1304 is the same as that in... Figure 7 The content described in the previous section is the same, so it is omitted.
[0332] The threshold adjustment unit 1304 selects a threshold to be used by the selection unit 1302 from the threshold data based on information representing the nature of the sound signal. Alternatively, the threshold adjustment unit 1304 may also adjust the threshold included in the threshold data based on information representing the nature of the sound signal.
[0333] Information indicating the nature of the sound signal can also be included in the input signal. Furthermore, the threshold adjustment unit 1304 can also obtain information indicating the nature of the sound signal from the input signal. Alternatively, the analysis unit 1301 can analyze the sound signal contained in the received input signal, derive the nature of the sound signal, and output information indicating the nature of the sound signal to the threshold adjustment unit 1304.
[0334] Information representing the nature of a sound signal can be obtained either before rendering begins or at any time during rendering.
[0335] Furthermore, the threshold adjustment unit 1304 may not be included in the sound signal processing device 1001, or it may function as a threshold adjustment unit 1304 in another communication device. In this case, the parsing unit 1301 or the selection unit 1302 may also obtain information representing the nature of the sound signal, threshold data corresponding to the nature, or information for adjusting the threshold data according to the nature from other communication devices via the communication IF 1403.
[0336] Figure 18 This is a flowchart representing another example of the selection process. Figure 19 This is another flowchart illustrating the selection process. Figure 18 and Figure 19 In this context, a threshold is set based on the properties of the direct sound. Specifically, in... Figure 18 In this process, the threshold adjustment unit 1304 determines the threshold from the threshold data based on the time difference (T) and the properties of the sound signal. Figure 19 In the process, the threshold adjustment unit 1304 adjusts the threshold determined from the threshold data based on the time difference (T) based on the properties of the sound signal.
[0337] The actions in each example are explained below. Additionally, regarding... Figure 14 The examples commonly use omitting explanations.
[0338] First, let me explain in Figure 18 The following is an example of the processing. Here, threshold data is pre-stored in memory 1404 according to each property of the direct tone. Thus, multiple threshold data corresponding to multiple properties are pre-stored in memory 1404. Furthermore, the threshold adjustment unit 1304 determines the threshold data to be used in the selection processing of the reflected tone from the multiple threshold data.
[0339] For example, the threshold adjustment unit 1304 obtains the properties of the direct tone based on the input signal (S211). The threshold adjustment unit 1304 may also obtain the properties of the direct tone that are associated with the input signal. Next, the threshold adjustment unit 1304 determines a threshold corresponding to the time difference (T) and the properties of the direct tone (S212).
[0340] In addition, such as Figure 19 As shown, the threshold adjustment unit 1304 can also adjust the threshold determined by the selection unit 1302 based on the properties of the direct tone (S221).
[0341] In any case, the input signal may include information representing the nature of the sound signal, information for adjusting the threshold according to the nature of the sound signal, or both. The threshold adjustment unit 1304 may also use one or both of these to adjust the threshold.
[0342] Furthermore, information indicating the nature of the sound signal, information used to adjust the threshold, or both, can be transmitted via an input signal different from the input signal containing the sound signal. In this case, information relating to an input signal different from the input signal can also be included in the input signal containing the sound signal, or the information relating to the input signal different from the input signal can be stored in memory 1404 along with information about the threshold.
[0343] exist Figure 18 and Figure 19 In the example, the threshold used in selecting the reflected tone is set based on the properties of the direct tone, i.e., the properties of the sound signal. This can be achieved as follows: Figure 18 Using pre-defined threshold data for each property, it is also possible to... Figure 19 In this way, the threshold can be adjusted based on the properties of the sound signal. Alternatively, the parameters of the threshold data can also be adjusted based on the properties of the sound signal.
[0344] Furthermore, the operation performed by the threshold adjustment unit 1304 can also be performed by the analysis unit 1301 or the selection unit 1302. For example, the analysis unit 1301 may obtain the properties of the sound signal. Alternatively, the selection unit 1302 may set the threshold based on the properties of the sound signal.
[0345] Next, the relationship between the properties of the sound signal and the threshold will be explained.
[0346] Two short sounds arriving at a listener's ear consecutively, if the time interval between them is sufficiently short, are perceived as a single sound. This phenomenon is called the priority effect. It is known that the priority effect occurs only for discontinuous, i.e., transient sounds (Non-Patent Document 1). Therefore, when the sound signal represents a stationary tone, the echo detection limit can be set lower compared to when the sound signal represents a non-stationary tone.
[0347] That is, based on the characteristics of this priority effect, for example, when the direct tone is a stable sound, the threshold is set to be relatively small. Alternatively, the higher the stability, the smaller the threshold can be set.
[0348] An example of processing when the sound signal is stationary will be explained. First, the threshold adjustment unit 1304 or the analysis unit 1301 determines the stationarity based on the amount of change in the frequency components of the sound signal over time. For example, if the amount of change is small, the stationarity is determined to be high. Conversely, if the amount of change is large, the stationarity is determined to be low. The result of the determination can be used to set a flag representing the level of stationarity, or a parameter representing stationarity can be set based on the amount of change.
[0349] Next, the threshold adjustment unit 1304 may adjust the threshold data or threshold based on information indicating the stability of the sound signal, such as a flag or parameter, and set the adjusted threshold data or threshold as the threshold data or threshold used in the selection unit 1302.
[0350] Alternatively, parameters for setting threshold data based on information representing the stability of the direct tone can be pre-stored in the memory 1404. In this case, the threshold adjustment unit 1304 can also determine the stability of the sound signal and set the threshold data used in selecting the reflected tone based on the information and parameters representing the stability.
[0351] Alternatively, multiple parameters of the threshold data may be pre-stored in the memory 1404 corresponding to multiple patterns of the direct tone's stability. In this case, the threshold adjustment unit 1304 may also determine the stability of the sound signal, select parameters of the threshold data based on the pattern of the direct tone's stability, and set the threshold data used in the selection of reflected tones based on the parameters of the threshold data.
[0352] In addition, the stability of a sound signal can be determined based on the amount of change in the frequency components of the sound signal each time a sound signal is input.
[0353] Alternatively, the stationarity of the sound signal can be determined based on information representing stationarity that has been pre-associated with the sound signal. That is, information representing the stationarity of the sound signal can be pre-associated with the sound signal and stored in the memory 1404. The parsing unit 1301 can also obtain the information representing stationarity associated with the sound signal each time an input sound signal is received. Furthermore, the threshold adjustment unit 1304 can adjust the threshold based on the information representing stationarity associated with the sound signal.
[0354] As another example of setting a threshold based on the properties of the sound signal, the application range of the echo detection limit can be set shorter when the sound signal represents a shorter sound (such as a click) compared to when the sound signal represents a longer sound. This processing is based on the characteristics of the priority effect.
[0355] It is known that, through the priority effect, two short sounds arriving consecutively at a listener's ear are perceived as a single sound if the time interval between them is sufficiently short. The upper limit of this time interval depends on the length of the sound. For example, the upper limit of this time interval is approximately 5 ms for a click sound, and sometimes as high as 40 ms for complex sounds such as human voices or music (Non-Patent Document 1).
[0356] Based on this priority effect, for example, in the case of sounds with shorter direct durations, a shorter duration threshold is set. Furthermore, the shorter the direct duration, the shorter the duration threshold is set.
[0357] Setting a shorter time threshold means setting a threshold corresponding to the echo detection limit based on the priority effect characteristics within a range where the time difference (T) between the direct tone and the reflected tone is small. Outside this range, no threshold corresponding to the echo detection limit based on the priority effect characteristics is set. That is, outside this range, the threshold is small. Therefore, setting a shorter time threshold for shorter sounds corresponds to setting a smaller threshold for shorter sounds.
[0358] As another example of setting a threshold based on the nature of the direct sound, the threshold can be set lower when the direct sound is an intermittent sound (speech, etc.) compared to when the direct sound is a continuous sound (music, etc.).
[0359] For example, when the direct sound corresponds to speech, there are repeated vocal and non-vocal parts, and as a masking effect, only the aftermasking effect occurs in the non-vocal part. On the other hand, when the direct sound is a continuous sound like musical content, both the aftermasking effect and the simultaneous masking effect based on the sound produced at that time occur. Therefore, the comprehensive masking effect is higher in the case of music, etc., than in the case of speech, etc.
[0360] Based on the masking effect characteristics described above, the threshold can be set higher in the case of music, etc., compared to the case of speech, etc. Conversely, the threshold can be set lower in the case of speech, etc., compared to the case of music, etc. That is, the threshold can be set lower when there are many interruptions in the direct sound.
[0361] As mentioned above, information indicating the properties of a direct tone can also include information about its smoothness, discontinuity, and duration. Furthermore, information indicating the properties of a direct tone can be any combination of these characteristics. Additionally, information indicating the properties of a direct tone can be information about the temporal variation of any one of these characteristics, or information about the temporal variation of any combination thereof. In other words, information indicating the properties of a direct tone can also be information about the temporal variation of the direct tone.
[0362] For example, as shown in the explanation of stationarity determination, information representing the properties of a direct tone can also be time-series data of frequency characteristics. Here, frequency characteristics can also be represented in conventional forms such as gain values for each frequency band, Fourier series of the signal over the time axis, or LPC coefficients or cepstral coefficients used to calculate the frequency envelope.
[0363] Furthermore, information representing the properties of a direct tone can also be presented as information representing the discontinuity of the direct tone, listing multiple sets of information (the approximate shape of the amplitude envelope) of the signal's amplitude stability over a time series. Here, the amplitude value can also be expressed as a ratio relative to a reference volume.
[0364] Furthermore, the information representing the properties of a direct tone can also be information related to the frequency characteristics of the direct tone. For example, the information representing the properties of a direct tone can also be information representing the stability of the frequency characteristics of the direct tone. Specifically, the information representing the properties of a direct tone can also be information that lists multiple groups of frequency characteristics of the signal during a period of small frequency variation in a time series (approximate shape of a spectrum). Here, the volume used as a reference for the aforementioned frequency characteristics can also be the aforementioned reference volume.
[0365] For example, information indicating the temporal variation of a direct tone is information representing the envelope of the direct tone. Information indicating the temporal variation of a direct tone can also be found in... Figure 12C The “minimum audible limit” described in [Example 4] is used when the threshold is a threshold. The signal compared to the minimum audible limit is the volume of the reflected sound.
[0366] The volume of the reflected sound is obtained through geometric calculations based on the positions of the sound source, the listener, and the reflecting object. Specifically, a reference volume of the reflected sound is obtained relative to a reference volume of the sound source. By using information about changes in the volume of the sound source as information representing the properties of the direct sound, the reference volume of the reflected sound is increased or decreased, allowing for an accurate determination of the volume of the reflected sound at any given moment. This is because changes in the volume of the sound source are reflected in changes in the volume of the reflected sound.
[0367] After adjusting the volume of the reflected sound, by comparing the volume of the reflected sound with the threshold, it is possible to more accurately and appropriately select the reflected sound that is audibly desired.
[0368] Of course, the same result can be obtained by adjusting the threshold based on the reciprocal of the change in the volume of the sound source, without adjusting the reference volume of the reflected sound, and then comparing the adjusted threshold with the reference volume of the reflected sound. That is, the reference volume of the reflected sound can be adjusted using information about the change in the volume of the sound source, and the threshold can also be adjusted using the same information. The adjustment of the reference volume of the reflected sound and the adjustment of the threshold are mutually corresponding.
[0369] Depending on the surface composition of the object reflecting the sound, the reflectivity of the sound (and the attenuation rate of the reflected sound) varies in each frequency band. Therefore, as described later, the reflectivity (attenuation rate) of the sound can also be associated with each frequency band for the object reflecting the sound. Based on this reflectivity information and the information from the spectrogram, it is possible to more accurately determine whether to select the reflected sound. For example, the following process can be performed.
[0370] Specifically, for example, information from the spectrogram indicates that frequency components in the high-frequency band are more dominant than those in the low-frequency band within a certain time interval. Additionally, for example, information from the reflectivity of sound indicates that the reflectivity of frequency components in the high-frequency band is extremely low compared to that in the low-frequency band.
[0371] In this case, even if the signal amplitude of the sound source is large on the time axis, the volume of the reflected sound is reduced by multiplying the frequency components represented by the information in the spectrogram with the attenuation rate of each frequency band represented by the information in the reflectivity, and the reflected sound may not be selected.
[0372] As mentioned above, information representing the properties of a direct tone can also be information representing the temporal variation of the direct tone. For example, information representing the properties of a direct tone can also represent values obtained by analyzing the direct tone over a predetermined time period.
[0373] Specifically, information representing the properties of a direct tone can also be obtained by calculating the average energy or average amplitude of the direct tone for each predetermined time length. Alternatively, information representing the properties of a direct tone can also be obtained by calculating the energy or average amplitude of the direct tone for each short time analysis length and then taking a weighted average of the energy or average amplitude for each long time analysis length that is longer than the short time analysis length.
[0374] More specifically, for example, information representing the temporal variation of the direct tone can also be obtained by calculating the energy or average amplitude of the direct tone for each pre-defined short time length (e.g., 5 ms, hereinafter referred to as an analysis frame). Alternatively, information representing the temporal variation of the direct tone can also be represented by a weighted average of the energy or average amplitude calculated over the past N-1 analysis frames.
[0375] Assuming the energy of the nth analysis frame is represented by E(n), the information I(n) representing the properties of the direct tone is obtained according to the following formula.
[0376] [Mathematical Expression 2] Here, the parameter a(i) represents the weighting coefficient. Typically, a(i) is set such that a(i) ≥ 0 and the sum of a(i) is 1. However, the method of setting a(i) is not limited to this.
[0377] Furthermore, information I(n) representing the properties of the direct tone is calculated every 5ms since the direct tone was received. That is, the temporal variation of information I(n) representing the properties of the direct tone can be calculated with low latency. Therefore, this method is suitable for applications requiring real-time performance.
[0378] Alternatively, the information I(n) representing the properties of direct sounds can be obtained using the following formula.
[0379] [Mathematical Expression 3] Here, the parameter b(i) represents the weighting coefficient. Typically, b(i) is set such that b(i) ≥ 0 and the sum of b(i) is 1. However, the method of setting b(i) is not limited to this.
[0380] In this formula, information I(n) representing the properties of the direct tone is recursively obtained. Therefore, the average energy over a long time period can be calculated with minimal computation.
[0381] Equations 1 and 2 above can be considered as filters where E(n) is the input signal and I(n) is the output signal. In this case, Equation 1 is a filter for the moving average (MA) model, and Equation 2 is a filter for the autoregressive (AR) model, both exhibiting the characteristics of low-pass filters. Alternatively, a filter for the ARMA model, which combines the two, can also be used.
[0382] Furthermore, the method for deriving information representing the temporal variation of the direct tone is not limited to the aforementioned calculation formula or filter; other known methods can also be used. As mentioned above, the information representing the temporal variation of the direct tone represents a value obtained by analyzing the direct tone over a predetermined time length. The direct tone can also be analyzed from a perspective other than average energy.
[0383] Furthermore, as mentioned above, the information representing the properties of a direct tone can also be information related to the frequency characteristics of the direct tone. This information related to the frequency characteristics of the direct tone can also be information calculated using those characteristics. For example, information related to the frequency characteristics of the direct tone can be obtained by averaging the low-frequency components of the direct tone over a predetermined analysis length to obtain the average energy of the low-frequency components.
[0384] Specifically, the low-frequency components of the direct tone are determined by applying a low-pass filter to the direct tone contained in the analysis frame length. Based on the energy or average amplitude of this low-frequency component, information representing the properties of the direct tone is derived in the same manner as in Equation 1 above.
[0385] Assume the energy of the low-frequency components in the nth analysis frame is determined by E L In the case of (n), the information I(n) representing the nature of the direct sound is obtained according to the following formula.
[0386] [Mathematical Expression 4] Here, the parameter c(i) represents the weighting coefficient. Typically, c(i) is set such that c(i) ≥ 0 and the sum of c(i) is 1. However, the method of setting c(i) is not limited to this.
[0387] Furthermore, information I(n) representing the properties of the direct tone is calculated every 5ms since the direct tone was received. That is, the temporal variation of information I(n) representing the properties of the direct tone can be calculated with low latency. Therefore, this method is suitable for applications requiring real-time performance.
[0388] In addition, similar to Equation 2, the information I(n) representing the properties of direct sounds can also be obtained according to the following formula.
[0389] [Mathematical Expression 5] Here, the parameter d(i) represents the weight coefficient. Typically, d(i) is set such that d(i) ≥ 0 and the sum of d(i) is 1. However, the method of setting d(i) is not limited to this.
[0390] In this formula, information I(n) representing the properties of the direct tone is recursively obtained. Therefore, the average energy over a long time period can be calculated with minimal computation.
[0391] Equations 3 and 4 above can be considered as filters with E(n) as the input signal and I(n) as the output signal. In this case, Equation 3 is a filter for the moving average (MA) model, and Equation 4 is a filter for the autoregressive (AR) model, both of which have the characteristics of low-pass filters. Alternatively, a filter for the ARMA model, which combines the two, can also be used.
[0392] In the above method for determining the low-frequency components of a direct tone, a filter with low-pass characteristics is used; however, the method for determining the low-frequency components of a direct tone is not limited to this. Furthermore, the method for deriving information representing the time variation of the direct tone is not limited to the aforementioned calculation formula or filter; other known methods can also be used. For example, the spectrum of a direct tone can be calculated by performing a frequency transformation on the direct tone. Furthermore, the energy or average amplitude of the low-frequency components of the spectrum can be calculated.
[0393] Furthermore, in the above, MA or AR models are used to derive information representing the temporal variation of direct tones. The coefficients of these models can be pre-set fixed values or time-varying values.
[0394] In addition, the relationship between the analysis frame length and the information update thread generation interval can also be as follows.
[0395] For example, when the analysis frame duration is TA (msec) and the information update thread generation interval is TU (msec), the value of N in Equations (1) and (3) above in the MA filter can also be approximately equal to the value given by TU / TA. Additionally, b(i) and d(i) (1≤i<N) in Equations (2) and (4) above in the AR filter can also be approximately equal to the filter's time constant, which is approximately equal to TU (msec).
[0396] The reason for this setting is that the filter is expected to converge during the information update interval.
[0397] On the other hand, in the above-described configuration, if the value of the information representing the temporal variation of the direct tone changes too drastically, I(n) can be pre-calculated. Furthermore, the pre-calculated I(n) can be applied to the selection processing of reflected tones. For example, in the processing of the frame at time t, I(t+tau) can be used. Here, tau is a value determined based on the convergence characteristics of the filter. In cases of slow convergence, the value of tau is larger compared to cases of fast convergence.
[0398] Furthermore, auditory masking (frequency masking) information calculated based on the direct tone can also be used as information representing the characteristics of the direct tone. The auditory masking information represents a threshold value for the amplitude in the frequency domain that is masked by the direct tone. It is also possible to compare the amplitude value of reflected tones in the same frequency domain with the threshold value, without selecting reflected tones with amplitude values smaller than the threshold. The amplitude value of reflected tones in the frequency domain can also be obtained by the analysis unit 1301 as information representing the characteristics of the reflected tones.
[0399] By setting the threshold used in selecting reflected sounds based on the properties of the direct sound, the reflected sounds that are audibly required can be appropriately selected, and auditory characteristics can be effectively reflected in the stereo sound reproduction system 1000. The processing of detecting the properties of the direct sound, determining the threshold based on the properties, and adjusting the threshold based on the properties can be performed either during the rendering process or before the rendering process begins.
[0400] For example, these processes can occur during virtual space creation (when the software is created), at the start of virtual space processing (when the software starts or rendering begins), or at timed intervals in information update threads that occur periodically during virtual space processing. Furthermore, virtual space creation can be timed to build the virtual space before the start of sound processing, or it can occur when virtual space information (spatial information) is acquired, or it can occur when the software acquires the information.
[0401] Here, in the information update thread, processing is performed to update the spatial information managed by the Spatial Information Management Departments 1201 and 1211.
[0402] The information update thread is responsible for tasks such as updating the position and orientation of the listener's avatar configured in the virtual space based on the position and orientation of the VR goggles worn by the listener, or updating the position of moving objects in the virtual space. This processing is provided within a processing thread that starts at a relatively low frequency of around tens of Hz.
[0403] In such low-frequency processing threads, the processing of information representing the properties of direct tones can still be performed. This is because the frequency of changes in the properties of direct tones is lower than the frequency of audio processing frames used for audio output. Therefore, the computational load of this processing can be relatively reduced. Furthermore, updating information at an unnecessarily fast frequency carries the risk of generating impulse noise. This risk can be avoided by updating information at a low frequency.
[0404] [Third variation of the threshold setting method] As another example of a method for setting the threshold, the threshold can also be set based on the computing resources (CPU power, memory resources, PC performance, or remaining battery power, etc.) used to process the reproduction of the virtual space. More specifically, the sensor 1405 of the sound signal processing device 1001 detects the amount of computing resources and sets a higher threshold when the amount of computing resources is low. As a result, the volume of more reflected sounds is lower than the threshold, thus reducing the reflected sounds that are processed by both ears and reducing the amount of computing power.
[0405] Alternatively, in cases where signal processing is performed by battery-powered devices such as smartphones or VR headsets, it may be desirable to prioritize long processing times and conserve computing resources. In such cases, the threshold can be set high without needing to monitor the amount or remaining amount of computing resources.
[0406] [Fourth variation of the threshold setting method] As another example of a method for setting a threshold, the audio signal processing device 1001 or the audio prompting device 1002 may have a threshold setting unit (not shown), allowing the administrator or listener of the virtual space to set the threshold.
[0407] For example, the listener wearing the sound prompt device 1002 could choose between an "energy-saving mode" (fewer reflected sounds and less computation) and a "high-performance mode" (more reflected sounds and more computation). Alternatively, the administrator of the stereo sound reproduction system 1000 or the producer of the stereo sound content could choose the mode. Furthermore, instead of a mode, a threshold or threshold data could be directly selected.
[0408] [The first variation of the rendering department's actions] Figure 20 This is a flowchart illustrating a first variation of the operation of the sound signal processing device 1001. Figure 20 The diagram shows the processing mainly performed by the rendering unit 1300 of the sound signal processing device 1001. In this modified example, volume compensation processing is added to the operation of the rendering unit 1300.
[0409] For example, the analysis unit 1301 acquires data (input signal) (S301). Next, the analysis unit 1301 analyzes the data (S302). Next, the selection unit 1302 determines whether to select reflected sounds based on the analysis results (S303). Next, the reproduction unit 1303 performs volume compensation processing based on the unselected reflected sounds (S304). Next, the reproduction unit 1303 performs audio processing on both direct and reflected sounds (S305). Finally, the reproduction unit 1303 outputs both direct and reflected sounds as audio (S306).
[0410] In the processes described above (S301 to S306), the processes other than the volume compensation process (S304) are common to the other examples described above, so their descriptions are omitted.
[0411] Volume compensation processing is performed for reflected sounds that were not selected in the selection process. For example, by not selecting reflected sounds in the selection process, a lack of volume perception occurs. Volume compensation processing can suppress the unpleasantness that accompanies this lack of volume perception. Two methods are disclosed as examples of methods for compensating for volume perception. Either method can be used.
[0412] First, the method of compensating for the sense of volume by increasing the volume of the direct tone will be explained. The reproduction unit 1303 generates the direct tone by increasing its volume by an amount corresponding to the volume of the unselected reflected tone. As a result, the sense of volume lost due to the lack of reflected tone generation is compensated.
[0413] When the reproduction unit 1303 increases the volume, it can also increase the volume according to the frequency characteristics of the reflected sound, one frequency component at a time. To enable this process, a predetermined attenuation rate for the volume of the reflected sound can be assigned to each frequency band. Thus, the frequency characteristics of the reflected sound can be derived.
[0414] Next, a method for compensating for the sense of volume by synthesizing reflected sounds into direct sounds will be explained. In this method, the reproduction unit 1303 adds unselected reflected sounds to the direct sounds to generate a direct sound, thereby compensating for the sense of volume caused by the lack of generated reflected sounds. The generated direct sound reflects the volume (amplitude), frequency, and delay of the unselected reflected sounds.
[0415] In the case of increasing the volume of the direct tone, the computational workload of the compensation process is very small, but only the volume is compensated. In the case of synthesizing the reflected tone into the direct tone, the computational workload of the compensation process is larger compared to the method of increasing the volume of the direct tone, but the characteristics of the reflected tone are compensated more accurately.
[0416] In both cases, no reflected tones are generated, only direct tones, thus reducing the overall computational load. In particular, the computational load required for binaural processing, which includes convolutional HRTF processing, is reduced, resulting in a significant reduction in overall computational load. This is because the computational load required for binaural processing is far greater than that required for the aforementioned compensation processing.
[0417] In addition, if the reason for not selecting the reflected sound is that the volume of the reflected sound is lower than the masking threshold, since the sense of volume will not be lost, the reflected sound can be removed without compensation.
[0418] [Second variation of the rendering department's actions] Figure 21 This is a flowchart illustrating a second variation of the operation of the sound signal processing device 1001. Figure 21The text primarily describes the processing performed by the rendering unit 1300 of the sound signal processing device 1001. In this modified example, the operation of the rendering unit 1300 includes left and right volume difference adjustment processing.
[0419] For example, the analysis unit 1301 analyzes the input signal (S401). Next, the analysis unit 1301 detects the direction of sound arrival (S402). Next, the selection unit 1302 adjusts the volume difference of the sound perceived by the left and right ears (S403). Furthermore, the selection unit 1302 adjusts the time difference (delay) of the sound arrival perceived by the left and right ears (S404). Based on the adjusted sound information, the selection unit 1302 determines whether to select the reflected sound (S405).
[0420] In the above-described processes (S401 to S405), the processes other than the left and right volume difference adjustment process (S403) and the delay adjustment process (S404) are common to the other examples described above, so their descriptions are omitted.
[0421] Figure 22 This is a diagram illustrating the configuration of avatars, sound source objects, and obstacle objects. For example, in the case where the listener's facing direction is 0 degrees, such as... Figure 22 As shown, if the polarity (e.g., positive or negative) of the direction of arrival of the direct tone (θ) and the direction of arrival of the reflected tone (γ) (the direction of the reflected tone (γ)) is different, the volume difference generated between the two ears is corrected.
[0422] Specifically, when the polarities of θ and γ are different, the ear that primarily (first) perceives the sound in the direct tone and the reflected tone is different. In this case, the selection unit 1302 performs a left-right volume difference adjustment process (S403) to adjust the volume of the direct tone according to the position of the ear that primarily perceives the reflected tone. For example, the selection unit 1302 attenuates the volume of the direct tone when it reaches the listener by multiplying the volume by (1.0 - 0.3sin(θ)) (0 ≤ θ ≤ 180°).
[0423] The selection unit 1302 determines whether to select the reflected sound by calculating the volume ratio of the corrected direct sound volume to the reflected sound volume as described above, and comparing the calculated volume ratio with a threshold. This corrects the volume difference between the two ears, more accurately derives the volume of the direct sound that affects the reflected sound, and more accurately determines whether to select the reflected sound.
[0424] In addition to adjusting the left and right volume difference (S403), the selection unit 1302 can also perform a delay adjustment process (S404) to match the position of the ear that perceives the reflected sound, thus delaying the arrival time of the direct sound. Specifically, the selection unit 1302 can also delay the arrival time of the direct sound by adding (a(sinθ+θ) / c) ms (where a is the radius of the head and c is the speed of sound) to the arrival time of the direct sound.
[0425] [The third variation of the rendering department's actions] The method for setting a threshold corresponding to the direction of arrival is explained.
[0426] Figure 23 This is another flowchart illustrating the selection process. Regarding... Figure 14 Examples of this type of writing commonly involve omitting explanatory notes. Figure 23 In the example, the selection unit 1302 uses a threshold corresponding to the direction of arrival to select the reflected sound.
[0427] Specifically, the selection unit 1302 calculates the direct sound arrival path (pd), the reflected sound arrival path (pr), and the orientation information D1 of the avatar, using the orientation of the avatar as a reference, the direct sound arrival direction (θ) and the reflected sound arrival direction (γ) (the direction of the reflected sound (γ)). That is, the selection unit 1302 detects the direct sound arrival direction (θ) and the reflected sound arrival direction (γ) (S231). The orientation of the avatar corresponds to the orientation of the listener. The orientation information D1 of the avatar may also be included in the input signal.
[0428] The selection unit 1302 uses three indices, including the direct direction of arrival (θ) and the reflected direction of arrival (γ), as well as the time difference (T), according to... Figure 15 The three-dimensional arrangement shown determines the threshold used in the selection process (S232).
[0429] As an example, to illustrate in such Figure 22 The method shown describes the threshold setting method used in the selection process when an avatar, a sound source object, and an obstacle object are configured.
[0430] The position information of the avatar, sound source object, and obstacle object, as well as the orientation information D1 of the avatar, are obtained from the input signal. Using this position information and orientation information D1, the direction of the direct sound (θ) and the direction of the sound image of the reflected sound are calculated when the orientation of the avatar is set to 0 degrees. Figure 22 In this case, the direction of the direct sound (θ) is about 20 degrees, and the direction of the sound image of the reflected sound (γ) is about 265 degrees (-95 degrees).
[0431] Next, refer to Figure 15The threshold data shown is stored in a three-dimensional arrangement. The threshold is determined from the arrangement region corresponding to the values of the two directions (θ) and (γ) and the value of the time difference (T) calculated by the analysis unit 1301. Even if there is no index corresponding to the calculated values of (θ), (γ), and (T), the threshold corresponding to the nearest index can be determined.
[0432] Alternatively, the threshold can be determined by interpolation, extrapolation, or other processing based on one or more thresholds corresponding to one or more indices close to the calculated values of (θ), (γ), and (T). For example, the threshold corresponding to (20 degrees, 265 degrees, T) can be determined based on four thresholds corresponding to the four indices (0 degrees, 225 degrees, T), (0 degrees, 270 degrees, T), (45 degrees, 225 degrees, T), and (45 degrees, 270 degrees, T).
[0433] The selection process based on the difference between the angle (θ) of the direction of arrival of the direct sound and the angle (γ) of the direction of arrival of the reflected sound is explained.
[0434] For example, it can also be pre-made and set up as follows: Figure 16 The threshold data shown is obtained by arranging the angle difference (Φ) between the direction of arrival of the direct tone (θ) and the direction of arrival of the reflected tone (γ) and the time difference (T) as a two-dimensional index. In this case, the angle difference (Φ) and the time difference (T) are referenced in the selection process. Alternatively, the angle difference (Φ) between the angle of arrival of the direct tone (θ) and the angle of arrival of the reflected tone (γ) can be calculated in the selection process, and the calculated angle difference (Φ) can be used to determine the threshold.
[0435] Alternatively, a threshold data can be set to be used as an index for arranging the combination of the angle difference (Φ), the direction of arrival of the direct sound (θ), and the time difference (T), or the combination of the angle difference (Φ), the direction of arrival of the reflected sound (γ), and the time difference (T).
[0436] Alternatively, it can be set as follows: Figure 15 The threshold data shown is obtained by arranging the values of (θ), (γ), and (T) as a three-dimensional index.
[0437] [The fourth variation of the rendering department's actions] The processing performed by the analysis unit 1301, selection unit 1302 and reproduction unit 1303 described above can also be performed as pipeline processing as described in Patent Document 3.
[0438] Figure 24 This is a block diagram illustrating a configuration example for pipeline processing in the rendering unit 1300.
[0439] Figure 24The rendering unit 1300 includes a reverberation processing unit 1311, an initial reflection processing unit 1312, a distance attenuation processing unit 1313, a selection unit 1314, a generation unit 1315, and a binocular processing unit 1316. These multiple components can also be... Figure 7 The rendering unit 1300 shown can be composed of multiple constituent elements, or it can be made of... Figure 5 It constitutes at least a portion of the multiple constituent elements of the sound signal processing device 1001 shown.
[0440] Pipeline processing refers to dividing the processing used to impart sound effects into multiple processes and executing these processes sequentially. Each process may perform signal processing on the audio signal or generate parameters used in signal processing.
[0441] The rendering unit 1300 can also perform reverberation processing, initial reflection processing, distance attenuation processing, and binaural processing as a pipeline process. However, these are just examples; pipeline processing may include other processing methods, or it may exclude some of them. For example, pipeline processing may also include diffraction processing and occlusion processing. Furthermore, reverberation processing, for example, can be omitted if it is not needed.
[0442] Furthermore, each process can be represented as a stage. Additionally, the results of each process, and the generated sound signals such as reflected sounds, can be represented as rendering elements. The multiple stages in the pipeline processing and their order are not limited to... Figure 24 The example shown.
[0443] Here, the parameters used in the selection process (arrival path, arrival time, and volume ratio related to direct and reflected tones) can also be calculated in one of the multiple stages used to generate the render item. That is, the parameters used in the selection of reflected tones are calculated in a part of the pipeline processing used to generate the render item. Alternatively, not all stages may be performed by the rendering unit 1300. For example, some stages may be omitted, or they may be performed outside of the rendering unit 1300.
[0444] The reverberation processing, initial reflection processing, distance attenuation processing, selection processing, generation processing, and binaural processing that may be included as stages in pipeline processing are described. Within each stage, metadata contained in the input signal can also be parsed to calculate the parameters used in generating the reflected tones.
[0445] In reverberation processing, the reverberation processing unit 1311 generates parameters used in the generation of a sound signal representing reverberant sound. Reverberant sound refers to the sound that arrives at the listener as reverberation after the direct tone. As an example, reverberant sound is the sound that arrives at the listener after a relatively late stage (e.g., from the arrival of the direct tone to about one hundred and several tens of ms) following the arrival of the initial reflected sound, as described later. It undergoes more reflections (e.g., dozens of times) than the initial reflected sound.
[0446] The reverberation processing unit 1311 refers to the sound signal and spatial information contained in the input signal and calculates the reverberation sound using a predetermined function prepared in advance as a function to generate the reverberation sound.
[0447] The reverberation processing unit 1311 can also apply known reverberation generation methods to the sound signal contained in the input signal to generate reverberation. An example of a known reverberation generation method is the Schroeder method, but known reverberation generation methods are not limited to the Schroeder method. Furthermore, the reverberation processing unit 1311 uses the shape and acoustic characteristics of the sound reproduction space represented by spatial information in the application of known reverberation generation methods. Therefore, the reverberation processing unit 1311 can calculate the parameters used to generate the reverberation.
[0448] In the initial reflection processing, the initial reflection processing unit 1312 calculates parameters for generating the initial reflection tone based on spatial information. The initial reflection tone is the reflection tone that reaches the listener after more than one reflection in a relatively early stage (e.g., about tens of milliseconds from the arrival of the direct tone) after the direct tone reaches the listener from the sound source object.
[0449] The initial reflection processing unit 1312 calculates, for example, the path of the reflected sound from the sound source object to the listener via the reflection object, by referring to the sound signal and metadata. For example, the shape of the three-dimensional sound field (space), the size of the three-dimensional sound field, the position of the reflecting object such as the structure, and the reflectivity of the reflecting object can also be used in the path calculation.
[0450] Furthermore, the initial reflection processing unit 1312 may also calculate the path of the direct tone. This path information may also be used as a parameter for the initial reflection processing unit 1312 to generate the initial reflected tone, or as a parameter for the selection unit 1314 to select the reflected tone.
[0451] In distance attenuation processing, the distance attenuation processing unit 1313 calculates the volume of the direct tone and the reflected tone reaching the listener based on the path lengths of the direct tone and the reflected tone. The volume of the direct tone and the reflected tone reaching the listener is attenuated proportionally to the distance of the path to the listener (inversely proportional to the distance) relative to the volume of the sound source. Therefore, the distance attenuation processing unit 1313 can calculate the volume of the direct tone by dividing the volume of the sound source by the path length of the direct tone, and can calculate the volume of the reflected tone by dividing the volume of the sound source by the path length of the reflected tone.
[0452] In the selection process, the selection unit 1314 selects the object to be reflected based on parameters calculated prior to the selection process. A selection method of this disclosure may also be used in the selection of the object to be reflected.
[0453] Selection processing can be performed on all reflected sounds, or, as described above, on only reflected sounds with high evaluation values. That is, reflected sounds with low evaluation values are automatically disqualified without any selection processing. For example, reflected sounds with very low volume can also be considered as having low evaluation values and are therefore disqualified.
[0454] Furthermore, for example, selection processing can be applied to all reflected sounds. Also, the evaluation value of the selected reflected sounds in the selection process can be determined, and reflected sounds with low evaluation values can be re-selected as not selected.
[0455] The selection and evaluation processes can be executed independently or in combination. When the selection and evaluation processes are executed in combination, one of the two processes can be executed first.
[0456] In the generation process, the generation unit 1315 generates direct tones and reflected tones. For example, the generation unit 1315 generates a direct tone based on the sound signal contained in the input signal, according to the arrival time and volume of the direct tone. Furthermore, regarding the reflected tone selected in the selection process, the generation unit 1315 generates a reflected tone based on the sound signal contained in the input signal, according to the arrival time and volume of the reflected tone.
[0457] In binaural processing, the binaural processing unit 1316 performs signal processing to make the sound signal of the direct tone perceived as sound arriving at the listener from the direction of the sound source object. Furthermore, the binaural processing unit 1316 performs signal processing to make the reflected tone selected by the selection unit 1314 perceived as sound arriving at the listener from the reflecting object.
[0458] For example, the binaural processing unit 1316 performs HRIR DB processing based on the position and orientation of the listener in the sound space, so that the sound reaches the listener from the position of the sound source object or the position of the obstacle object.
[0459] Additionally, HRIR (Head-Related Impulse Responses) describes the response characteristics when a single impulse is generated. Specifically, HRIR is the response characteristic obtained by transforming the head-related transfer function from its frequency domain representation to its time domain representation using a Fourier transform. This head-related transfer function represents the changes in sound produced by surrounding objects, including the auricle, head, and shoulders, as a transfer function. The HRIR DB is a database containing such information.
[0460] Furthermore, the position and orientation of the listener in the sound space can be, for example, the position and orientation of a virtual listener in a virtual sound space. Alternatively, the position and orientation of the virtual listener in the virtual sound space can change in accordance with the movement of the listener's head. Furthermore, the position and orientation of the virtual listener in the virtual sound space can also be determined based on information obtained from sensor 1405.
[0461] The program, spatial information, HRIR DB, threshold data or other parameters used in the above processing are obtained from the memory 1404 of the sound signal processing device 1001 or from outside the sound signal processing device 1001.
[0462] Furthermore, pipeline processing may also include other processing. Additionally, the rendering unit 1300 may include a processing unit (not shown) for performing other processing included in pipeline processing. For example, the rendering unit 1300 may also include a diffraction processing unit and an occlusion processing unit.
[0463] The diffraction processing unit performs processing to generate a sound signal representing a sound containing diffracted tones, which are caused by an obstacle object in a three-dimensional sound field (space) located between the listener and the sound source object. A diffracted tone is a sound that reaches the listener from the sound source object by bypassing an obstacle object when such an obstacle object exists between the sound source object and the listener.
[0464] The diffraction processing unit, for example, refers to the sound signal and metadata to calculate the path of the diffracted sound from the sound source object, bypassing the obstacle object, to reach the listener, and generates the diffracted sound based on the path. In the path calculation, the positions of the sound source object, the listener, and the obstacle object in the three-dimensional sound field (space), as well as the shape and size of the obstacle object, can also be used.
[0465] When a sound source object exists on the opposite side of an obstacle object, the occlusion processing unit generates a sound signal of the sound leaking from the sound source object through the obstacle object based on spatial information and information such as the material of the obstacle object.
[0466] [Example of a sound source object] In the above, the positional information assigned to the sound source object represents the position of the sound source object as a "point" in the virtual space. That is, in the above, the sound source is defined as a "point sound source".
[0467] On the other hand, a sound source in virtual space can also be defined as an object with length, size, and shape, that is, a non-point sound source that extends spatially. In this case, the distance between the listener and the sound source and the direction of sound arrival are uncertain. Therefore, reflected sounds caused by such sound sources do not need to be analyzed by the analysis unit 1301, or are limited to being selected by the selection unit 1302 regardless of the analysis result. As a result, sound quality degradation that may occur due to failure to select reflected sounds can be avoided.
[0468] Alternatively, a representative point, such as the object's center of gravity, can be determined, and the processing disclosed herein can be applied assuming that sound originates from that point. In this case, the threshold can also be adjusted based on the spatial extension information of the sound source.
[0469] [Examples of direct and reflected sounds] For example, a direct sound is a sound that is not reflected by a reflecting object, while a reflected sound is a sound that is reflected by a reflecting object. A direct sound can also be a sound that reaches the listener from the sound source without being reflected by a reflecting object, and a reflected sound can also be a sound that reaches the listener from the sound source after being reflected by a reflecting object.
[0470] Furthermore, direct tone and reflected tone are not limited to the sound that reaches the listener; they can also be the sound before it reaches the listener. For example, direct tone can also be the sound output from the sound source, or in other words, the sound of the sound source.
[0471] Figure 25 This is a diagram illustrating the transmission and diffraction of sound. For example... Figure 25 As shown, sometimes the direct sound does not reach the listener because of an obstacle object between the sound source and the listener. In this case, the sound emitted from the sound source, passing through the obstacle object, and reaching the listener can also be considered as the direct sound. Furthermore, the sound emitted from the sound source, diffracted through the obstacle object, and reaching the listener can also be considered as the reflected sound.
[0472] Furthermore, the two sounds compared in the selection process are not limited to the direct and reflected sounds of a sound emitted from a single sound source. For example, sound selection can also be performed by comparing two reflected sounds of a sound emitted from a single sound source. In this case, the direct sound in this disclosure can be replaced with the sound that arrives at the listener first, and the reflected sound in this disclosure can be replaced with the sound that arrives at the listener later.
[0473] [Example of bitstream construction] Bitstreams may contain, for example, audio signals and metadata. Audio signals are audio data representing sound, including information related to the frequency and intensity of the sound. Furthermore, metadata contains spatial information related to the space of the sound field, i.e., the sound space.
[0474] For example, spatial information is information relating to the space in which a listener is located when receiving sound based on a sound signal. Specifically, spatial information is information related to a predetermined location (location position) used to position a sound image at a specific location in sound space (e.g., a three-dimensional sound field), that is, information used to enable the listener to perceive sound arriving from a direction corresponding to that predetermined location. Spatial information may include, for example, information about the sound source and location information indicating the listener's position.
[0475] Sound source object information refers to the information about the sound source object that generates sound based on the sound signal. That is, sound source object information is information related to the object (sound source object) that reproduces the sound signal, and it is information related to a virtual sound source object configured in a virtual sound space. Here, the virtual sound space can also correspond to the real space where the object generating the sound is configured, and the sound source object in the virtual sound space can also correspond to the object generating sound in the real space.
[0476] Sound source object information can also represent the location of a sound source object configured in the sound space, the orientation of the sound source object, the directionality of the sound emitted by the sound source object, whether the sound source object is a living being, and whether the sound source object is a moving object. For example, a sound signal can be associated with one or more sound source objects represented by the sound source object information.
[0477] Bitstreams, for example, have a data structure consisting of metadata (control information) and sound signals.
[0478] The audio signal and metadata can be contained in a single bitstream or in multiple separate bitstreams. Furthermore, the audio signal and metadata can be contained in a single file or in multiple separate files.
[0479] Bitstreams can exist either per audio source or per playback time. When bitstreams exist per playback time, multiple bitstreams can be processed in parallel simultaneously.
[0480] Metadata can be assigned to each bitstream individually, or it can be assigned to multiple bitstreams together as information to control them. In this case, multiple bitstreams can also share the metadata. Alternatively, metadata can be assigned at each playback time.
[0481] In the presence of multiple bitstreams or multiple files, information indicating associated bitstreams or files may be included in more than one bitstream or more than one file. Alternatively, information indicating associated bitstreams or files may be included in each of the individual bitstreams or each of the individual files.
[0482] Here, associated bitstreams or associated files refer, for example, to bitstreams or files that may be used simultaneously during audio processing. Additionally, it may also include bitstreams or files that contain information indicating associated bitstreams or associated files.
[0483] Here, the information representing the associated bitstream or file can be, for example, an identifier representing the associated bitstream or file. Alternatively, the information representing the associated bitstream or file can be, for example, the filename, URL (Uniform Resource Locator), or URI (Uniform Resource Identifier).
[0484] In this case, the acquisition unit can also determine and acquire the associated bitstream or associated file based on information representing the associated bitstream or associated file. Alternatively, information representing the associated bitstream or associated file can be included in the bitstream or file, and also in other bitstreams or other files.
[0485] Here, the file containing information representing the associated bitstream or associated file can also be a control file such as a declaration file for content distribution.
[0486] In addition, all or part of the metadata can be obtained from outside the audio signal bitstream. For example, metadata for controlling the audio and metadata for controlling the video can be obtained from outside the bitstream, or metadata for both can be obtained from outside the bitstream.
[0487] Furthermore, metadata for controlling the image may also be included in the bitstream acquired by the stereo sound reproduction system 1000. In this case, the stereo sound reproduction system 1000 may also output the metadata for controlling the image to a display device that displays the image or a stereo image reproduction device that reproduces the stereo image.
[0488] [Example of information contained in metadata] Metadata can also be information used in the description of a scene represented by a sound space. Here, a scene is a term that refers to the collection of all elements of a sound space, including three-dimensional images and sound events, modeled by a sound signal reproduction system using metadata.
[0489] That is, metadata includes not only information used to control audio processing, but also information used to control video processing. Metadata can contain only one of the information used to control audio processing or the information used to control video processing, or it can contain both.
[0490] The stereo sound reproduction system 1000 processes sound signals using metadata contained in the bitstream and interactive listener location information acquired through appending, to generate virtual sound effects. These sound effects can include initial reflection processing, obstacle removal, diffraction processing, masking, and reverberation processing, as well as other sound processing using metadata. For example, sound effects such as distance attenuation, localization, or Doppler effects can be added.
[0491] In addition, information can be attached to the metadata to toggle the on / off of all or some of the additional sound effects, or priority information for multiple processing of sound effects.
[0492] In addition, as an example, metadata includes information related to the sound space, including sound source objects and obstacle objects, and information related to the positioning location used to locate the sound image in a specified position within the sound space (i.e., to make the listener perceive the sound coming from a specified direction).
[0493] Here, an obstacle object is an object that may affect the listener's perception of sound by blocking or reflecting it before the sound emitted by the sound source reaches the listener. Besides stationary objects, obstacle objects can also include moving bodies such as animals or machines. Animals can also be people.
[0494] Furthermore, when multiple sound source objects exist in the sound space, for any given sound source object, the other sound source objects may become obstacle objects. That is, objects that do not emit sound, such as building materials or inanimate objects (i.e., non-sound-emitting objects), as well as sound source objects that emit sound, can all become obstacle objects.
[0495] The metadata contains all or part of the information representing the shape of the sound space, the shape and location of obstacle objects in the sound space, the shape and location of sound source objects in the sound space, and the location and orientation of the listener in the sound space.
[0496] The sound space can be either enclosed or open. Furthermore, the metadata can also include information about the reflectivity of obstacles within the sound space that can reflect sound. For example, the floor, walls, or ceiling that form the boundary of the sound space can also be considered obstacles.
[0497] Reflectivity is the energy ratio of reflected sound to incident sound, and it can be set for each frequency band of the sound. Of course, reflectivity can also be set uniformly regardless of the frequency band of the sound. In addition, when the sound space is an open space, parameters such as attenuation rate, diffraction tone, and initial reflection tone can be set uniformly, for example.
[0498] Metadata can also include information other than reflectance as parameters related to obstacle objects or sound source objects. For example, metadata can also include information related to the raw materials of the object as parameters related to both the sound source object and the non-sound-producing object. Specifically, the raw material-related information in the metadata can include reflectance and other information such as diffraction, transmissivity, and sound absorption, and can also establish a correspondence between each piece of information identifying the raw material and raw material-related parameters such as reflectance, diffraction, and sound absorption.
[0499] Information related to a sound source object may also include volume, radiation characteristics (directivity), reproduction conditions, the number and type of sound sources in an object, and information representing the sound source region within the object. Reproduction conditions may, for example, specify whether the sound is a continuously flowing sound or an event-triggered sound. The sound source region within an object can be set based on the relative position of the listener and the object, or it can be set using the object as a reference.
[0500] For example, when the sound source area is set according to the relative position of the listener and the object, from the listener's perspective, the listener can perceive sound A from the right side of the object and sound B from the left side of the object.
[0501] Furthermore, when using an object as a reference to define a sound source region, it is possible to fix which area of the object emits which sound. For example, when the listener views the object from the front, the listener can perceive high frequencies from the right side of the object and low frequencies from the left side. And when the listener views the object from the back, the listener can perceive low frequencies from the right side of the object and high frequencies from the left side.
[0502] Spatial metadata can also include the time up to the initial reflection, reverberation time, and the ratio of direct to diffuse sound. When the ratio of direct to diffuse sound is zero, the listener can perceive only the direct sound.
[0503] [Brief Summary] This implementation method is briefly summarized here.
[0504] When analyzing the relationship between direct and reflected sounds, with the direct sound set as the preceding sound and the reflected sound set as the following sound, in the case of a relationship that produces a priority effect, that is, when the reflected sound is below the echo detection limit, the reflected sound will not be perceived. Therefore, even if the reflected sound is deleted, the auditory impact on the listener is minimal.
[0505] Figure 26 This diagram illustrates an example of the positional relationship between the listener and the obstacle object in this embodiment. Figure 27 This diagram illustrates another example of the positional relationship between the listener and the obstacle object in this embodiment. Additionally, Figure 26 and Figure 9 The positional relationships shown are the same. Figure 27 and Figure 10 The positional relationships shown are the same. Additionally, Figure 28 This is an example of the echo detection limit threshold in this embodiment. Additionally, Figure 28 The echo detection limit threshold shown is Figure 12C An example of threshold data is shown.
[0506] For example, when comparing Figure 26 Positional relationship and Figure 27 When considering positional relationships, Figure 26 The volume of reflected sound heard by the listener under the positional relationship is greater than Figure 27 The volume of reflected sound heard by a listener in a certain position is small. This is because... Figure 26 The path length of the reflected sound under the positional relationship is greater than Figure 27 The path length of the reflected sound is long under the positional relationship.
[0507] Therefore, when judging solely by the volume of the reflected sound, Figure 26 The reflected sound ratio shown Figure 27 The reflected sound shown has little impact on hearing. However, if... Figure 26 The arrival time of the reflected sound at the listening position shown is... Figure 27 Comparing the arrival times of the reflected sound at the listening position as shown, then Figure 26 The situation shown is even later.
[0508] Therefore, if judged from the perspective of the echo detection limit, then as Figure 28 As shown, Figure 27 The reflected sound shown is below the echo detection limit, so it will not be perceived as a reflected sound by the listener. Figure 26 The reflected sound shown is perceived as a reflected sound by the listener because it is above the echo detection limit.
[0509] In this embodiment, this situation is used to determine the auditory importance of the reflected sound, so that unimportant reflected sounds are not reproduced, thereby reducing the amount of computation related to the processing of reflected sounds.
[0510] The above is a brief summary of this implementation method.
[0511] In addition, to make the above determination, it is necessary to calculate the time difference between the direct tone and the reflected tone, as well as the volume ratio between the direct tone and the reflected tone.
[0512] Performing such calculations could increase the computational load. In particular, calculating the volume ratio of direct to reflected sound sometimes requires a cumulative calculation process using various parameters in virtual space, which could lead to an increase in computational load.
[0513] Therefore, the following section will further explain in detail a sound signal processing method that can appropriately reduce computational load in the sound space.
[0514] (Implementation Method 2) Hereinafter, Embodiment 2 will be described. The description will focus on the differences from Embodiment 1, and the description of the commonalities will be omitted or simplified.
[0515] [The Structure of the Rendering Department] First, the configuration of the rendering unit 2300 in this embodiment will be explained. Figure 29 This is a block diagram showing an example of the configuration of the rendering unit 2300 in this embodiment.
[0516] The rendering unit 2300 includes a resolution unit 2301, a selection unit 2302, and a reproduction unit 2303. Furthermore, the audio signal processing apparatus of this embodiment is an example of a decoding apparatus, which includes a decoder, and the decoder includes the rendering unit 2300. That is, it can be said that the audio signal processing apparatus of this embodiment includes a resolution unit 2301, a selection unit 2302, and a reproduction unit 2303. The rendering unit 2300 performs additional audio processing on the audio data contained in the input signal and outputs it.
[0517] Similar to Implementation 1, the input signal consists of, for example, spatial information, sensor information, and sound data. The spatial information also includes physical information such as the reflection coefficient, transmission coefficient, and diffraction coefficient of the object (obstacle object).
[0518] Furthermore, in this embodiment, reflected sounds are mainly used as an example of indirect sounds for explanation, but the same treatment is applied even when indirect sounds are used instead of reflected sounds. Additionally, as an example, indirect sounds are reflected sounds or diffracted sounds, etc.
[0519] The analysis unit 2301 can perform all or part of the processing performed by the analysis unit 1301 in Embodiment 1.
[0520] Similar to the analysis unit 1301 in Embodiment 1, the analysis unit 2301 analyzes the sound signal contained in the input signal and the spatial information received from the spatial information management units 1201 and 1211. Therefore, the analysis unit 2301 calculates the information required for generating direct and reflected sounds in the reproduction unit 2303, as well as the information required for selecting whether to generate reflected sounds. The method by which the analysis unit 2301 calculates this information is as described in Embodiment 1.
[0521] Furthermore, the analysis unit 2301 is as described in the analysis unit 1301 of Embodiment 1. Figure 8 As performed in S101, the input signal is analyzed. That is, the analysis unit 2301 analyzes the input signal input to the sound signal processing device of this embodiment and detects direct sounds and reflected sounds that may be generated in the sound space.
[0522] Upon detecting such direct and reflected sounds, the analysis unit 2301 generates a sound signal representing the reflected sound and a sound signal representing the direct sound based on spatial information and sound data.
[0523] More specifically, the analysis unit 2301 generates sound signals representing reflected sound and sound signals representing direct sound based on the location information of the sound source object, the location information of the object (obstacle object), the location information and physical information of the listener, and sound data contained in the spatial information.
[0524] That is, the analysis unit 2301 generates a sound signal generated in the virtual space based on spatial information and sound data. More specifically, the sound signal generated by the analysis unit 2301 is assigned attribute information for determining the attributes of the sound signal; that is, the analysis unit 2301 generates a sound signal containing such attribute information. The analysis unit 2301 also generates attribute information. A sound signal containing attribute information is generated for each sound generated in the virtual space. The attributes of the sound signal include information indicating whether the sound represented by the sound signal is a direct tone or a reflected tone. In this embodiment, as an example, the attribute is information indicating whether the sound represented by the sound signal is a direct tone or a reflected tone.
[0525] In addition, for simplicity, sometimes a sound signal with the attribute representing reflected sound (indirect sound) is recorded as a sound signal representing reflected sound (indirect sound), and a sound signal with the attribute representing direct sound is recorded as a sound signal representing direct sound.
[0526] Furthermore, the attribute information can also include information necessary for radiating the sound signal into the sound space, such as gain information, gain characteristics for each bandwidth, positional information, and directivity information. That is, this necessary information can also be retained in the attribute information. Additionally, the attribute information can be associated with the sound signal as metadata. For example, the gain characteristics of the sound signal for each bandwidth contained in the attribute information can be determined based on spatial information contained in the input information. Information representing frequency characteristics indicating auditory sensitivity can be determined based on spatial information contained in the input information, and in particular, can be determined as information associated with the listener's persona.
[0527] A sound that reaches the listener's head directly from a sound source is a direct sound, while a sound that reaches the listener's head after being reflected or diffracted by a reflective object from the same sound source is an indirect sound (reflected sound or diffracted sound).
[0528] In addition, in this embodiment, the analysis unit 2301 generates a sound signal representing a reflected tone (indirect tone) and a sound signal representing a direct tone related to the reflected tone (indirect tone).
[0529] Furthermore, the direct sound in relation to an indirect sound refers to a direct sound originating from the same sound source as the indirect sound.
[0530] The sound signal whose attribute represents information about the reflected tone (indirect tone) contains information about the direct tone sound signal related to that reflected tone (indirect tone).
[0531] The analysis unit 2301 can store the generated sound signals in its own memory. Furthermore, the analysis unit 2301 can generate multiple sound signals and store them in the memory.
[0532] Furthermore, similar to Embodiment 1, the analysis unit 2301 can calculate values related to the path to the listener's location (i.e., the listening position) and the time taken to reach the direct sound and reflected sound, respectively. Similarly, the analysis unit 2301 can calculate information representing the relationship between the direct sound and the reflected sound, such as values related to the time difference between the time it takes for the direct sound to reach the listening position from the sound source location (i.e., the sound source position) and the time it takes for the reflected sound to reach the listening position (the time difference between the direct sound and the reflected sound).
[0533] Furthermore, the analysis unit 2301 includes a first analysis unit 2301a, a second analysis unit 2301b, and a third analysis unit 2301c.
[0534] The first analysis unit 2301a calculates a first volume and a second volume. The first volume is the volume corresponding to the direction from the sound source position to the listening position, and the second volume is the volume corresponding to the direction from the sound source position to the position of the reflector, i.e., the reflector position. Furthermore, the sound corresponding to the direction from the sound source position to the listening position is defined as the first sound; that is, the first volume is the volume of the sound corresponding to the direction from the sound source position to the listening position, i.e., the first sound. Conversely, the sound corresponding to the direction from the sound source position to the reflector position is defined as the second sound; that is, the second volume is the volume of the sound corresponding to the direction from the sound source position to the reflector position, i.e., the second sound.
[0535] Thus, the first analysis unit 2301a calculates the first volume of the first sound and the second volume of the second sound. First, using Figure 30 Explain the first and second sounds.
[0536] Figure 30 This is a diagram used to illustrate the first sound, the second sound, the direct sound, and the reflected sound in this embodiment.
[0537] like Figure 30 As shown, sound source 100 emits sound. The sound emitted from sound source 100 diffuses to its surroundings. As described above, direct sound is the sound that reaches the listener's head directly from sound source 100, while reflected sound is the sound that reaches the listener's head after being reflected by a reflective object (such as a wall) after being output from sound source 100.
[0538] The first sound is the sound emitted from sound source 100 and becomes a direct sound. More specifically, the first sound is the sound immediately after it is emitted from sound source 100, and is the sound that has not yet reached the listener. Figure 30 In the diagram, the first sound is represented by a solid line. The first sound travels from the sound source 100 toward the listener, becoming a direct sound that is heard by the listener.
[0539] The second sound is the sound emitted from sound source 100 and becomes a reflected sound. More specifically, the second sound is the sound immediately after it is emitted from sound source 100 and is moving towards the reflecting object. The second sound is the sound that has not yet reached the reflecting object (the reflecting object) and the listener. Figure 30 In the diagram, the second sound is represented by a double-dotted line. The second sound propagates from the sound source 100 towards the reflecting object (e.g., a wall), is reflected by the reflecting object, and thus becomes... Figure 30 The reflected sound, indicated by the single-dotted line, is heard by the listener.
[0540] Reflected sound is the sound emitted from sound source 100 that is reflected by a reflective object and travels from the position of the reflective object to the position of the receiver.
[0541] The position of the reflector is determined based on the positional information of both the reflector object and the sound source object; for example, it could also be the position of the sound image of the reflected sound. More specifically, such as... Figure 30 As shown, the position where the sound emitted from the sound source is reflected by the reflecting surface can also be set as the position of the sound image of the reflected sound, that is, the position of the reflecting object.
[0542] Furthermore, the position of the reflector can be determined independently of the position of the sound source object. For example, it could be the center of the reflector object or any point on its surface. While this coarse determination of the reflected position reduces the precision of the reflected sound's directionality, it significantly saves computational resources.
[0543] The second sound travels in the direction from the sound source to the reflector, typically towards the point where the sound emitted from the sound source is reflected by the reflector, but it can also be towards any of the reflector locations mentioned above. Similarly, the reflected sound travels in the direction from the reflector to the listening position, typically from the point where the sound emitted from the sound source is reflected by the reflector towards the listening position, but it can also be towards any of the reflector locations mentioned above.
[0544] Furthermore, the first analysis unit 2301a calculates the first volume of the first sound and the second volume of the second sound. In this embodiment, the first analysis unit 2301a uses the directivity of sound, which is an example of the property of an object existing in the propagation of sound, to calculate the first volume and the second volume.
[0545] Here, the object existing in the propagation of sound is the sound source 100, and the directivity of sound, as an example of the property of the object existing in the propagation of sound, is the directivity of the sound emitted from the sound source 100. That is, the first analysis unit 2301a calculates the directivity of the sound emitted from the sound source 100 by performing analysis processing of the input signal. For example, the directivity of sound includes the spatial information contained in the input signal.
[0546] Furthermore, using Figure 31 and Figure 32A Explain the directionality of sound.
[0547] Figure 31 This is a diagram illustrating the directionality of the sound in this embodiment. Figure 32A This is a diagram illustrating the relationship between the directivity of the sound in this embodiment and the first sound, the second sound, the direct sound, and the reflected sound. Furthermore, Figure 32A yes Figure 31 The directionality of the sound shown Figure 30 The diagrams shown overlap.
[0548] The directionality of sound, for example, is determined by... Figure 31 The directivity diagram shown is used for definition. Data representing the directivity of sound can also be stored in SOFA (Spatially Oriented Format for Acoustics) format.
[0549] In data representing the directionality of sound, the scale indicates the volume of the sound, for example, expressed in dB values. The further out of the circle, the louder the sound; the closer to the center, the quieter the sound. Figure 31 The region enclosed by the curve shown represents the directivity of the sound source 100; that is, the curve represents the volume of the sound in the direction in which the sound travels. Furthermore, the volume is not limited to dB values and can be represented by other values as well.
[0550] Furthermore, in data representing the directionality of sound, the scale can also be described as the amount of sound attenuation. That is, the further out of the circle, the less attenuation, and the further out of the circle, the more attenuation.
[0551] The frontal direction is the direction in which sound travels towards 0°. For example... Figure 31 and Figure 32A As shown, for example, the first sound is a sound that travels in a forward direction. The first analysis unit 2301a calculates the first volume of the first sound as 0 dB based on the directivity of the sound and the direction in which the first sound travels.
[0552] In addition, such as Figure 31 and Figure 32A As shown, for example, the second sound is a sound that travels in a 270° direction. The first analysis unit 2301a calculates the second volume of the second sound as -8 dB based on the sound's directivity and the direction in which the second sound travels.
[0553] Thus, the first analysis unit 2301a calculates the first volume and the second volume based on the directivity of the sound, which is an example of the properties of an object existing in the propagation of sound.
[0554] Furthermore, in the explanation up to this point, it has been stated that the second volume level corresponds to the volume level in the direction from the sound source position towards the position of the reflector, i.e., the position of the reflector. Hereafter, we will use... Figure 32B Here is an example illustrating the calculation method for this direction. Figure 32BThis diagram illustrates an example of the method for calculating the second volume in this embodiment. When the virtual image (mirror image) 110 of the sound source 100 is placed at a position mirrored by a wall (reflector), and the point where the line connecting the position of the virtual image (mirror image) 110 and the listener's position intersects the reflector is designated as the reflection point (reflector position), the direction connecting the sound source 100 to this reflection point can also be calculated as the direction of the second sound. In this case, the second volume can also be calculated by applying the direction of the second sound to the directivity of the sound source. Alternatively, the second volume can of course be calculated by applying the direction of the reflected sound (from the virtual image (mirror image) 110 towards the listener) to the directivity of the aforementioned virtual image (mirror image) 110. In this case, the directivity of the aforementioned virtual image (mirror image) 110 is naturally mirrored with the directivity of the sound source. Furthermore, the direction from the sound source position to the reflector position, i.e., the reflector position, can be calculated without using a mirror image. That is, the result is the same whether the second volume is calculated using a mirror image or not. To determine the direction and volume from the sound source location towards the reflection point (reflector location), the calculated results of the direction and volume from the mirror image location of the sound source towards the listening location can also be used. In other words, setting the direction from the sound source location towards the reflection point (reflector location) to the direction corresponding to the second volume is essentially the same as setting the direction from the mirror image location of the sound source towards the listening location to the direction corresponding to the second volume.
[0555] Furthermore, the first analysis unit 2301a calculates and obtains the first volume and the second volume using the directionality of the sound, but is not limited to this. For example, the first analysis unit 2301a can also calculate the first volume and the second volume by inferring the volume in each direction based on the shape data associated with the input signal. For example, in the case of the shape data being a trumpet, it can also calculate the first volume and the second volume, which are stronger in the direction of the blowhole and whose volume decreases according to a predetermined rule whenever the direction is deviated from.
[0556] Furthermore, the first analysis unit 2301a calculates the volume ratio of the first volume to the second volume, calculated based on the directivity of the sound. For example, the first analysis unit 2301a uses the reference volume of the sound emitted by the sound source 100 to calculate the volume ratio. This volume ratio is sometimes referred to as Dir. As described above, spatial information includes information related to the sound source object contained in the sound space; the information related to the sound source object includes information necessary for radiating sound data into the sound space; and the information necessary for radiating sound data into the sound space includes information about the reference volume. Therefore, the first analysis unit 2301a calculates the reference volume of the sound emitted by the sound source 100 by performing analysis processing of the input signal, thereby calculating the volume ratio.
[0557] For example, in data representing the directivity of sound, the scale is the amount of sound attenuation and expressed in dB values. The first volume is represented by (reference volume + attenuation in the direction the first sound travels), and the second volume is represented by (reference volume + attenuation in the direction the second sound travels). Furthermore, Dir, as the volume ratio, is calculated by (reference volume + attenuation in the direction the second sound travels) - (reference volume + attenuation in the direction the first sound travels), i.e., (attenuation in the direction the second sound travels - attenuation in the direction the first sound travels).
[0558] Furthermore, when the first and second volumes are represented by quantities other than dB values, if the first volume is set to A and the second volume is set to B, then Dir satisfies the following equation 5.
[0559] Dir = B / A … (Equation 5) Next, the second analysis unit 2301b will be explained.
[0560] The second analysis unit 2301b performs processing using the reflection coefficient. For example, the second analysis unit 2301b calculates the reflection coefficient of a reflector, which is an example of the property of an object existing in the propagation of sound, by performing analysis processing of the input signal. Here, the object existing in the propagation of sound is a reflector that reflects the second sound, i.e., a reflector object (wall). The second analysis unit 2301b calculates the reflection coefficient of this reflector object and performs processing using the calculated reflection coefficient. In addition, sometimes this reflection coefficient is recorded as Ref.
[0561] The third analysis unit 2301c calculates the direct sound path length to the listener's location (i.e., the listening position) and the reflected sound path length to the listening position. More specifically, the third analysis unit 2301c calculates the direct sound path length and the reflected sound path length based on the positions of objects present in the sound propagation. Here, the positions of objects present in the sound propagation include the position of the sound source object, the position of the reflecting object (wall), and the position of the listener.
[0562] The third analysis unit 2301c performs analysis processing on the input signal to calculate the position of the sound source object, the position of the reflecting object (wall), and the position of the listener. Based on the calculated positions of the sound source object, the reflecting object (wall), and the listener, the third analysis unit 2301c calculates the direct sound path length and the reflected sound path length geometrically.
[0563] Then, the third analysis unit 2301c calculates the path length ratio of the direct tone path length to the reflected tone path length. The path length ratio is the direct tone path length / the reflected tone path length.
[0564] The selection unit 2302 can perform all or part of the processing performed by the selection unit 1302 in Embodiment 1. Furthermore, the selection unit 2302 determines whether the reproduction unit 2303 should output (reproduce) an output signal based on the sound signal generated by the analysis unit 2301. That is, the selection unit 2302 first specifies and acquires one of a plurality of sound signals (e.g., a sound signal representing a reflected sound) generated by the analysis unit 2301. Then, the selection unit 2302 determines whether to select the reflected sound represented by the acquired sound signal. If the reflected sound is selected, the reproduction unit 2303 outputs an output signal based on the sound signal representing that reflected sound.
[0565] The selection unit 2302 has an acquisition unit 2302a and a determination unit 2302b. The determination unit 2302b includes a first determination unit 2302b1, a second determination unit 2302b2, and a third determination unit 2302b3.
[0566] The acquisition unit 2302a acquires the sound signals generated by the analysis unit 2301 and stored in the memory of the analysis unit 2301. For example, the acquisition unit 2302a acquires the sound signals representing reflected sounds (indirect sounds) and the sound signals representing the direct sounds related to the reflected sounds (indirect sounds). In addition, the acquisition unit 2302a acquires the values related to the time difference between the direct sounds and the reflected sounds calculated by the analysis unit 2301.
[0567] The determination unit 2302b determines whether to select the reflected sound represented by the sound signal acquired by the acquisition unit 2302a. If the determination unit 2302b determines that the reflected sound should be selected, the determination unit 2302b outputs the acquired sound signal to the reproduction unit 2303, the reproduction unit 2303 acquires the output sound signal, and outputs an output signal based on the acquired sound signal.
[0568] The determination unit 2302b can determine whether to select a reflected sound, for example, through the following process. The determination unit 2302b determines whether to select a reflected sound based on the time difference obtained by the acquisition unit 2302a, the properties of the object present in the propagation of the sound, the first volume of the first sound, and the second volume of the second sound.
[0569] More specifically, the property of the object present in the propagation of sound is the directionality of the sound. The first determination unit 2302b1 determines whether to select the reflected sound based on the time difference, a first volume calculated based on the directionality of the sound, and a second volume calculated based on the directionality of the sound. When performing this process, the first determination unit 2302b1 can obtain the first volume and the second volume calculated by the first analysis unit 2301a based on the directionality of the sound.
[0570] Furthermore, for example, the first determination unit 2302b1 may also determine whether to select reflected sound based on the time difference and the volume ratio of the first volume to the second volume calculated according to the directionality of the sound. When performing this process, the first determination unit 2302b1 may obtain the volume ratio calculated by the first analysis unit 2301a.
[0571] Alternatively, for example, the determination unit 2302b may use the reflection coefficient of the reflector as a property of the object present in the propagation of sound to determine whether to select the reflected sound.
[0572] More specifically, the second determination unit 2302b2 may also use the reflection coefficient calculated by the second analysis unit 2301b to determine whether to select the reflected sound.
[0573] Alternatively, for example, the determination unit 2302b may also determine whether to select reflected sound based on the position of an object present in the propagation of sound.
[0574] More specifically, the positions of the objects present in the propagation of sound are the position of the sound source object, the position of the reflecting object (wall), and the position of the listener. At this time, the third determination unit 2302b3 can also determine whether to select the reflected sound based on the path length ratio of the direct sound path length to the reflected sound path length calculated by the third analysis unit 2301c based on the position of the object.
[0575] The reproduction unit 2303 can perform all or part of the processing performed by the reproduction unit 1303 in Embodiment 1. Furthermore, the reproduction unit 2303 acquires the sound signal output from the selection unit 2302 and outputs an output signal based on the acquired sound signal.
[0576] The reproduction unit 2303 processes the acquired sound signal by performing binaural filtering and other methods to generate and output an output signal. Binaural filtering is implemented, for example, by processing the acquired sound signal using a head-related transfer function.
[0577] In addition, the reproduction unit 2303 can also synthesize and output the sound signal representing the direct tone acquired by the acquisition unit 2302a and the generated output signal.
[0578] Furthermore, the reproduction unit 2303 can also generate and output an output signal by performing both binaural filtering and diffusion filtering on the sound signal output from the selection unit 2302. Diffusion filtering, for example, improves the realism of the indirect sound by diffusing the reflected sound (indirect sound) represented by the sound signal into the acquired sound signal. Additionally, diffusion filtering uses a filter that realistically simulates the audible strength of the sound diffusion represented by the acquired sound signal (i.e., simulates the audible strength of the sound diffusion perceived by the listener). Finite-pulse filters and / or infinite-pulse filters are used in diffusion filtering.
[0579] Hereinafter, an example of the operation of the sound signal processing method performed by the sound signal processing apparatus (more specifically, the rendering unit 2300) of this embodiment will be described.
[0580] [Example of actions in the rendering department] Figure 33 This is a flowchart illustrating an example of the operation of the sound signal processing apparatus in this embodiment. Figure 33 This section primarily illustrates the processing performed by the rendering unit 2300 included in the sound signal processing apparatus of this embodiment. Furthermore, details related to Embodiment 1 are omitted or simplified here. Figure 8 Explanation of commonalities.
[0581] First, the analysis unit 2301 performs analysis processing on the input signal (S101a). More specifically, the analysis unit 2301 analyzes the input signal and detects direct and reflected sounds that may occur in the sound space. Upon detecting such direct and reflected sounds, the analysis unit 2301 generates a sound signal representing the reflected sound and a sound signal representing the direct sound based on spatial information and sound data. The analysis unit 2301 stores the generated sound signals in its memory.
[0582] In addition, the analysis unit 2301 analyzes the input signal and calculates values related to the path to the listening position, the time taken to reach the position, and the time difference between the direct tone and the reflected tone for both the direct tone and the reflected tone.
[0583] First, the analysis unit 2301 calculates the characteristics of the direct tone represented by the generated sound signal and the reflected tone represented by the generated sound signal. Specifically, it calculates the arrival time of the direct tone and the reflected tone when they reach the listener (listening position). Furthermore, the method for calculating these arrival times can be the method shown in Embodiment 1.
[0584] Furthermore, the analysis unit 2301 calculates the time difference (T) between the direct tone and the reflected tone. In addition, the method for calculating the time difference (T) can be the method shown in Embodiment 1.
[0585] In addition, unlike step S101 of embodiment 1, in step S101a, it is not necessary to calculate the volume ratio (L).
[0586] In addition, the first analysis unit 2301a can calculate the first volume and the second volume, and can also calculate the volume ratio of the first volume to the second volume.
[0587] Furthermore, the second analysis unit 2301b can calculate the reflection coefficient of the reflective object (wall). As described above, the second sound is a reflected sound that is reflected by the reflective object.
[0588] Furthermore, the third analysis unit 2301c can calculate the direct tone path length and the reflected tone path length, and can calculate the path length ratio of the direct tone path length to the reflected tone path length.
[0589] The selection unit 2302 performs selection of the reflected sound (selection processing) (S102a). That is, the selection unit 2302 selects whether the reproduction unit 2303 reproduces the output signal representing the reflected sound produced by the analysis unit 2301.
[0590] First, the acquisition unit 2302a acquires the sound signal generated by the analysis unit 2301 and stored in the memory. For example, the acquisition unit 2302a acquires at least one of the sound signal representing the reflected tone (indirect tone) and the sound signal representing the direct tone related to the reflected tone (indirect tone), and acquires both here. In addition, the acquisition unit 2302a acquires the value related to the time difference (T) between the direct tone and the reflected tone calculated by the analysis unit 2301.
[0591] As mentioned above, the reflected sound and the direct sound related to that reflected sound are sounds from the same sound source.
[0592] The determination unit 2302b determines whether to select a reflected sound through the following process, for example. The determination unit 2302b determines whether to select a reflected sound based on the time difference obtained by the acquisition unit 2302a, the properties of the object present in the propagation of the sound, the first volume of the first sound, and the second volume of the second sound.
[0593] Furthermore, the determination unit 2302b also uses a threshold determined based on the time difference (T) between the direct tone and the reflected tone to determine whether to select the reflected tone. If the determination unit 2302b determines that the reflected tone should be selected, it outputs a sound signal representing the reflected tone to the reproduction unit 2303. The reproduction unit 2303 acquires the output sound signal and outputs an output signal based on that sound signal. In other words, this is equivalent to the determination unit 2302b selecting the reproduction unit 2303 to output an output signal based on the acquired sound signal representing the reflected tone.
[0594] The threshold is a value determined based on the time difference (T) between the direct tone and the reflected tone; in other words, it is a value dependent on the time difference (T), as shown in the threshold data of Implementation 1. The threshold data is, for example, represented as the threshold at which the reflected tone is perceived or not in a graph where the horizontal axis has the value of the time difference (T) between the direct tone and the reflected tone, and the vertical axis has the volume ratio of the direct tone and the reflected tone.
[0595] More specifically, the threshold data representing the threshold is Figures 11-13 The data shown is as follows.
[0596] Selection unit 2302 performs selection processing as described above.
[0597] The reproduction unit 2303 acquires the sound signal output from the selection unit 2302 (more specifically, the determination unit 2302b) and outputs an output signal based on the sound signal (S103a). Here, the reproduction unit 2303 combines and outputs the sound signal representing the direct tone acquired by the acquisition unit 2302a and the generated output signal (the sound signal representing the reflected tone).
[0598] Furthermore, examples 1 through 7 are used to illustrate the parsing and selection processes in more detail.
[0599] <Example 1> First, let's explain the first example of parsing and selection processing.
[0600] Figure 34 This is a flowchart illustrating the first example of the parsing and selection processes in this embodiment. Furthermore, details related to Embodiment 1 are omitted or simplified here. Figure 14 Explanation of commonalities.
[0601] First, the selection unit 2302 selects the reflected sound detected by the analysis unit 2301 (S310). That is, the acquisition unit 2302a of the selection unit 2302 selects the sound signal generated by the analysis unit 2301 and stored in the memory, and acquires the selected sound signal. For example, the acquisition unit 2302a selects the sound signal representing the reflected sound and acquires the selected sound signal. At this time, the acquisition unit 2302a can also acquire the sound signal representing the direct sound related to the reflected sound. Furthermore, this step S310 is the same process as step S201.
[0602] Furthermore, the analysis unit 2301 (more specifically, the first analysis unit 2301a) calculates the volume ratio (Dir) of the first volume and the second volume (S311). More specifically, firstly, the first analysis unit 2301a calculates the first volume and the second volume. For example, the first analysis unit 2301a uses the directivity of sound, which is an example of the property of an object existing in the propagation of sound, to calculate the first volume and the second volume. Furthermore, the first analysis unit 2301a calculates the volume ratio (Dir) of the first volume and the second volume calculated based on the directivity of sound.
[0603] The analysis unit 2301 calculates the time difference (T) between the direct tone and the reflected tone (S312). That is, as described above, the analysis unit 2301 calculates a value related to the time difference between the direct tone and the reflected tone.
[0604] Additionally, the selection unit 2302 (determination unit 2302b) uses threshold data to determine a threshold corresponding to the time difference (T) calculated in step S312 (S313). Furthermore, this step S313 is the same process as step S204.
[0605] Furthermore, the determination unit 2302b (first determination unit 2302b1) determines whether to select a reflected sound based on the time difference obtained by the acquisition unit 2302a, the properties of the object present in the sound propagation, the first volume of the first sound, and the second volume of the second sound. More specifically, the first determination unit 2302b1 determines whether to select a reflected sound based on the time difference, the first volume calculated based on the sound's directivity, and the second volume calculated based on the sound's directivity. More specifically, the first determination unit 2302b1 determines whether to select a reflected sound based on the time difference and the volume ratio (Dir) of the first volume and the second volume calculated based on the sound's directivity. In this example, the first determination unit 2302b1 determines whether the volume ratio (Dir) calculated by the first analysis unit 2301a is above a threshold determined corresponding to the time difference (T) (S314).
[0606] If the volume ratio (Dir) is less than the threshold (No in step S314), the selection unit 2302 does not select the reflected tone (S341). That is, in this case, the selection unit 2302 does not select the reflected tone specified in step S310. Furthermore, this step S341 is the same process as step S207.
[0607] Here, use Figure 35 The processing of step S314 will be explained.
[0608] Figure 35 This is a diagram illustrating the process of step S314 in this embodiment. Figure 35 Two volume ratios (Dir) and echo detection thresholds are shown. Additionally, Figure 35 The echo detection limit threshold shown is Figure 12C An example of threshold data is shown.
[0609] The two volume ratios are represented by solid circles and dashed circles. If the volume ratio calculated in step S311 is equal to the volume ratio represented by the solid circle, it is determined that the volume ratio is above the threshold (echo detection limit threshold). Conversely, if the volume ratio calculated in step S311 is equal to the volume ratio represented by the dashed circle, it is determined that the volume ratio is below the threshold (echo detection limit threshold).
[0610] Figure 35 The processing in step S314 described herein does not use the volume ratio (L) of direct tone to reflected tone used in step S205 of embodiment 1. As mentioned above, calculating this volume ratio (L) may increase the computational load, but in this embodiment, since the volume ratio (L) is not used, the computational load and computational burden can be appropriately reduced.
[0611] Furthermore, a volume ratio (Dir) less than a threshold means that the second volume of the second sound is sufficiently smaller than the first volume of the first sound. Even if such a second sound is reflected by a reflective object or other reflective surface and reaches the listening position, the volume of the reflected sound heard by the listener is very small. Therefore, even if the listener cannot hear the reflected sound, they rarely perceive any dissonance. That is, even if the reflected sound is removed, the auditory impact on the listener is minimal.
[0612] That is, the sound signal processing method of this embodiment can suppress the dissonance felt by the listener and can appropriately reduce the amount of computation and the computational load.
[0613] reuse Figure 34 Please provide an explanation.
[0614] If the volume ratio (Dir) is above a threshold ("Yes" in step S314), the analysis unit 2301 (second analysis unit 2301b) performs the following processing. That is, the second analysis unit 2301b determines the reflection coefficient (Ref) and multiplies the calculated volume ratio (Dir) by the determined reflection coefficient (Ref) (S321). As a result, the second analysis unit 2301b calculates the value (Dir×Ref) resulting from multiplying the volume ratio (Dir) and the reflection coefficient (Ref).
[0615] Here, the second analysis unit 2301b calculates and determines the reflection coefficient (Ref) of a reflector, which is an example of the properties of an object existing in the propagation of sound. The object existing in the propagation of sound is a reflector that reflects the second sound, i.e., the reflector object (wall).
[0616] Furthermore, the second analysis unit 2301b calculates the value (Dir×Ref) by multiplying the calculated volume ratio (Dir) by the calculated reflection coefficient (Ref).
[0617] The determination unit 2302b (the second determination unit 2302b2) determines whether to select a reflected sound based on the time difference obtained by the acquisition unit 2302a, the calculated reflection coefficient (Ref), the first volume calculated based on the sound directivity, and the second volume calculated based on the sound directivity. More specifically, the second determination unit 2302b2 determines whether to select a reflected sound based on the time difference and Equation 6.
[0618] Dir×Ref …(Equation 6) In this example, the second determination unit 2302b2 determines whether the value (Dir×Ref) calculated by the second analysis unit 2301b is above the threshold determined corresponding to the time difference (T) (S322).
[0619] If the value (Dir×Ref) is less than the threshold (No in step S322), the selection unit 2302 does not select the reflected tone (S341). That is, in this case, the selection unit 2302 does not select the reflected tone specified in step S310.
[0620] Here, use Figure 36 The processing of step S322 will be explained.
[0621] Figure 36 This is a diagram illustrating the process of step S322 in this embodiment. Figure 36 This displays the volume ratio (Dir), value (Dir×Ref), and echo detection threshold. Additionally, Figure 36 The echo detection limit threshold shown is Figure 12C An example of threshold data is shown.
[0622] Figure 36 The volume ratio (Dir) shown is the volume ratio (Dir) that is determined to be "yes" in step S314, that is, the volume ratio (Dir) above the threshold. Figure 36 The value shown (Dir×Ref) is for Figure 36 The value is obtained by multiplying the volume ratio (Dir) by the reflection coefficient (Ref). In step S322, the calculated value (Dir×Ref) is equivalent to... Figure 36 Given the value (Dir×Ref), it is determined that the value (Dir×Ref) is less than the threshold (echo detection limit threshold).
[0623] In step S322, the volume ratio (L) is not used, just as in step S314. Therefore, the computational load and computational complexity can be appropriately reduced.
[0624] Furthermore, a value (Dir×Ref) less than the threshold indicates that the volume of the reflected sound, where the second sound is reflected by a reflective object or other reflective surface, is sufficiently small compared to the volume of the first sound. Even when the reflected sound reaches the listening position, the volume of the reflected sound heard by the listener is very low. Therefore, even if the listener cannot hear the reflected sound, they rarely perceive any dissonance. In other words, even if the reflected sound is removed, the auditory impact on the listener is minimal.
[0625] That is, the sound signal processing method of this embodiment can suppress the dissonance felt by the listener and can appropriately reduce the amount of computation and the computational load.
[0626] reuse Figure 34 Please provide an explanation.
[0627] If the value (Dir×Ref) is above the threshold ("Yes" in step S322), the analysis unit 2301 (the third analysis unit 2301c) performs the following processing. That is, the third analysis unit 2301c calculates the volume ratio (L1) of the direct tone to the reflected tone based on the position of the object present in the propagation of the sound (S331). In addition, this volume ratio (L1) of the direct tone to the reflected tone is equivalent to the volume ratio of the direct tone heard by the listener to the volume of the reflected tone heard by the listener. Furthermore, this volume ratio (L1) of the direct tone to the reflected tone is equivalent to the volume ratio of the direct tone reaching the listener to the volume of the reflected tone reaching the listener.
[0628] The method for calculating the volume ratio (L1) in the third analysis unit 2301c is as follows.
[0629] First, the third analysis unit 2301c calculates the direct sound path length and the reflected sound path length. More specifically, the third analysis unit 2301c calculates the direct sound path length and the reflected sound path length based on the positions of objects present in the propagation of sound. That is, the positions of objects present in the propagation of sound are the position of the sound source object, the position of the reflecting object (wall), and the position of the listener.
[0630] Furthermore, the third analysis unit 2301c calculates the path length ratio of the direct sound path length to the reflected sound path length. The path length ratio is the direct sound path length / reflected sound path length. Generally speaking, the longer the distance from the sound source to the listening position, the lower the sound volume, i.e., there is attenuation due to distance. The path length ratio is used to represent this attenuation caused by distance.
[0631] Furthermore, the third analysis unit 2301c calculates the volume ratio (L1) of the direct tone to the reflected tone by multiplying the value (Dir×Ref) calculated by the second analysis unit 2301b in step S321 by the path length ratio. That is, the volume ratio (L1) of the direct tone to the reflected tone satisfies the following equation 7.
[0632] L1 = (Dir × Ref) × path length ratio… (Equation 7) The volume ratio (Dir) is the ratio of the volume of the first sound immediately after it is emitted from the sound source (e.g., sound source 100) to the volume of the second sound. (Dir×Ref) represents the ratio of the volume of the first sound to the volume of the second sound after it is reflected by a reflector. To represent the first and second sounds reaching the listening position as direct and reflected sounds, the volume ratio of direct sound to reflected sound (L1) is calculated by multiplying this ratio by the ratio of path lengths that account for volume attenuation due to distance to the listener. Alternatively, the volume ratio of direct sound to reflected sound (L1) can also be described as a volume ratio that takes into account attenuation due to distance.
[0633] The determination unit 2302b (third determination unit 2302b3) determines whether to select a reflected sound based on the position of an object present in the propagation of the sound. More specifically, the third determination unit 2302b3 determines whether to select a reflected sound based on the path length ratio of the direct sound path length to the reflected sound path length, calculated according to the position of the object present in the propagation of the sound. In this example, the third determination unit 2302b3 determines whether the volume ratio (L1) of the direct sound to the reflected sound calculated by the third analysis unit 2301c is above a threshold determined corresponding to the time difference (T) (S332).
[0634] If the volume ratio (L1) is less than the threshold (No in step S332), the selection unit 2302 does not select the reflected tone (S341). That is, in this case, the selection unit 2302 does not select the reflected tone specified in step S310.
[0635] Here, use Figure 37 The processing of step S332 will be explained.
[0636] Figure 37 This is a diagram illustrating the process of step S332 in this embodiment. Figure 37 The values (Dir×Ref), the volume ratio of direct to reflected sound (L1), and the echo detection threshold are shown. Additionally, Figure 37 The echo detection limit threshold shown is Figure 12C An example of threshold data is shown.
[0637] Figure 37The value shown (Dir×Ref) is the value (Dir×Ref) that was determined to be "yes" in step S322, that is, the value above the threshold (Dir×Ref). Figure 37 The volume ratio (L1) of the direct tone to the reflected tone shown is relative to... Figure 37 The value shown (Dir×Ref) is multiplied by the path length ratio to obtain the value. In step S332, when the calculated volume ratio (L1) is equivalent to Figure 37 In the case of the volume ratio (L1) shown, it is determined that the volume ratio (L1) is less than the threshold (echo detection limit threshold).
[0638] In step S332, the volume ratio (L) is not used, similar to step S314. Compared to the volume ratio (L), the volume ratio (L1) can be calculated with less computation. Therefore, the computational load and computational complexity can be appropriately reduced.
[0639] Furthermore, a volume ratio (L1) less than the threshold is equivalent to the following: if the second sound is reflected by a reflective object or other reflective surface and reaches the listening position, the volume of this reflected sound is sufficiently small compared to the first volume of the first sound. Even when this reflected sound reaches the listening position, the volume of the reflected sound heard by the listener is very small. Therefore, even if the listener cannot hear the reflected sound, they rarely experience any dissonance. That is, even if the reflected sound is removed, the auditory impact on the listener is minimal.
[0640] That is, the sound signal processing method of this embodiment can suppress the dissonance felt by the listener and can appropriately reduce the amount of computation and the computational load.
[0641] reuse Figure 34 Please provide an explanation.
[0642] If the volume ratio (L1) is above the threshold (Yes in step S332), the selection unit 2302 selects the reflected sound as the reflected sound of the generating object (S342). That is, in this case, the selection unit 2302 selects the reflected sound specified in step S310. Furthermore, this step S342 is the same process as step S206.
[0643] Next, the selection unit 2302 determines whether an unspecified reflected sound exists (S343). If an unspecified reflected sound exists ("Yes" in step S343), the selection unit 2302 repeats the above process (steps S310 to S343). If no unspecified reflected sound exists ("No" in step S343), the selection unit 2302 ends the process. Furthermore, step S343 is the same process as step S208.
[0644] And, to carry out Figure 33 The processing of step S103a.
[0645] <Example 2> Next, we will explain the second example of parsing and selection processing.
[0646] Figure 38 This is a flowchart illustrating the second example of the analysis and selection processes in this embodiment. Furthermore, explanations of the commonalities with the first example are omitted or simplified here.
[0647] In the second example, steps S310 to S314 are performed. Furthermore, in the second example, steps S321 to S332 and step S342 performed in the first example are not performed.
[0648] Furthermore, steps S341 and S343 are performed.
[0649] <Example 3> Next, we will explain the third example of parsing and selection processing.
[0650] Figure 39 This is a flowchart illustrating the third example of the analysis and selection processes in this embodiment. Furthermore, explanations of the commonalities with the first example are omitted or simplified here.
[0651] In the third example, steps S310 to S322 are performed. Furthermore, in the third example, steps S331 and S332 performed in the first example are not performed.
[0652] Then, proceed with steps S341 to S343.
[0653] <Example 4> Next, we will explain the fourth example of parsing and selection processing.
[0654] Figure 40 This is a flowchart illustrating the fourth example of the analysis and selection processes in this embodiment. Furthermore, explanations of the commonalities with the first example are omitted or simplified here.
[0655] In the fourth example, steps S310 to S314 and steps S331 to S343 are performed. Furthermore, in the fourth example, steps S321 and S322 performed in the first example are not performed.
[0656] In addition, in step S331 of example 4, the volume ratio (L1) of the direct tone to the reflected tone satisfies the following equation 8.
[0657] L1 = Dir × path length ratio… (Equation 8) <Example 5> Next, we will explain the fifth example of parsing and selection processing.
[0658] Figure 41 This is a flowchart illustrating the fifth example of the analysis and selection processes in this embodiment. Furthermore, explanations of the commonalities with the first example are omitted or simplified here.
[0659] In the fifth example, step S310 is performed.
[0660] Furthermore, the analysis unit 2301 (second analysis unit 2301b) determines the reflection coefficient (Ref) (S411). Here, the second analysis unit 2301b calculates and determines the reflection coefficient (Ref) of a reflector, which is an example of the properties of an object existing in the propagation of sound. The object existing in the propagation of sound is a reflector that reflects the second sound, i.e., the reflector object (wall).
[0661] The second analysis unit 2301b calculates the volume taking into account the reflection coefficient by multiplying the reference volume by the reflection coefficient. More specifically, firstly, the second analysis unit 2301b calculates the reference volume using the same method as the first analysis unit 2301a. Additionally, the reference volume is sometimes denoted as N. Then, the second analysis unit 2301b calculates the volume taking into account the reflection coefficient by multiplying the calculated reference volume (N) by the reflection coefficient (Ref) calculated in step S411. The volume taking into account the calculated reflection coefficient satisfies Equation 9.
[0662] N×Ref …(Equation 9) Next, the second analysis unit 2301b calculates the volume ratio of the reference volume (N) to the volume (N×Ref) considering the reflection coefficient (S412). This volume ratio is recorded as NRef, which is the value obtained by dividing the volume (N×Ref) considering the reflection coefficient by the reference volume (N), satisfying Equation 10.
[0663] NRef = (N × Ref) / N … (Equation 10) The analysis unit 2301 calculates the time difference (T) between the direct tone and the reflected tone (S312). The selection unit 2302 (determination unit 2302b) uses threshold data to determine the threshold corresponding to the time difference (T) (S313).
[0664] Furthermore, the determination unit 2302b (second determination unit 2302b2) determines whether the calculated volume ratio (NRef) is above a threshold (S413). This threshold is equivalent to the echo detection limit threshold. Figure 12C An example of threshold data is shown.
[0665] If the volume ratio (NRef) is less than the threshold (No in step S413), the selection unit 2302 does not select the reflected tone (S341). That is, in this case, the selection unit 2302 does not select the reflected tone specified in step S310.
[0666] Furthermore, if the volume ratio (NRef) is above the threshold (Yes in step S413), the selection unit 2302 determines whether an unspecified reflected sound exists (S343). If an unspecified reflected sound exists (Yes in step S343), the selection unit 2302 repeats the above process (steps S310 to S413). If no unspecified reflected sound exists (No in step S343), the selection unit 2302 ends the process.
[0667] Here, the processing of step S413 is studied.
[0668] As described above, the threshold is a value determined by the determination unit 2302b based on the time difference (T) in step S313.
[0669] Furthermore, to calculate the volume ratio (NRef), a volume (N×Ref) considering the reference volume (N), the reflection coefficient (Ref), and the reflection coefficient is used. The reference volume can be considered as the volume of the sound that reaches the listening position without causing reflections, etc., while the volume considering the reflection coefficient can be considered as the volume of the sound that reaches the listening position after being reflected by a reflecting object. That is, here, the reference volume is equivalent to the volume of the sound that becomes a direct tone, that is, equivalent to the first volume of the first sound, and the volume considering the reflection coefficient is equivalent to the volume of the sound that becomes a reflected tone, that is, equivalent to the second volume of the second sound.
[0670] Based on the above, it can also be said that in the processing of step S413, the determination unit 2302b (the second determination unit 2302b2) performs the following processing: based on the time difference, the reflection coefficient as an example of the properties of an object existing in the propagation of sound, the first volume and the second volume, it determines whether to select the reflected sound.
[0671] <Example 6> Next, we will explain the sixth example of parsing and selection processing.
[0672] Figure 42 This is a flowchart illustrating the sixth example of the analysis and selection processes in this embodiment. Furthermore, the explanation of the commonalities with the fifth example is omitted or simplified here.
[0673] In the 6th example, steps S310 to S412 are performed.
[0674] Furthermore, the process of step S413 is performed. Here, if the volume ratio (NRef) is above the threshold ("Yes" in step S413), the analysis unit 2301 (the third analysis unit 2301c) calculates the value (L2) obtained by multiplying the determined reflection coefficient (Ref) by the path length ratio (S511).
[0675] The reflection coefficient (Ref) is determined by the second analysis unit 2301b in step S411, and the third analysis unit 2301c can obtain the determined reflection coefficient (Ref) before step S511.
[0676] In addition, the third parsing unit 2301c can calculate the path length ratio before step S511.
[0677] Furthermore, the determination unit 2302b (the third determination unit 2302b3) determines whether the calculated value (L2) is above a threshold determined corresponding to the time difference (T) (S512). This threshold is equivalent to the echo detection limit threshold. Figure 12C An example of threshold data is shown.
[0678] If the value (L2) is less than the threshold (No in step S512), the selection unit 2302 does not select the reflected tone (S341). That is, in this case, the selection unit 2302 does not select the reflected tone specified in step S310.
[0679] Furthermore, if the value (L2) is above the threshold ("Yes" in step S512), the selection unit 2302 selects the reflected sound as the reflected sound of the generated object (S342). That is, in this case, the selection unit 2302 selects the reflected sound specified in step S310.
[0680] Next, the selection unit 2302 determines whether there is an unspecified reflected sound (S343). If there is an unspecified reflected sound ("Yes" in step S343), the selection unit 2302 repeats the above process (steps S310 to S342). If there is no unspecified reflected sound ("No" in step S343), the selection unit 2302 ends the process.
[0681] Here, the processing of step S512 is studied.
[0682] If the reflection coefficient (Ref) is small enough, the volume of the reflected sound received by the listener will be very low even if the second sound arrives at the listening position as a reflected sound. In step S512, a value (L2) is further compared with a threshold, which is the reflection coefficient (Ref) multiplied by the path length ratio used to account for volume attenuation caused by distance.
[0683] In other words, a value (L2) less than the threshold means that even if the second sound is reflected by a reflective object or other reflective surface and reaches the listening position, the volume of the reflected sound heard by the listener is very low. Therefore, even if the listener cannot hear the reflected sound, they rarely feel any dissonance. That is, even if the reflected sound is removed, the auditory impact on the listener is small.
[0684] <Example 7> Next, we will explain the seventh example of parsing and selection processing.
[0685] Figure 43 This is a flowchart illustrating the seventh example of the analysis and selection processes in this embodiment. Furthermore, explanations of the commonalities with the first example are omitted or simplified here.
[0686] In the 7th example, the process of step S310 is performed.
[0687] The first analysis unit 2301a calculates the first volume and the second volume (S611). For example, the first analysis unit 2301a uses the directivity of sound, which is an example of the property of an object existing in the propagation of sound, to calculate the first volume and the second volume.
[0688] The second analysis section 2301b calculates and determines the reflection coefficient (Ref) of the reflector as an example of the properties of an object present in the propagation of sound (S612). The object present in the propagation of sound is the reflector that reflects the second sound, i.e., the reflector object (wall).
[0689] The analysis unit 2301 calculates the volume ratio (L3) of the direct tone to the reflected tone based on the calculated first volume, the calculated second volume, and the determined reflection coefficient (Ref) (S613). Furthermore, this volume ratio (L3) is equivalent to the volume ratio of the direct tone received by the listener to the volume of the reflected tone received by the listener. More specifically, this volume ratio (L3) is equivalent to the volume ratio of the direct tone reaching the listener to the volume of the reflected tone reaching the listener. Here, the first volume is set as A, the second volume as B, and the volume ratio (L3) is expressed by the following equation 11.
[0690] L3 = {(B × Ref) / A} × path length ratio… (Equation 11) (B×Ref) represents the volume of the second sound after it is reflected by the reflector. This volume is divided by the first volume (A) of the first sound and multiplied by the path length ratio to account for the volume attenuation caused by the distance to the listener, thus calculating the volume ratio (L3) of the direct sound to the reflected sound.
[0691] Furthermore, the determination unit 2302b determines whether the volume ratio (L3) of the direct tone to the reflected tone calculated by the analysis unit 2301 is above the threshold (S614).
[0692] If the volume ratio (L3) is less than the threshold (No in step S614), the selection unit 2302 does not select the reflected tone (S341). That is, in this case, the selection unit 2302 does not select the reflected tone specified in step S310.
[0693] If the volume ratio (L3) is above the threshold (Yes in step S614), the selection unit 2302 selects the reflected sound as the reflected sound of the generating object (S342). That is, in this case, the selection unit 2302 selects the reflected sound specified in step S310.
[0694] Furthermore, the threshold is a value corresponding to the time difference (T), which includes (B×Ref) / A in Equation 11. Therefore, it can also be said that in step S614, the determination unit 2302b determines whether to select the reflected sound based on the time difference and (B×Ref) / A.
[0695] Furthermore, in the above embodiments, in Figure 29 Examples of the configuration of the rendering unit 2300 are shown, but it is not limited to these examples. Here, we use... Figure 44 Let's explain the rendering section of another example.
[0696] Figure 44 This is a block diagram illustrating the configuration of another example of the rendering unit 3300.
[0697] The rendering unit 3300 includes an initial reflection processing unit 3310, a directional processing unit 3320, a distance attenuation processing unit 3330, a detection limit determination unit 3340, and a reproduction unit 3350.
[0698] The initial reflection processing unit 3310 has a second resolution unit 2301b and a second determination unit 2302b2. Furthermore, the initial reflection processing unit 3310 can perform the same processing as the resolution unit 2301 in Embodiment 2. The directionality processing unit 3320 has a first resolution unit 2301a and a first determination unit 2302b1. The distance attenuation processing unit 3330 has a third resolution unit 2301c and a third determination unit 2302b3.
[0699] The processing performed by the rendering unit 3300 can be performed, for example, as part of the pipeline processing described in Patent Document 3.
[0700] exist Figure 44The system includes multiple processing units (initial reflection processing unit 3310, directivity processing unit 3320, distance attenuation processing unit 3330, detection limit determination unit 3340, and reproduction unit 3350), which perform the processing for imparting sound effects executed by the rendering unit 3300. Pipeline processing refers to performing multiple processes sequentially on each sound.
[0701] However, these processes are just one example; pipeline processing may include other processes as well, or it may not include some processes at all. Furthermore, the order of the processes can also differ. Each processing unit can also be represented as a stage. Additionally, sound signals representing the results of each process, such as generated reflected sounds, can also be represented as rendering items.
[0702] In this disclosure, when the determination unit included in each processing unit decides not to select the reflected sound (e.g., in... Figure 34 If the determination in step S314, etc., is "no", the processing in subsequent processing units for the reflected sound can be skipped.
[0703] Even if the order of the various processing units is changed, the effect of this disclosure can still be achieved by having each processing unit make its own determination.
[0704] Here's an example of a better order. The initial reflection processing unit 3310 and the directivity processing unit 3320 can be processed before the distance attenuation processing unit 3330. This is because, in order for the distance attenuation processing unit 3330 to calculate the volume at the listening position, it is necessary to obtain parameters related to the directivity of the sound emitted from the sound source object and parameters related to the reflection coefficient in advance.
[0705] When the directional processing unit 3320 processes the signal before the distance attenuation processing unit 3330, if the first determination unit 2302b1 of this disclosure determines that the reflected sound will not be generated, the processing in subsequent processing units, namely the distance attenuation processing unit 3330, the detection limit determination unit 3340, and the reproduction unit 3350, can be omitted regarding the reflected sound. As a result, the processing unit, which processes the signal earlier in the pipeline processing, can determine whether to select the sound generated in the sound space, thus appropriately reducing the computational load and computational complexity.
[0706] The following is a summary of this implementation method.
[0707] The sound signal processing method of this embodiment is a sound signal processing method executed by a sound signal processing device, including a parsing step, an acquisition step, a determination step, and a reproduction step.
[0708] In the analysis step, a time difference is calculated. This time difference is the time it takes for the sound emitted from sound source 100, as a direct tone, to travel from the sound source's location (i.e., the sound source position) to the listener's location (i.e., the listening position) and the time it travels after being reflected by a reflector (i.e., the reflected tone) to the listening position. In the acquisition step, the calculated time difference and the sound signal representing the reflected tone are acquired. In the determination step, based on the acquired time difference, the properties of the object present in the sound propagation, the volume corresponding to the direction from the sound source location to the listening position (i.e., the first volume), and the volume corresponding to the direction from the sound source location to the reflector's location (i.e., the reflector's position) (i.e., the second volume), a determination is made as to whether to select the reflected tone represented by the acquired sound signal. In the reproduction step, an output signal based on a sound signal representing the selected reflected tone is output.
[0709] Therefore, based on the time difference, the property, the first volume, and the second volume, it is selected whether to output an output signal based on the sound signal representing the reflected sound. That is, it is appropriately selected whether to output an output signal based on the sound signal. By not outputting an output signal, the computational complexity and computational load can be reduced. In other words, a sound signal processing method that can appropriately reduce the computational complexity and computational load can be implemented.
[0710] In particular, in this embodiment, as shown in the first example of parsing and selection processing, the volume ratio (L) of direct tone to reflected tone used in step S205 of embodiment 1 is not used. As mentioned above, calculating this volume ratio (L) may increase the computational load, but in this embodiment, since the volume ratio (L) is not used, the computational load and computational burden can be appropriately reduced.
[0711] Furthermore, in this embodiment, a threshold is used to determine whether to select the reflected sound. As shown in Example 1, when the volume ratio (Dir), value (Dir×Ref), or volume ratio (L1) is less than the threshold, the volume of the reflected sound heard by the listener is very low, and the listener rarely feels dissonance even if they cannot hear the reflected sound. That is, the sound signal processing method of this embodiment can suppress dissonance felt by the listener and can appropriately reduce the computational load.
[0712] In the sound signal processing method of this embodiment, the direction from the sound source position to the position of the reflector, i.e., the position of the reflector, is calculated based on the direction from the position of the mirror image of the sound source formed by the reflector to the listening position.
[0713] Therefore, it is possible to calculate the direction from the sound source location to the reflector location based on the direction from the mirror image location of the sound source to the listening location.
[0714] In the sound signal processing method of this embodiment, the property is the directionality of the sound. In the determination step, it is determined whether to select the reflected sound based on the obtained time difference, the first volume calculated based on the directionality, and the second volume calculated based on the directionality.
[0715] Therefore, based on this time difference and the first and second volume levels calculated according to the directivity, it is determined whether to output an output signal. That is, it is possible to more appropriately select whether to output an output signal based on this sound signal. In other words, it is possible to implement a sound signal processing method that can more appropriately reduce the computational load.
[0716] In the sound signal processing method of this embodiment, in the determination step, it is determined whether to select reflected sound based on the obtained time difference and the volume ratio of the first volume calculated based on directivity to the second volume calculated based on directivity.
[0717] Therefore, based on this time difference and the volume ratio of the first volume to the second volume calculated according to directivity, it is selected whether to output an output signal. That is, it is possible to more appropriately select whether to output an output signal based on this sound signal. In other words, it is possible to realize a sound signal processing method that can more appropriately reduce the amount of computation and computational load.
[0718] In the sound signal processing method of this embodiment, the property is the reflection coefficient of the reflector.
[0719] Therefore, as shown in Example 5 of the analytical and selection processing, based on the time difference, the reflection coefficient of the reflector, the first volume, and the second volume, it is selected whether to output an output signal based on the sound signal representing the reflected sound. That is, it is more appropriate to select whether to output an output signal based on the sound signal. In other words, it is possible to realize a sound signal processing method that can more appropriately reduce the computational load.
[0720] In the sound signal processing method of this embodiment, the properties are the directivity of the sound and the reflection coefficient of the reflector. In the determination step, a determination is made as to whether to select the reflected sound based on the obtained time difference, the reflection coefficient, the first volume calculated based on the directivity, and the second volume calculated based on the directivity.
[0721] Therefore, based on the time difference, the reflection coefficient, and the volume ratio of the first volume to the second volume calculated according to the directivity, it is determined whether to output an output signal. That is, it is possible to more appropriately select whether to output an output signal based on the sound signal. In other words, it is possible to realize a sound signal processing method that can more appropriately reduce the computational load.
[0722] In the sound signal processing method of this embodiment, when the first volume is A, the second volume is B, and the reflection coefficient is Ref, in the determination step, a determination is made on whether to select the reflected sound based on the obtained time difference and the following formula. (B×Ref) / A.
[0723] Therefore, based on this time difference and the above formula, it is possible to choose whether to output an output signal. That is, to more appropriately choose whether to output an output signal based on this sound signal. In other words, it is possible to realize a sound signal processing method that can more appropriately reduce the amount of computation and computational load.
[0724] In the sound signal processing method of this embodiment, when the volume ratio of the first volume and the second volume is set to Dir and the reflection coefficient is set to Ref, in the determination step, a determination is made on whether to select the reflected sound based on the obtained time difference and the following formula. Dir×Ref.
[0725] Therefore, based on this time difference and the above formula, it is possible to choose whether to output an output signal. That is, to more appropriately choose whether to output an output signal based on this sound signal. In other words, it is possible to realize a sound signal processing method that can more appropriately reduce the amount of computation and computational load.
[0726] In the sound signal processing method of this embodiment, in the determination step, it is further determined whether to select reflected sound based on the position of the object present in the sound propagation.
[0727] Therefore, the decision to output an output signal is made based on the positions of the sound source 100, the reflector, and the listener, all of which are objects present in the propagation of sound. In other words, it allows for a more appropriate selection of whether to output an output signal based on the sound signal. This enables a sound signal processing method that can more appropriately reduce computational complexity and workload.
[0728] In the sound signal processing method of this embodiment, in the determination step, it is determined whether to select the reflected sound based on the path length ratio of the direct sound path length to the reflected sound path length. The direct sound path length is the path length of the direct sound to the listening position calculated based on the position of the object, and the reflected sound path length is the path length of the reflected sound to the listening position calculated based on the position of the object.
[0729] Therefore, the decision to output a signal is also based on the path length ratio. That is, it allows for a more appropriate selection of whether to output a signal based on the audio signal. In other words, it enables an audio signal processing method that can more appropriately reduce computational complexity and workload.
[0730] The computer program in this embodiment is a computer program used to enable a computer to execute the above-described sound signal processing method.
[0731] Therefore, the computer can execute the above-mentioned sound signal processing method according to the computer program.
[0732] The sound signal processing apparatus according to this embodiment includes a parsing unit 2301, an acquisition unit 2302a, a determination unit 2302b, and a reproduction unit 2303.
[0733] The analysis unit 2301 calculates the time difference, which is the time difference between the moment when the sound emitted from the sound source 100 as a direct tone travels from the sound source position (i.e., the sound source location) to the moment when the sound is reflected by a reflector and becomes a reflected tone, reaching the listening position. The acquisition unit 2302a acquires the calculated time difference and the sound signal representing the reflected tone. The determination unit 2302b determines whether to select the reflected tone represented by the acquired sound signal based on the acquired time difference, the properties of the object present in the sound propagation, the volume corresponding to the direction from the sound source position to the listening position (i.e., a first volume), and the volume corresponding to the direction from the sound source position to the reflector position (i.e., the reflector location) (i.e., a second volume). The reproduction unit 2303 outputs an output signal based on a sound signal representing the selected reflected tone.
[0734] Therefore, based on the time difference, the property, the first volume, and the second volume, it is selected whether to output an output signal based on the sound signal representing the reflected sound. That is, it is appropriately selected whether to output an output signal based on the sound signal. When no output signal is output, the computational load and computational complexity can be reduced. In other words, a sound signal processing device that can appropriately reduce the computational load and computational complexity can be realized.
[0735] (Replenish) Furthermore, the solutions disclosed herein are not limited to specific implementation methods and can be implemented with various modifications.
[0736] For example, in an implementation, a process that is performed by a specific component may be performed by other components instead of that specific component. Furthermore, the order of multiple processes may be changed, or multiple processes may be performed in parallel.
[0737] Furthermore, the ordinal numbers 1, 2, etc., used in the description can be appropriately replaced, removed, or newly assigned. These ordinal numbers do not necessarily correspond to a meaningful order and can also be used for element identification.
[0738] Furthermore, for example, in comparisons of thresholds, "above the threshold" and "greater than the threshold" can be interchanged. Similarly, "below the threshold" and "less than the threshold" can be interchanged. Additionally, for example, "time" and "moment" can be interchanged.
[0739] Furthermore, in the process of selecting more than one target sound from multiple sounds, if no sound meets the conditions, then none of the sounds may be selected as target sounds. That is, the process of selecting more than one target sound from multiple sounds may also include the case of not selecting any target sound.
[0740] Additionally, for example, at least one of the first element, the second element, and the third element may correspond to the first element, the second element, the third element, or any combination thereof.
[0741] Furthermore, in Example 5, it was explained that the determination unit 2302b (the second determination unit 2302b2) performs a process to determine whether to select the reflected sound based on the time difference, the reflection coefficient (an example of the properties of an object present in the propagation of sound), the first volume, and the second volume, but it is not limited to this. For example, the process of step S412 may be omitted, and the determination unit 2302b (the second determination unit 2302b2) may perform the following process instead of step S413. The second determination unit 2302b2 performs a process to determine whether the determined reflection coefficient (Ref) is above a threshold.
[0742] The following is a summary of the process.
[0743] A sound signal processing method, executed by a sound signal processing device, includes: a parsing step, calculating a time difference, the time difference being the time difference between the moment when a sound emitted from a sound source as a direct sound travels from the location of the sound source (i.e., the sound source position) to the moment when the sound is reflected by a reflective object and becomes a reflected sound, and arrives at the listening position; an acquisition step, acquiring the calculated time difference and a sound signal representing the reflected sound; a determination step, based on the acquired time difference and the properties of objects present in the propagation of the sound, determining whether to select the reflected sound represented by the acquired sound signal; and a reproduction step, outputting an output signal based on a sound signal representing the selected reflected sound.
[0744] In this process, volume ratio (L) is not used. Therefore, the computational load and computational complexity can be appropriately reduced.
[0745] Furthermore, if the reflection coefficient (Ref) is less than the threshold and sufficiently small, even if the second sound arrives at the listening position as a reflected tone, the volume of the reflected tone heard by the listener will be very low. Therefore, even if the listener cannot hear the reflected tone, they will rarely feel any dissonance. That is, even if the reflected tone is removed, the auditory impact on the listener will be minimal.
[0746] That is, the audio signal processing method described in this paper can suppress the dissonance felt by the listener and can appropriately reduce the amount of computation and the computational load.
[0747] Furthermore, for example, the embodiments described involve implementing the solutions based on this disclosure as a sound signal processing apparatus, encoding apparatus, or decoding apparatus. However, the solutions based on this disclosure are not limited thereto, and can also be implemented as software for performing sound signal processing methods, encoding methods, or decoding methods.
[0748] For example, the program for performing the aforementioned sound signal processing, encoding, or decoding methods can also be pre-stored in ROM. Furthermore, the CPU can also execute the actions according to this program.
[0749] Alternatively, the program used to perform the aforementioned audio processing, encoding, or decoding methods can be stored in a computer-readable recording medium. Furthermore, the computer can also record the program stored in the recording medium into its RAM and operate according to that program.
[0750] Furthermore, the aforementioned components can typically be implemented as integrated circuits, i.e., LSIs, which have input and output terminals. They can be formed as a single chip or as a chip incorporating all or some of the components of the implementation. Depending on the level of integration, LSIs can also be classified as ICs, system LSIs, super LSIs, or very large-scale LSIs.
[0751] Furthermore, it is not limited to LSIs; dedicated circuits or general-purpose processors can also be used. Additionally, FPGAs that can be programmed after LSI manufacturing, or reconfigurable processors capable of reconfiguring the connection or configuration of circuit units within the LSI, can also be used. Moreover, if advancements in semiconductor technology or other derived technologies lead to integrated circuit technologies that replace LSIs, then these technologies can certainly be used for the integration of constituent elements. This could include applications in biotechnology, among others.
[0752] Furthermore, an FPGA or CPU can download all or part of the software used to implement the sound signal processing, encoding, or decoding methods described in this disclosure via wireless or wired communication. Additionally, all or part of the software for updates can be downloaded via wireless or wired communication. Moreover, the digital signal processing described in this disclosure can be executed by storing the downloaded software in a memory using an FPGA or CPU and operating based on the stored software.
[0753] At this time, devices equipped with FPGAs or CPUs can also be connected to the signal processing device wirelessly or via wired connection, or connected to the signal processing server via a network. Furthermore, the device and the signal processing device or signal processing server can perform the audio signal processing, encoding, or decoding methods described in this disclosure.
[0754] For example, the audio signal processing apparatus, encoding apparatus, or decoding apparatus of this disclosure may also include an FPGA or a CPU. Furthermore, the audio signal processing apparatus, encoding apparatus, or decoding apparatus may also include an interface for obtaining software from an external source to operate the FPGA or CPU, and a memory for storing the obtained software. Moreover, the FPGA or CPU may execute the signal processing described in this disclosure based on the stored software operations.
[0755] Alternatively, the server may provide software related to the audio processing, encoding, or decoding processes disclosed herein. Furthermore, a terminal or device may operate as the audio signal processing, encoding, or decoding apparatus described in this disclosure by installing the software. Alternatively, the terminal or device may install the software by connecting to the server via a network.
[0756] Alternatively, a device other than the terminal or device may obtain the software installation data via a network connection to a server, and install the software on the terminal or device by providing the software installation data to the terminal or device through this other device. Another example of software could be VR software or AR software used to enable the terminal or device to execute the sound signal processing method described in the embodiments.
[0757] Furthermore, in the above embodiments, each component may be constructed using dedicated hardware, or implemented by executing software programs suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.
[0758] The above description describes apparatuses and the like with respect to one or more embodiments based on the implementation methods, but the solutions available in this disclosure are not limited to the implementation methods. Solutions that can be obtained by applying various modifications to the implementation methods that can be conceived by those skilled in the art, as well as solutions constructed by combining the constituent elements of different modifications, are also included within the scope of one or more solutions.
[0759] Industrial applicability This disclosure includes, for example, solutions that can be applied in a sound processing apparatus, an encoding apparatus, a decoding apparatus, or a terminal or device having any of these apparatuses.
[0760] Label Explanation 100 sound sources 110 Virtual Image (Mirror Image) 1000 Stereo Sound Reproduction System 1001 Sound signal processing device (audio processing device) 1002 Sound Prompt Device 1100, 1120, 1500 encoding devices Input data for 1101 and 1113 1102 Encoder 1103 Encoded Data 1104, 1114, 1404, 1503 memory 1110, 1130 Decoding Devices 1111 sound signal 1112, 1200, 1210 decoders 1121 Sending Department 1122 Send signal 1131 Receiving Department 1132 Received signal Spatial Information Management Department, 1201 & 1211 1202 Audio Data Decoder Rendering Departments: 1203, 1213, 1300, 2300, 3300 Analysis Departments 1301 and 2301 1302, 1314, 2302 Selection Department Reproduction section 1303, 2303, 3350 1304 Threshold Adjustment Section 1311 Reverb Processing Department 1312, 3310 Initial Reflection Processing Section 1313, 3330 Distance Attenuation Processing Unit 1315 Production Department 1316 Binocular Processing Unit 1401 Speaker 1402 and 1501 processors 1403, 1502 Communication IF 1405 sensor 2301a First Analysis Section 2301b, Section 2 of Analysis 2301c Third Analysis Section 2302a Acquisition Department 2302b Judgment Department 2302b1 First Judgment Department 2302b2 Second Judgment Department 2302b3 Third Judgment Section 3320 Directivity Processing Unit 3340 Detection Limit Judgment Department
Claims
1. A sound signal processing method, executed by a sound signal processing device, wherein, include: The analysis steps include calculating the time difference, which is the time difference between the time it takes for the sound emitted from the sound source to travel as a direct sound from the location of the sound source to the location of the listener, i.e., the listening position, and the time it takes for the sound to travel as a reflected sound after being reflected by a reflective object to the listening position. The step involves obtaining the calculated time difference and the sound signal representing the reflected sound; The determination step, based on the acquired time difference, the properties of the object present in the sound propagation, a first volume, and a second volume, determines whether to select the reflected sound represented by the acquired sound signal. The first volume is the volume corresponding to the direction from the sound source location to the listening location, and the second volume is the volume corresponding to the direction from the sound source location to the position of the reflecting object. The reproduction step outputs an output signal based on a sound signal representing the reflected sound that has been determined to be selected.
2. The sound signal processing method according to claim 1, wherein, The direction from the sound source location to the position of the reflector is calculated based on the direction from the position of the mirror image of the sound source formed by the reflector to the listening position.
3. The sound signal processing method according to claim 1, wherein, The property in question is the directivity of the sound. In the determination step, based on the obtained time difference, the first volume calculated according to the directivity, and the second volume calculated according to the directivity, it is determined whether to select the reflected sound.
4. The sound signal processing method according to claim 3, wherein, In the determination step, based on the obtained time difference and the volume ratio of the first volume calculated according to the directivity to the second volume calculated according to the directivity, it is determined whether to select the reflected sound.
5. The sound signal processing method according to claim 1, wherein, The property is the reflectivity of the reflector.
6. The sound signal processing method according to claim 1, wherein, The properties are the directivity of the sound and the reflection coefficient of the reflector. In the determination step, based on the obtained time difference, the reflection coefficient, the first volume calculated according to the directivity, and the second volume calculated according to the directivity, it is determined whether to select the reflected sound.
7. The sound signal processing method according to claim 6, wherein, When the first volume is A, the second volume is B, and the reflection coefficient is Ref, In the determination step, based on the obtained time difference and Equation 1, it is determined whether to select the reflected sound. (B×Ref) / A … (Equation 1).
8. The sound signal processing method according to claim 6, wherein, When the volume ratio of the first volume to the second volume is Dir, and the reflection coefficient is Ref, In the determination step, based on the obtained time difference and Equation 2, it is determined whether to select the reflected sound. Dir×Ref …(Equation 2) 9. The sound signal processing method according to any one of claims 1 to 8, wherein, In the determination step, it is also determined whether to select the reflected sound based on the position of the object present in the propagation of the sound.
10. The sound signal processing method according to claim 9, wherein, In the determination step, the decision is made on whether to select the reflected sound based on the path length ratio of the direct sound path length to the reflected sound path length. The direct sound path length is the path length of the direct sound to the listening position calculated based on the position of the object, and the reflected sound path length is the path length of the reflected sound to the listening position calculated based on the position of the object.
11. A computer program for causing a computer to execute the sound signal processing method according to any one of claims 1 to 10.
12. A sound signal processing device, wherein, have: The analysis unit calculates the time difference, which is the time difference between the time when the sound emitted from the sound source as a direct sound reaches the listener's location (i.e., the sound source position) and the time when the sound is reflected by a reflector and reaches the listener's location (i.e., the listening position). The acquisition unit acquires the calculated time difference and the sound signal representing the reflected sound; The determination unit, based on the acquired time difference, the properties of the object present in the sound propagation, a first volume, and a second volume, determines whether to select the reflected sound represented by the acquired sound signal, wherein the first volume is the volume corresponding to the direction from the sound source position to the listening position, and the second volume is the volume corresponding to the direction from the sound source position to the position of the reflecting object, i.e., the position of the reflecting object; and The reproduction unit outputs an output signal based on a sound signal representing the reflected sound that has been selected.
Citation Information
Patent Citations
Signal processor
JP2019022049A
Apparatus and method for rendering a sound scene using pipeline stages
WO2021180938A1