Audio signal processing method, computer program, and audio signal processing device
By determining the properties of direct and reflected sound signals in virtual reality and augmented reality systems, unnecessary computational steps are reduced, solving the problems of computational load and complexity in virtual reality and augmented reality, extending battery life, and improving sound quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
- Filing Date
- 2024-10-04
- Publication Date
- 2026-04-28
AI Technical Summary
Existing technologies struggle to adequately reduce the computational load and complexity of sound signal processing in virtual reality and augmented reality, especially when dealing with multiple sound sources and reflected sounds, leading to shortened battery life and wasted computing resources.
By determining the attributes of sound signals, direct sound and reflected sound are distinguished, and only reflected sound that meets specific conditions is further processed and output, reducing unnecessary computational steps and workload.
It effectively reduces the computational load and computational complexity in virtual reality and augmented reality systems, extends battery life, and improves sound quality, adapting to immersive audio experiences in multi-source and complex environments.
Smart Images

Figure CN121942218A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to methods for processing sound signals, etc. Background Technology
[0002] In recent years, goods and services utilizing ER (Extended Reality), including VR (Virtual Reality), AR (Augmented Reality), and MR (Mixed Reality), have become increasingly popular. Consequently, the importance of sound signal processing technologies—which assign acoustic effects corresponding to the environment of a virtual sound source to the sound emitted in virtual or real space, thereby providing listeners with immersive audio—has increased.
[0003] Additionally, the listener can also be a listener or a user. Furthermore, technologies related to the sound signal processing method disclosed herein are shown in Patent Document 1, Patent Document 2, Patent Document 3, and Non-Patent Document 1.
[0004] Existing technical documents Patent documents Patent Document 1: Japanese Patent No. 6288100 Patent Document 2: Japanese Patent Application Publication No. 2019-22049 Patent Document 3: International Publication No. 2021 / 180938 Non-patent literature Non-Patent Literature 1: BCJ Moore, *An Outline of Auditory Psychology*, Chengxin Bookstore, April 20, 1994, Chapter 6: Spatial Perception, p. 225 Summary of the Invention
[0005] The problem that the invention aims to solve However, in the technology shown in Patent Document 1, it is sometimes difficult to properly reduce the amount of computation and the computational load.
[0006] Therefore, the purpose of this disclosure is to provide a sound signal processing method that can appropriately reduce computational load and computational complexity.
[0007] Methods used to solve problems One aspect of this disclosure is a sound signal processing method executed by a sound signal processing apparatus, comprising: an acquisition step, acquiring a sound signal, the sound signal including attribute information that determines the attributes of the sound signal; a determination step, in which, if the attribute determined by the attribute information contained in the acquired sound signal is information representing an indirect tone, a first determination process is performed to determine whether the acquired sound signal satisfies the first condition, and a second determination process is performed to determine whether the acquired sound signal satisfies the second condition, which is different from the first condition; and a reproduction step, in which, if the acquired sound signal satisfies the first condition and the second condition, an output signal based on the acquired sound signal is output.
[0008] One aspect of this disclosure is a sound signal processing method executed by a sound signal processing apparatus, comprising: an acquisition step, acquiring a sound signal representing an indirect tone; a rendering pipeline step, performing one or more first processing, a first determination processing, a second determination processing, and one or more second processing different from the one or more first processing on the acquired sound signal; a reproduction step, outputting an output signal based on the acquired sound signal; a first gain accumulation step, calculating a first gain / loss amount by accumulating the gain / loss amount determined by the one or more first processing, wherein the gain / loss amount determined by the one or more first processing is a gain / loss amount that amplifies the amplitude of the acquired sound signal; and a second gain accumulation step, calculating a gain / loss amount by accumulating the gain / loss amount determined by the one or more second processing. The second increment / decrease amount, determined by the one or more second processes, is the increment / decrease amount that amplifies the amplitude of the acquired sound signal. In the first determination process, if the calculated first increment / decrease amount is above a first threshold, the acquired sound signal is determined to satisfy the first condition. If the acquired sound signal is determined not to satisfy the first condition, the second determination process is not performed. If the acquired sound signal is determined to satisfy the first condition, the second determination process is performed. In the second determination process, if the value corresponding to the calculated second increment / decrease amount is above a second threshold that is different from the first threshold, the acquired sound signal is determined to satisfy the second condition. If the acquired sound signal is determined to satisfy the second condition, the reproduction step is performed.
[0009] In addition, a computer program of one aspect of this disclosure causes a computer to perform the above-described sound signal processing method.
[0010] Additionally, one aspect of the sound signal processing apparatus disclosed herein includes: an acquisition unit that acquires a sound signal, the sound signal including attribute information that determines an attribute of the sound signal; a determination unit that, when the attribute determined by the attribute information included in the acquired sound signal is information representing an indirect tone, performs a first determination process to determine whether the acquired sound signal satisfies a first condition and a second determination process to determine whether the acquired sound signal satisfies a second condition different from the first condition; and a reproduction unit that, when the acquired sound signal satisfies the first condition and the second condition, outputs an output signal based on the acquired sound signal.
[0011] In addition, these inclusive or specific technical solutions can also be implemented by non-transitory recording media such as systems, devices, methods, integrated circuits, computer programs, or computer-readable CD-ROMs, or by any combination of systems, devices, methods, integrated circuits, computer programs, and recording media.
[0012] Invention Effects According to one aspect of the sound signal processing method disclosed herein, the computational load and computational complexity can be appropriately reduced. Attached Figure Description
[0013] Figure 1 This is a diagram illustrating an example of direct and reflected tones generated in sound space.
[0014] Figure 2 This is a diagram illustrating an example of the stereo sound reproduction system of Embodiment 1.
[0015] Figure 3A This is a block diagram illustrating an example of the configuration of the encoding device in Embodiment 1.
[0016] Figure 3B This is a block diagram illustrating an example of the configuration of the decoding device in Embodiment 1.
[0017] Figure 3C This is a block diagram illustrating another configuration example of the encoding device according to Embodiment 1.
[0018] Figure 3D This is a block diagram illustrating another configuration example of the decoding device according to Embodiment 1.
[0019] Figure 4A This is a block diagram illustrating a configuration example of the decoder in Implementation Method 1.
[0020] Figure 4B This is a block diagram illustrating another configuration example of the decoder in Implementation 1.
[0021] Figure 5This is a diagram illustrating an example of the physical configuration of the sound signal processing device in Embodiment 1.
[0022] Figure 6 This is a diagram illustrating an example of the physical configuration of the encoding device in Embodiment 1.
[0023] Figure 7 This is a block diagram showing an example of the configuration of the rendering unit in Embodiment 1.
[0024] Figure 8 This is a flowchart illustrating an example of the operation of the sound signal processing device in Embodiment 1.
[0025] Figure 9 It is a diagram that shows the relative positional relationship between the listener and the obstacle object.
[0026] Figure 10 It is a diagram that shows the relative positional relationship between the listener and the obstacle object.
[0027] Figure 11 It is a graph showing the relationship between the time difference and threshold of direct and reflected tones.
[0028] Figure 12A This is a diagram that represents part of an example of how threshold data is set.
[0029] Figure 12B This is a diagram that represents part of an example of how threshold data is set.
[0030] Figure 12C This is a diagram that represents part of an example of how threshold data is set.
[0031] Figure 13 This is a diagram illustrating an example of how a threshold is set.
[0032] Figure 14 This is a flowchart representing an example of a selection process.
[0033] Figure 15 It is a graph showing the relationship between the direction of the direct tone, the direction of the reflected tone, the time difference, and the threshold.
[0034] Figure 16 It is a graph showing the relationship between angle difference, time difference, and threshold.
[0035] Figure 17 This is a block diagram representing another example of the rendering unit.
[0036] Figure 18 This is a flowchart representing another example of the selection process.
[0037] Figure 19 This is another flowchart representing the selection process.
[0038] Figure 20 This is a flowchart illustrating a first variation of the operation of the sound signal processing apparatus according to Embodiment 1.
[0039] Figure 21 This is a flowchart illustrating a second variation of the operation of the sound signal processing device in Embodiment 1.
[0040] Figure 22 This is a diagram showing a configuration example of avatars, sound source objects, and obstacle objects.
[0041] Figure 23 This is another flowchart representing the selection process.
[0042] Figure 24 This is a block diagram illustrating a configuration example for pipeline processing in the rendering unit.
[0043] Figure 25 It is a diagram showing the transmission and diffraction of sound.
[0044] Figure 26 This is a block diagram showing an example of the configuration of the rendering unit in Embodiment 2.
[0045] Figure 27 This is a flowchart illustrating an example of the operation of the sound signal processing device in Embodiment 2.
[0046] Figure 28 This is a graph representing the threshold data of the first threshold in Implementation Method 2.
[0047] Figure 29 This is a block diagram illustrating an example of the configuration of the rendering unit in Embodiment 3.
[0048] Figure 30 This is a block diagram illustrating an example of the configuration of the rendering unit in Embodiment 4.
[0049] Figure 31 This is a graph illustrating the energy of the head-related transfer function in Implementation Method 4.
[0050] Figure 32 This is a graph illustrating the energy of the head-related transfer function in Implementation Method 4.
[0051] Figure 33 This is a graph illustrating the energy of the head-related transfer function in Implementation Method 4.
[0052] Figure 34 This is a graph illustrating the energy of the head-related transfer function in Implementation Method 4.
[0053] Figure 35 This is a flowchart illustrating an example of the operation of the sound signal processing device in Embodiment 4.
[0054] Figure 36 This is a block diagram illustrating an example of the configuration of the rendering unit in Embodiment 5.
[0055] Figure 37 This is a diagram showing a table used to illustrate the effects of the change in the first threshold and the invalidation process in Implementation 5. Detailed Implementation
[0056] (The insights that form the basis of this disclosure) Previous studies have explored sound signal processing techniques that imbue virtual sound sources with acoustic effects based on the environment of the space in virtual or real spaces and provide immersive audio to listeners.
[0057] Such a sound signal processing technique is disclosed in Patent Document 1. More specifically, Patent Document 1 discloses a technique for detecting the importance of an audio signal (sound signal) and not outputting audio signals with low detected importance. In this way, by not outputting audio signals with low importance, it is expected that the computational load and computational complexity will be appropriately reduced in this sound signal processing technique.
[0058] Additionally, reflected sound sometimes becomes important in sound space (virtual or real space).
[0059] Figure 1 This diagram illustrates an example of direct and reflected sounds generated in a sound space. In sound processing that uses sound to represent the characteristics of a virtual space, in order to represent the breadth of the space and the material of the walls, as well as to accurately determine the location of the sound source (sound image localization), it is effective to reproduce not only direct sounds but also reflected sounds.
[0060] For example, in such Figure 1 When listening to sound within a rectangular room, a sound source produces six primary reflections corresponding to the six walls. The reproduction of these reflections provides clues for a proper understanding of the space and the sound image. Furthermore, for each reflection, secondary reflections are produced on surfaces other than the one that produced the reflection. These reflections also serve as perceptually valid clues.
[0061] However, even considering only the case of secondary reflection, a sound source will produce 1 direct tone and 36 (6+6×5) reflected tones, resulting in 37 vocal lines. Processing these vocal lines requires a considerable amount of computation.
[0062] Furthermore, in recent years, applications related to the concept of the Metaverse, such as virtual meetings, virtual shopping, or virtual concerts, have inevitably involved multiple sound sources, thus requiring a much larger amount of computation.
[0063] Furthermore, listeners in virtual spaces use headphones or VR goggles to receive audio. To provide stereo sound to such listeners, binaural processing is performed on each sound ray, assigning sound pressure levels and phase differences between the two ears to reproduce the direction of arrival and the sense of distance. Therefore, the computational load is extremely high if all the reflected sounds are to be reproduced.
[0064] On the other hand, small rechargeable batteries are sometimes used as batteries for VR goggles worn by listeners experiencing virtual space, due to their convenience. To extend battery life, the computational load required for the processing described above should ideally be low. Therefore, it is desirable to reduce the number of sound lines generated on a scale of several hundred without compromising sound localization and spatial accuracy.
[0065] Furthermore, in systems that reproduce sound, there are sometimes 6 degrees of freedom (DoF) allowed for the listener's position (i.e., the listening position as the listener's location) and orientation. In such cases, the positional relationship between the listener, the sound source, and the object reflecting the sound cannot be determined if it is not during reproduction (rendering). Therefore, the reflected sound also cannot be determined if it is not during reproduction. Consequently, it is difficult to predetermine the reflected sound of the object being processed.
[0066] Therefore, the appropriate selection and output (reproduction) of one or more reflected sounds from multiple reflected sounds generated in the sound space during reproduction is beneficial for the appropriate reduction of computational load.
[0067] Furthermore, controlling whether to select a sound corresponds to determining whether to select a sound, and more specifically, to determining whether to select and output (reproduce) a sound. Additionally, selecting a sound can mean selecting it as the target sound for processing, or selecting it as a non-target sound.
[0068] Furthermore, while Patent Document 1 detects the importance of an audio signal, more specifically the importance of the direct tone represented by that audio signal, it does not investigate the importance of reflected tones. Therefore, in cases such as Figure 1 When indirect sounds such as reflected sounds are generated, the computational load and computational complexity will increase, making it sometimes difficult to appropriately reduce the computational load and computational complexity.
[0069] Therefore, there is a need for sound signal processing methods that can appropriately reduce computational load in the sound space.
[0070] Therefore, the first aspect of the sound signal processing method disclosed herein is a sound signal processing method executed by a sound signal processing device, comprising: an acquisition step, acquiring a sound signal, the sound signal containing attribute information that determines the attributes of the sound signal; a determination step, in which, if the attribute determined by the attribute information contained in the acquired sound signal is information representing an indirect tone, a first determination process is performed to determine whether the acquired sound signal satisfies a first condition, and a second determination process is performed to determine whether the acquired sound signal satisfies a second condition different from the first condition; and a reproduction step, in which, if the acquired sound signal satisfies the first condition and the second condition, an output signal based on the acquired sound signal is output.
[0071] Therefore, a first determination process and a second determination process are performed on a sound signal with the attribute of indirect tone, and an output signal based on the sound signal obtained under the condition that the sound signal satisfies the first and second conditions is output. That is, it is appropriately determined whether to output an output signal based on the sound signal with the attribute of indirect tone. When no output signal is output, the computational complexity and computational load are reduced. That is, a sound signal processing method that can appropriately reduce the computational complexity and computational load can be implemented.
[0072] The second method of sound signal processing disclosed herein is the first method of sound signal processing, wherein the indirect tone is a reflected tone.
[0073] This enables a sound signal processing method that can use reflected sound as indirect sound.
[0074] The third method of sound signal processing disclosed herein is a method of sound signal processing of the first or second method, wherein, if the attribute determined by the attribute information contained in the acquired sound signal is information representing a prescribed tone different from the indirect tone, the second determination process is not performed in the determination step.
[0075] Therefore, the second judgment process is not performed on sound signals with the specified tone attribute, thus further reducing the computational load. In other words, a sound signal processing method that can appropriately reduce the computational load is realized.
[0076] The fourth method of sound signal processing disclosed herein is the third method of sound signal processing, wherein the specified tone is a direct tone.
[0077] Thus, a sound signal processing method that can use direct tones as specified tones is realized.
[0078] The fifth method of sound signal processing disclosed herein is a sound signal processing method of any one of the first to fourth methods, wherein, in the first determination process, if the amplitude value of the obtained sound signal is above a first threshold, it is determined that the obtained sound signal satisfies the first condition; and in the second determination process, if the volume ratio of the direct tone and the indirect tone when the direct tone and the indirect tone reach the listener's location (i.e., the listening location) is above a second threshold determined based on the arrival time difference between the direct tone and the indirect tone, it is determined that the obtained sound signal satisfies the second condition.
[0079] Therefore, in the first determination process, if the amplitude value is above a first threshold, the sound signal is determined to satisfy the first condition; in the second determination process, if the volume ratio is above a second threshold, the sound signal is determined to satisfy the second condition. An output signal based on the sound signal obtained under the condition of satisfying this second condition is output. That is, it more appropriately determines whether to output an output signal based on the sound signal. In other words, it enables a sound signal processing method that can more appropriately reduce computational complexity and computational load.
[0080] The sixth method of sound signal processing disclosed herein is a sound signal processing method of any one of the first to fifth methods, wherein the reproduction step includes: a first reproduction step, outputting a first output signal based on the sound signal whose attribute is a specified tone different from the indirect tone; and a second reproduction step, outputting a second output signal based on the sound signal whose attribute is the indirect tone, wherein in the second reproduction step, the second output signal is output by performing diffusion filtering on the acquired sound signal, wherein the diffusion filtering improves the realism of the indirect tone by diffusing the indirect tone, and in the first reproduction step, the diffusion filtering is not performed on the acquired sound signal, and the first output signal is output.
[0081] Therefore, by performing diffusion filtering on the sound signal with the attribute of indirect tone and then outputting it, the listener can hear the indirect tone with higher sound quality. Furthermore, since diffusion filtering is not performed on the sound signal with the attribute of a specific tone, the computational load and computational complexity are further reduced. In other words, a sound signal processing method that can appropriately reduce the computational load and computational complexity can be implemented.
[0082] The seventh method of sound signal processing disclosed herein is a sound signal processing method of any one of the first to sixth methods, wherein the second determination process is performed after the first determination process is performed.
[0083] Thus, a sound signal processing method is realized that can perform a second decision process after the first decision process.
[0084] Performing the determination process in this order produces the special effect described below. The main processing of the first determination process is determining the relationship between the amplitude value of the sound signal to be the target and a predetermined threshold, so its computational load is extremely light. On the other hand, since the second determination process also needs to calculate the volume ratio, arrival time difference, etc., between the indirect tone to be the target and the direct tone related to that indirect tone, its computational load is significantly greater than that of the first determination process. Therefore, in the first determination process with its light computational load, multiple input signals are first filtered, and only the remaining signals are determined in the second determination process. As a result, the overall computational load of the determination steps can be greatly reduced. Therefore, in this disclosure, whose original purpose is to reduce the computational load of sound signal processing, the order of this process is of utmost importance.
[0085] The eighth aspect of this disclosure is a sound signal processing method executed by a sound signal processing apparatus, comprising: an acquisition step, acquiring a sound signal representing an indirect tone; a rendering pipeline step, performing one or more first processing, a first determination processing, a second determination processing, and one or more second processing different from the one or more first processing on the acquired sound signal; a reproduction step, outputting an output signal based on the acquired sound signal; a first gain accumulation step, calculating a first gain / loss amount by accumulating the gain / loss amount determined by the one or more first processing, wherein the gain / loss amount determined by the one or more first processing is a gain / loss amount that amplifies the amplitude of the acquired sound signal; and a second gain accumulation step, calculating a gain / loss amount by accumulating the gain / loss amount determined by the one or more second processing. The second increment / decrease amount, determined by the one or more second processes, is the increment / decrease amount that amplifies the amplitude of the acquired sound signal. In the first determination process, if the calculated first increment / decrease amount is above a first threshold, the acquired sound signal is determined to satisfy the first condition. If the acquired sound signal is determined not to satisfy the first condition, the second determination process is not performed. If the acquired sound signal is determined to satisfy the first condition, the second determination process is performed. In the second determination process, if the value corresponding to the calculated second increment / decrease amount is above a second threshold that is different from the first threshold, the acquired sound signal is determined to satisfy the second condition. If the acquired sound signal is determined to satisfy the second condition, the reproduction step is performed.
[0086] Therefore, the sound signal representing an indirect tone is subjected to a first determination process and a second determination process, and an output signal based on the sound signal obtained under the condition that the sound signal satisfies the first and second conditions is output. That is, it is appropriately determined whether to output an output signal based on the sound signal representing the indirect tone. When no output signal is output, the computational complexity and computational load are reduced. That is, a sound signal processing method that can appropriately reduce the computational complexity and computational load can be implemented.
[0087] In the ninth aspect of the sound signal processing method of this disclosure, in the eighth aspect of the sound signal processing method, the second increment / decrease calculated in the second gain accumulation step includes the increment / decrease when diffusion filtering is performed, which improves the realism of the indirect tone by diffusion. The increment / decrease when diffusion filtering is performed is the increment / decrease that amplifies the amplitude of the acquired sound signal. The value is the ratio of the volume of the direct tone involved in the indirect tone to the second increment / decrease calculated in the second gain accumulation step. The second threshold is determined based on the arrival time difference between the direct tone and the indirect tone.
[0088] Therefore, in the second determination process, when the aforementioned ratio is greater than or equal to the second threshold, it is determined that the sound signal representing the indirect tone satisfies the second condition, and an output signal based on the sound signal obtained under the condition of satisfying such the second condition is output. That is, it more appropriately determines whether to output an output signal based on the sound signal representing the indirect tone. In other words, it is possible to realize a sound signal processing method that can more appropriately reduce the computational load and computational complexity.
[0089] The 10th method of sound signal processing disclosed herein is the 8th method of sound signal processing, wherein the second increment / decrease calculated in the second gain accumulation step includes an increment / decrease of the amplitude based on the sound quality adjustment function set by the listener's selection, the value being the ratio of the volume of the direct tone involved in the indirect tone to the second increment / decrease calculated in the second gain accumulation step, and the second threshold is determined based on the arrival time difference between the direct tone and the indirect tone.
[0090] Thus, listeners can listen to sounds with their preferred sound quality by making a selection that aligns with their own preferences.
[0091] The 11th method of this disclosure is a sound signal processing method of any one of the 8th to 10th methods, wherein the first threshold is a value related to volume.
[0092] Therefore, it is possible to implement a sound signal processing method that uses a volume-related value as the first threshold.
[0093] The audio signal processing method of the 12th aspect of this disclosure is an audio signal processing method of any one of the 8th to 11th aspects, wherein, in the second gain accumulation step, a parameter representing the increase or decrease amount determined by the one or more second processes is obtained, the increase or decrease amount determined by the one or more second processes is the increase or decrease amount that amplifies the amplitude of the obtained audio signal, and the second increase or decrease amount is calculated based on the multiple parameters obtained.
[0094] Therefore, a second increment or decrement corresponding to the above parameters is calculated, thereby enabling a more appropriate determination of whether to output an output signal based on an indirect tone. In other words, a sound signal processing method that can more appropriately reduce computational complexity and workload can be implemented.
[0095] The 13th method of sound signal processing disclosed herein is a sound signal processing method of any one of the 8th to 12th methods, comprising: a modification step of setting the first threshold; and an invalidation step of performing invalidation processing to invalidate the second determination processing. In the first determination processing, if the calculated first increment or decrement is greater than or equal to the set first threshold, it is determined that the acquired sound signal satisfies the first condition. If the invalidation processing has been performed, and it is determined that the acquired sound signal satisfies the first condition, the reproduction step is performed.
[0096] Therefore, by eliminating the second decision-making process through invalidation, the computational load and computational complexity required for the second decision-making process can be reduced. In other words, a sound signal processing method that can more appropriately reduce computational load and computational complexity can be implemented.
[0097] The sound signal processing method of the 14th aspect of this disclosure, in the sound signal processing method of the 13th aspect, has an area where the storage region representing the value of the first threshold set in the change step is adjacent to the area where the storage region storing the signal indicating the implementation of the invalidation process is located.
[0098] Therefore, not only can the virtual space administrator or listener easily visually grasp the set values, but also, by having adjacent memory regions, a special effect is created where the values of both can be set simultaneously (through a single memory access process). In other words, by configuring the value representing the aforementioned first threshold and the signal indicating the implementation of the aforementioned invalidation process in the high-order bit field and low-order bit field of a region accessible through a single address, writing this configured series of data through a single memory access process allows both data to be set simultaneously.
[0099] In the audio signal processing method of the 15th aspect of this disclosure, in the audio signal processing method of the 13th aspect, the first threshold set when the invalidation processing is performed is greater than the first threshold set when the invalidation processing is not performed.
[0100] Therefore, a sound signal processing method is available that can set the size of the first threshold based on whether invalidation processing is performed.
[0101] The computer program of the 16th aspect of this disclosure is used to cause a computer to execute the sound signal processing method of any one of the 1st to 15th aspects.
[0102] Therefore, the computer can execute the above-mentioned sound signal processing method according to the computer program.
[0103] The sound signal processing apparatus of the 17th aspect of this disclosure includes: an acquisition unit that acquires a sound signal, the sound signal including attribute information that determines the attributes of the sound signal; a determination unit that, when the attribute determined by the attribute information included in the acquired sound signal is information representing an indirect tone, performs a first determination process to determine whether the acquired sound signal satisfies a first condition and a second determination process to determine whether the acquired sound signal satisfies a second condition different from the first condition; and a reproduction unit that, when the acquired sound signal satisfies the first condition and the second condition, outputs an output signal based on the acquired sound signal.
[0104] Therefore, the first and second determination processes are performed on the sound signal with the attribute of indirect tone, and an output signal based on the sound signal obtained under the condition that the sound signal satisfies the first and second conditions is output. That is, it is appropriately determined whether to output an output signal based on the sound signal with the attribute of indirect tone. When no output signal is output, the computational load and computational complexity are reduced. That is, a sound signal processing apparatus that can appropriately reduce the computational load and computational complexity can be realized.
[0105] (Implementation Method 1) (Example of a stereo sound reproduction system) Figure 2 This diagram illustrates an example of a stereo sound reproduction system 1000. Specifically, Figure 2 This describes a stereo sound reproduction system 1000 as an example of a system capable of applying the sound processing or decoding techniques disclosed herein. Stereo sound is also referred to as immersive audio. The stereo sound reproduction system 1000 includes a sound signal processing device 1001 and a sound prompting device 1002.
[0106] The sound signal processing device 1001 is also manifested as an audio processing device, which performs audio processing on the sound signal emitted by the virtual sound source to generate an audio-processed sound signal for the listener. The sound signal is not limited to speech; any audible sound is acceptable. Audio processing, for example, is signal processing performed on the sound signal to reproduce the effects that the sound undergoes from the sound source to the listener.
[0107] The sound signal processing device 1001 performs sound processing based on spatial information describing the reasons for the aforementioned effects. Spatial information includes, for example, information indicating the location of the sound source, the listener, and surrounding objects; information indicating the shape of the space; and parameters related to sound propagation. The sound signal processing device 1001 may be, for example, a PC (Personal Computer), a smartphone, a tablet computer, or a game console.
[0108] The processed audio signal is presented to the listener by the audio prompt device 1002. The audio prompt device 1002 is connected to the audio signal processing device 1001 via wireless or wired communication. The processed audio signal generated by the audio signal processing device 1001 is transmitted to the audio prompt device 1002 via wireless or wired communication.
[0109] When the sound prompting device 1002 is composed of multiple devices, such as a device for the right ear and a device for the left ear, the multiple devices simultaneously provide prompts through communication between the multiple devices or through communication between each of the multiple devices and the sound signal processing device 1001. The sound prompting device 1002 may be, for example, a headset, earplugs, head-mounted display, or a surround sound system composed of multiple fixed speakers, worn on the listener's head.
[0110] Furthermore, the stereo sound reproduction system 1000 can also be used in combination with image prompting devices or stereoscopic image prompting devices that visually provide ER experiences, including AR / VR. For example, the space processed by spatial information is a virtual space, where the positions of sound sources, listeners, and objects are virtual positions of virtual sound sources, virtual listeners, and virtual objects within the virtual space. This space can also be represented as a sound space. Additionally, spatial information can also be represented as sound spatial information.
[0111] also, Figure 2 This is an example of a system configuration where the sound signal processing device 1001 and the sound prompting device 1002 are different devices, but the stereo sound reproduction system 1000, which can apply the sound processing method (sound signal processing method) or decoding method of this disclosure, is not limited to this. Figure 2The configuration can be as follows. For example, the sound signal processing device 1001 can be included in the sound prompting device 1002, and the sound prompting device 1002 performs both sound processing and sound prompting.
[0112] Alternatively, the sound signal processing device 1001 and the sound prompting device 1002 may share the implementation of the sound processing described in this disclosure. Furthermore, a portion or all of the sound processing described in this disclosure may be implemented via a server connected to the sound signal processing device 1001 or the sound prompting device 1002 through a network.
[0113] Furthermore, the audio signal processing apparatus 1001 can also perform audio processing by decoding a bitstream generated by encoding at least a portion of the audio signal and spatial information data used for audio processing. Therefore, the audio signal processing apparatus 1001 can also be exemplified as a decoding device.
[0114] (Example of an encoding device) Figure 3A This is a block diagram illustrating an example of the configuration of the encoding device 1100. Specifically, Figure 3A The configuration of encoding device 1100, which is an example of the encoding device of this disclosure, is shown.
[0115] Input data 1101 is encoded object data containing spatial information and / or audio signals input to encoder 1102. Details regarding the spatial information will be explained later.
[0116] Encoder 1102 encodes the input data 1101 to generate encoded data 1103. Encoded data 1103 is, for example, a bitstream generated through encoding processing.
[0117] The memory 1104 stores the encoded data 1103. The memory 1104 may be, for example, a hard disk or an SSD (Solid-State Drive), or other types of memory.
[0118] Furthermore, in the above description, the bitstream generated through encoding processing was listed as an example of encoded data 1103 stored in memory 1104, but the encoded data 1103 can also be data other than a bitstream. For example, the encoding device 1100 may also store transformed data generated by converting the bitstream into a specified data format in memory 1104. The transformed data may, for example, be a file or multiplexed stream corresponding to more than one bitstream.
[0119] Here, the file is a file with a file format such as ISOBMFF (ISO Base Media File Format). Furthermore, the encoded data 1103 can also be in the form of multiple packets generated by splitting the aforementioned bitstream or file.
[0120] For example, the bitstream generated by encoder 1102 can be transformed into data different from the bitstream. In this case, encoding device 1100 has a transformation unit (not shown), which can perform the transformation processing either by the transformation unit or by a CPU (Central Processing Unit), which is an example of a processor described later.
[0121] (Example of a decoding device) Figure 3B This is a block diagram illustrating an example of the configuration of the decoding device 1110. Specifically, Figure 3B The configuration of decoding device 1110, which is an example of the decoding device of this disclosure, is shown.
[0122] Memory 1114 stores, for example, the same data as the encoded data 1103 generated by encoding device 1100. The stored data is read from memory 1114 and input as input data 1113 into decoder 1112. Input data 1113 is, for example, a bitstream intended for decoding. Memory 1114 can be, for example, a hard disk or SSD, or other type of storage.
[0123] Alternatively, the decoding device 1110 may not input the data read from the memory 1114 as input data 1113 into the decoder 1112 as is, but instead transform the read data and input the transformed data as input data 1113 into the decoder 1112. The data before transformation may be, for example, multiplexed data containing more than one bitstream. Here, the multiplexed data may also be a file with a file format such as ISOBMFF.
[0124] Furthermore, the data before conversion can also be multiple packets generated by splitting the aforementioned bitstream or file. Alternatively, data different from the bitstream can be read from memory 1114 and converted into a bitstream. In this case, the decoding device 1110 may also include a conversion unit (not shown), and the conversion processing may be performed by the conversion unit, or by a CPU, as an example of a processor described later.
[0125] Decoder 1112 decodes the input data 1113 and generates an audio signal 1111 representing a prompt to the listener.
[0126] (Another example of an encoding device) Figure 3C This is a block diagram illustrating another configuration example of an encoding device. Specifically, Figure 3C This illustrates the configuration of encoding device 1120, which is another example of the encoding device disclosed herein. Figure 3C In China, for the sake of Figure 3A The same constituent elements are assigned to the same constituent elements as Figure 3A The same labels are used for these constituent elements, and the descriptions are omitted for these constituent elements.
[0127] Encoding device 1100 stores encoded data 1103 in memory 1104. On the other hand, encoding device 1120 differs from encoding device 1100 in that it has a transmitting unit 1121 that transmits encoded data 1103 to the outside.
[0128] The transmitting unit 1121 transmits a transmission signal 1122 generated based on encoded data 1103 or data transformed from encoded data 1103 into other data formats to other devices or servers. The data used in generating the transmission signal 1122 may be, for example, a bit stream, multiplexed data, file, or packet as described in the encoding device 1100.
[0129] (Another example of a decoding device) Figure 3D This is a block diagram illustrating another configuration example of a decoding device. Specifically, Figure 3D This illustrates the configuration of decoding device 1130, another example of a decoding device disclosed herein. Figure 3D In China, for the sake of Figure 3B The same constituent elements are assigned to the same constituent elements as Figure 3B The same labels are used for these constituent elements, and the descriptions are omitted for these constituent elements.
[0130] Decoding device 1110 reads input data 1113 from memory 1114. On the other hand, decoding device 1130 differs from decoding device 1110 in that it has a receiving unit 1131 that receives input data 1113 from the outside.
[0131] The receiving unit 1131 receives the received signal 1132 to obtain received data, and outputs the input data 1113 input to the decoder 1112. The received data can be the same as the input data 1113 input to the decoder 1112, or it can be data in a different format than the input data 1113.
[0132] If the format of the received data differs from the format of the input data 1113, the receiving unit 1131 may convert the received data into the input data 1113. Alternatively, the receiving data may be converted into the input data 1113 by a conversion unit (not shown) or CPU of the decoding device 1130. The received data may be, for example, a bit stream, multiplexed data, a file, or a packet as described in the encoding device 1120.
[0133] (Example of a decoder) Figure 4A This is a block diagram illustrating an example of the configuration of decoder 1200. Specifically, Figure 4A Indicates as Figure 3B or Figure 3D The decoder 1200 is an example of the decoder 1112 in the example.
[0134] Input data 1113 is the encoded bitstream, which contains encoded audio data as the encoded audio signal and metadata used in audio processing.
[0135] The Spatial Information Management Unit 1201 acquires the metadata contained in the input data 1113 and parses the metadata. The metadata contains information describing the elements that act on sound and are configured in the sound space. The Spatial Information Management Unit 1201 manages the spatial information used in sound processing obtained by parsing the metadata and provides the spatial information to the Rendering Unit 1203.
[0136] Furthermore, in this disclosure, the information used in sound processing is represented as spatial information, but other representations may also be used. For example, the information used in sound processing may be represented as sound spatial information or scene information. In addition, when the information used in sound processing changes over time, the spatial information input to the rendering unit 1203 may also be represented as information such as spatial state, sound spatial state, or scene state.
[0137] Furthermore, spatial information can be managed on a per-sound-space or per-scene basis. For example, when multiple distinct rooms are represented as virtual spaces, these rooms can be managed as separate scenes. Additionally, even within the same space, spatial information can be managed as different scenes depending on the presented conditions.
[0138] Therefore, multiple spatial information can be managed for multiple sound spaces or multiple scenes. In the management of multiple spatial information, each spatial information can be assigned an identifier to distinguish between the multiple spatial information.
[0139] Spatial information data can also be included in the bitstream, which is one example of input data 1113. Alternatively, the bitstream may contain identifiers for spatial information, and the spatial information data may be obtained from an information source outside the bitstream. Specifically, when the bitstream only contains identifiers for spatial information, the identifiers can be used during rendering to obtain spatial information data stored in the device's memory or an external server as input data 1113.
[0140] Furthermore, the information managed by the Spatial Information Management Department 1201 is not limited to the information contained in the bitstream. For example, the input data 1113 may include data on the characteristics and structure of the representation space obtained from software or servers providing VR or AR, as data not included in the bitstream.
[0141] Furthermore, the input data 1113 may also include data representing the characteristics and location of the listener or object. Additionally, the input data 1113 may include information about the listener's location obtained by sensors possessed by the terminal, including the decoding devices (1110, 1130), and may also include information representing the terminal's location inferred based on the information obtained from the sensors.
[0142] That is, the spatial information management unit 1201 can also communicate with external systems or servers to obtain spatial information and the listener's location (i.e., listening location). The spatial information management unit 1201 can obtain clock synchronization information from external systems and perform clock synchronization processing with the rendering unit 1203.
[0143] Furthermore, the space described above can be a virtually formed space, i.e., VR space, or a real space or a virtual space corresponding to a real space, i.e., AR space or MR space. Additionally, the virtual space can also be represented as a sound field or acoustic space. Furthermore, the information indicating position described above can be information such as coordinate values representing the position within the space, information representing the relative position with respect to a defined reference position, or information representing the motion or acceleration of the position within the space.
[0144] The audio data decoder 1202 decodes the encoded audio data contained in the input data 1113 to obtain the audio signal.
[0145] The encoded audio data acquired by the stereo sound reproduction system 1000 is, for example, a bitstream encoded in a format specified by MPEG-H 3D Audio (ISO / IEC 23008-3). Furthermore, MPEG-H 3D Audio is merely one example of an encoding method that can be used when generating encoded audio data contained within a bitstream. The encoded audio data can also be a bitstream encoded in other encoding methods.
[0146] For example, the encoding method can also be an irreversible codec such as MP3 (MPEG-1 Audio Layer-3), AAC (Advanced Audio Coding), WMA (Windows Media Audio), AC3 (Audio Codec-3), or Vorbis. Alternatively, the encoding method can be a reversible codec such as ALAC (Apple Lossless Audio Codec) or FLAC (Free Lossless Audio Codec).
[0147] Alternatively, any encoding method other than those described above can be used. For example, PCM (pulse code modulation) data can be used as encoded audio data. In this case, for example, if the number of quantization bits of the PCM data is N, the decoding process can be a process of converting the N-bit binary number into a number form (e.g., floating-point form) that the rendering unit 1203 can process.
[0148] The rendering unit 1203 acquires the sound signal and spatial information, uses the spatial information to perform sound processing on the sound signal, and outputs the sound processed sound signal (sound signal 1111).
[0149] Before rendering begins, the Spatial Information Management Unit 1201 reads the metadata of the input signal, detects the rendering items such as objects and sounds defined by the spatial information, and sends them to the Rendering Unit 1203. After rendering begins, the Spatial Information Management Unit 1201 monitors the changes in spatial information and the listener's location over time, updates and manages the spatial information, and sends the updated spatial information to the Rendering Unit 1203.
[0150] The rendering unit 1203 generates and outputs a sound signal with added audio processing based on the sound signal contained in the input data 1113 and the spatial information received from the spatial information management unit 1201.
[0151] Spatial information update processing and sound signal output processing with added audio processing can also be performed by the same thread. Spatial information management unit 1201 and rendering unit 1203 can assign processing to their respective independent threads. When spatial information update processing and sound signal output processing with added audio processing are performed in different threads, the start frequency of each thread can be set separately, or the processing can be performed in parallel.
[0152] When the spatial information management unit 1201 and the rendering unit 1203 perform processing in different independent threads, computing resources can be allocated preferentially to the rendering unit 1203. As a result, sound output processing that does not allow for even small delays, such as generating noise like a popping sound with a delay of 1 sample (0.02 msec), can be performed safely.
[0153] At this time, the allocation of computing resources to the spatial information management unit 1201 is restricted. However, compared to the output processing of audio signals, the updating of spatial information is a low-frequency process (e.g., updating the listener's facial orientation), and therefore does not need to be instantaneous like the output processing of audio signals. Therefore, even if the allocation of computing resources is restricted, it will not have a significant impact on the sound quality.
[0154] Spatial information updates can be performed periodically at preset times or intervals, or they can be performed when preset conditions are met. Furthermore, spatial information updates can be performed manually by the listener or the administrator of the sound space, or they can be triggered by changes in external systems.
[0155] For example, the listener can operate the controller to update the spatial information when their avatar's standing position is instantly distorted, or when it moves forward or backward. Alternatively, the spatial information can be updated when the virtual space administrator implements a performance that suddenly changes the environment of the venue. In these cases, the thread used to update the spatial information managed by the Spatial Information Management Department 1201 can be started either periodically or as a single interrupt.
[0156] The information update thread, which performs spatial information update processing, handles tasks such as updating the position or orientation of the listener's avatar within the virtual space based on the position or orientation of the VR goggles worn by the listener, and updating the position of moving objects within the virtual space. This is provided within a relatively low-frequency processing thread, typically around tens of Hz. Processing reflecting the properties of direct tones can also be performed within this low-frequency processing thread. This is because the frequency of changes in the properties of direct tones is lower than the frequency of audio processing frames used for audio output. This approach reduces the computational load on the processing and avoids the risk of impulsive noise that would result from updating information at unnecessarily high frequencies.
[0157] Figure 4B This is a block diagram representing another example of a decoder. Specifically, Figure 4B Indicates as Figure 3B or Figure 3D Another example of decoder 1112 is the construction of decoder 1210.
[0158] Figure 4B The point is that input data 1113 does not contain encoded audio data but rather unencoded audio signals. Figure 4A Different. Input data 1113 includes a bitstream containing metadata and an audio signal.
[0159] Spatial Information Management Department 1211 due to its relationship with Figure 4A The same applies to the Spatial Information Management Department 1201, so the explanation is omitted.
[0160] Rendering Department 1213 due to its relationship with Figure 4A The rendering unit 1203 is the same, so the description is omitted.
[0161] Alternatively, decoders 1112, 1200, and 1210 can also be represented as audio processing units that perform audio processing. Furthermore, decoders 1110 and 1130 can also be audio signal processing devices 1001, or can be represented as audio processing devices.
[0162] (Physical structure of a sound signal processing device) Figure 5 This diagram illustrates an example of the physical configuration of the sound signal processing device 1001. Additionally, Figure 5 The sound signal processing device 1001 can also be Figure 3B Decoding device 1110 or Figure 3D Decoding device 1130. Figure 3B or Figure 3D The multiple constituent elements shown can also be obtained through Figure 5 The various components shown are installed. Furthermore, a portion of the components described herein can also be incorporated into the sound prompting device 1002.
[0163] Figure 5 The sound signal processing device 1001 includes a processor 1402, a memory 1404, a communication interface 1403, a sensor 1405, and a speaker 1401.
[0164] The processor 1402 is, for example, a CPU, a DSP (Digital Signal Processor), or a GPU (Graphics Processing Unit). The audio processing or decoding processing of this disclosure can also be implemented by the CPU, DSP, or GPU executing a program stored in memory 1404. Furthermore, the processor 1402 is, for example, a circuit that performs information processing. The processor 1402 can also be a dedicated circuit that performs signal processing of sound signals, including the audio processing of this disclosure.
[0165] The memory 1404 may be composed of, for example, RAM (Random Access Memory) or ROM (Read Only Memory). The memory 1404 may also include magnetic recording media such as a hard disk or semiconductor memory such as an SSD. Furthermore, the memory 1404 may be internal memory built into the CPU or GPU. In addition, the memory 1404 may store spatial information managed by the spatial information management units 1201 and 1211. Furthermore, it may also store threshold data, which will be described later.
[0166] The communication IF1403 is, for example, a communication module corresponding to communication methods such as Bluetooth (registered trademark) or WIGIG (registered trademark). The audio signal processing device 1001 communicates with other communication devices, for example, via the communication IF1403, to obtain the bitstream of the decoded object. The obtained bitstream is stored, for example, in the memory 1404.
[0167] The communication IF1403, for example, consists of signal processing circuitry and an antenna corresponding to the communication method. The communication method is not limited to Bluetooth (registered trademark) and WIGIG (registered trademark), but can also be LTE (Long Term Evolution), NR (New Radio), or Wi-Fi (registered trademark), etc.
[0168] Furthermore, the communication method is not limited to the wireless communication methods mentioned above. The communication method can also be a wired communication method such as Ethernet (registered trademark), USB (Universal Serial Bus), or HDMI (registered trademark) (High-Definition Multimedia Interface).
[0169] Sensor 1405 performs sensing to infer the listener's position and orientation. Specifically, sensor 1405 infers the listener's position and / or orientation based on the detection results of one or more of the following: position, orientation, movement, velocity, angular velocity, and acceleration of a part or part of the body or the whole body, and generates position / or orientation information representing the listener's position and / or orientation.
[0170] Alternatively, the sensor 1405 may be an external device of the sound signal processing device 1001. A part of the body may also be the listener's head, etc. The position / or orientation information may be information indicating the listener's position and / or orientation in real space, or it may be information indicating the displacement of the listener's position and / or orientation based on a predetermined point in time. Furthermore, the position / or orientation information may also be information indicating the relative position and / or orientation to the stereo sound reproduction system 1000 or the external device equipped with the sensor 1405.
[0171] Sensor 1405 can be, for example, a camera or a ranging device such as LiDAR (Light Detection and Ranging). Sensor 1405 can also capture images of the listener's head movements and detect these movements by processing the captured images. Furthermore, a device that uses wireless communication in any frequency band, such as millimeter waves, to perform position estimation can also be used as sensor 1405.
[0172] Furthermore, the sound signal processing device 1001 can also obtain location information from an external device equipped with sensor 1405 via communication IF 1403. In this case, the sound signal processing device 1001 may also not include sensor 1405. Here, the external device is, for example, a... Figure 2 The sound prompting device 1002 described herein may be a stereoscopic image reproduction device worn on the head of the listener. In this case, the sensor 1405 is configured by combining various sensors such as a gyroscope sensor and an accelerometer sensor.
[0173] For example, as the speed of the listener's head movement, sensor 1405 can detect the angular velocity of rotation about at least one of three mutually orthogonal axes in the sound space, and can also detect the acceleration of displacement about at least one of the three axes.
[0174] For example, as a measure of head movement in the listener, sensor 1405 can detect rotational motion about at least one of three mutually orthogonal axes in the sound space, and displacement about at least one of these three axes. Specifically, sensor 1405 detects 6DoF position (x, y, z) and angle (yaw, pitch, roll) as the listener's position. Sensor 1405 is constructed by combining various sensors used for motion detection, such as gyroscopes and accelerometers.
[0175] Alternatively, sensor 1405 can be implemented using a camera or a GPS (Global Positioning System) receiver, which detects the listener's location. Location information obtained by inferring location using LiDAR or similar devices as sensor 1405 can also be used. For example, sensor 1405 can be built into a smartphone in the case where the stereo sound reproduction system 1000 is implemented in a smartphone.
[0176] Furthermore, sensor 1405 may also include a temperature sensor such as a thermocouple that detects the temperature of the sound signal processing device 1001. Additionally, sensor 1405 may also include a battery included in the sound signal processing device 1001, or a sensor that detects the remaining amount of the battery connected to the sound signal processing device 1001.
[0177] The loudspeaker 1401 includes, for example, a drive mechanism and an amplifier, such as a diaphragm, a magnet, or a voice coil, which transmits the processed sound signal as a sound prompt to the listener. The loudspeaker 1401 activates the drive mechanism based on the sound signal amplified by the amplifier (more specifically, a waveform signal representing the waveform of the sound), causing the diaphragm to vibrate. The diaphragm, vibrating in response to the sound signal, generates sound waves, which propagate through the air and reach the listener's ear, allowing the listener to perceive the sound.
[0178] Additionally, an example is given here of a sound signal processing device 1001 that includes a speaker 1401 and prompts the sound signal after sound processing via the speaker 1401, but the sound signal prompting mechanism is not limited to the above configuration.
[0179] For example, an audio-processed sound signal can be output to an external sound prompting device 1002 connected via a communication module. Communication via the communication module can be either wired or wireless. Furthermore, as another example, the sound signal processing device 1001 has a terminal for outputting an analog sound signal, to which a cable for an earphone or similar device is connected, and a prompting sound signal is emitted from the earphone or similar device.
[0180] In the above-described cases, the sound prompting device 1002 may also be a headset, earphone, head-mounted display, neck speaker, or wearable speaker worn on the head or part of the listener's body. Alternatively, the sound prompting device 1002 may also be a surround sound system composed of multiple fixed speakers. Furthermore, the sound prompting device 1002 can also reproduce sound signals.
[0181] (Physical structure of the encoding device) Figure 6 This is a diagram illustrating an example of the physical configuration of the encoding device 1500. Figure 6 The encoding device 1500 can also be Figure 3A Encoding device 1100 or Figure 3C The encoding device 1120 can also Figure 3A or Figure 3C The multiple constituent elements shown are composed of Figure 6 The installation of the multiple components shown.
[0182] Figure 6 The encoding device 1500 includes a processor 1501, a memory 1503, and a communication IF 1502.
[0183] The processor 1501 is, for example, a CPU, DSP, or GPU. The encoding processing of this disclosure can also be implemented by the CPU, DSP, or GPU executing a program stored in memory 1503. Furthermore, the processor 1501 is, for example, a circuit that performs information processing. The processor 1501 can also be a dedicated circuit that performs signal processing on an audio signal, including the encoding processing of this disclosure.
[0184] The memory 1503 may be composed of, for example, RAM or ROM. The memory 1503 may also include magnetic recording media, such as a hard disk, or semiconductor memory, such as an SSD. Furthermore, the memory 1503 may also be internal memory embedded in a CPU or GPU.
[0185] The communication IF1502 is, for example, a communication module corresponding to communication methods such as Bluetooth (registered trademark) or WIGIG (registered trademark). The encoding device 1500 communicates with other communication devices, for example, via the communication IF1502, and transmits the encoded bit stream.
[0186] The communication IF1502, for example, consists of signal processing circuitry and an antenna corresponding to the communication method. The communication method is not limited to Bluetooth (registered trademark) and WIGIG (registered trademark), but can also be LTE, NR, or Wi-Fi (registered trademark), etc. Furthermore, the communication method is not limited to wireless communication. The communication method can also be wired communication methods such as Ethernet (registered trademark), USB, or HDMI (registered trademark).
[0187] The communication module, for example, consists of a signal processing circuit and an antenna corresponding to the communication method. In the examples above, Bluetooth (registered trademark) or WIGIG (registered trademark) are cited as communication methods, but it can also correspond to communication methods such as LTE (Long Term Evolution), NR (New Radio), or Wi-Fi (registered trademark). Furthermore, the communication IF may not be a wireless communication method as described above, but rather a wired communication method such as Ethernet (registered trademark), USB (Universal Serial Bus), or HDMI (registered trademark) (High-Definition Multimedia Interface).
[0188] [The Composition of the Rendering Department] Figure 7 This is a block diagram illustrating an example of the configuration of the rendering unit 1300. Specifically, Figure 7 Indicates and Figure 4A and Figure 4B An example of the detailed configuration of the rendering unit 1300 corresponding to the rendering units 1203 and 1213.
[0189] The rendering unit 1300 consists of a resolution unit 1301, a determination unit 1302, and a reproduction unit 1303. It performs additional audio processing on the audio data contained in the input signal and outputs it.
[0190] The input signal may consist of spatial information, sensor information, and sound data. The input signal may also contain a bitstream of sound data and metadata (control information), in which case spatial information may also be included in the metadata.
[0191] Spatial information is information related to the sound space (three-dimensional sound field) formed by the stereo sound reproduction system 1000. It consists of information related to the objects contained in the sound space and information related to the listener. Among the objects, there are sound source objects that emit sound and become sound sources, and non-sound-emitting objects that do not emit sound. Sound source objects can also be simply represented as sound sources.
[0192] Non-sound-producing objects can act as obstacles reflecting the sound emitted by a sound source object, but there are also cases where a sound source object acts as an obstacle reflecting the sound emitted by other sound source objects. Obstacle objects can also be represented as reflecting objects.
[0193] As information shared by both the sound source object and the non-sound-producing object, it includes location information, shape information, and the attenuation rate of the volume when the object reflects the sound.
[0194] Position information is represented by coordinate values along three axes in Euclidean space, such as the X, Y, and Z axes, but it doesn't necessarily have to be three-dimensional. For example, position information can also be two-dimensional, represented by coordinate values along only the X and Y axes. The position information of an object is determined by the representative position of its shape, represented by a mesh or voxels.
[0195] Shape information can also include information related to the material of the surface.
[0196] The attenuation rate can be represented by a real number above 0 and below 1, or by a negative decibel value. Since volume is not amplified by reflection in real space, the attenuation rate is set to a negative decibel value. However, for example, to present the sense of horror in an unreal space, an attenuation rate above 1, i.e., a positive decibel value, can be deliberately set.
[0197] Furthermore, the attenuation rate can be set to a different value for each of the multiple frequency bands, or it can be set independently for each frequency band. Additionally, when setting the attenuation rate for each type of material on the object's surface, the corresponding attenuation rate value can be used based on information related to the surface material.
[0198] Furthermore, spatial information can also include information indicating whether an object is a living organism or whether it is a moving object. If the object is a moving object, the position indicated by the location information can also change over time. In this case, information about the changed position or the amount of change is transmitted to the rendering unit 1300.
[0199] Information related to the sound source object includes not only the information shared by the sound source object and the non-sound-producing object, but also the sound data and the information required to project the sound data into the sound space. Sound data represents information related to the frequency and intensity of the sound, and is data that reflects the sound perceived by the listener.
[0200] The audio data is typically a PCM signal, but it can also be data compressed using encoding methods such as MP3. In this case, the signal needs to be decoded at least before it reaches the reproduction unit 1303, so the rendering unit 1300 may also include a decoding unit (not shown). Alternatively, the signal can also be decoded by the audio data decoder 1202.
[0201] For a sound source object, one sound data point or multiple sound data points can be set. Additionally, recognition information can be assigned to each sound data point, and information related to the sound source object can also include this recognition information.
[0202] Information required to project sound data into the sound space may include, for example, reference volume information used as a reference in the reproduction of sound data, information representing the properties (also called characteristics) of the sound data, information related to the location of the sound source object, and information related to the orientation of the sound source object (i.e., information related to the directivity of the sound emitted by the sound source object).
[0203] Reference volume information can be, for example, the effective value of the amplitude of the sound data at the sound source location when the sound data is radiated into the sound space, or it can be represented as a decibel (dB) value in floating point.
[0204] For example, when the reference volume is 0dB, it can also mean that the volume of the signal level represented by the sound data is not increased or decreased, but the sound is radiated into the sound space at the original volume from the position indicated by the information related to the location of the sound source object. Alternatively, when the reference volume is -6dB, it can also mean that the volume of the signal level represented by the sound data is set to approximately half, and the sound is radiated into the sound space from the position indicated by the information related to the location of the sound source object.
[0205] The reference volume information can be assigned to each sound data point individually, or it can be assigned to multiple sound data points uniformly.
[0206] Information representing the properties of sound data can be, for example, information related to the volume of the sound source, and information representing the time-series variation of the sound source's volume.
[0207] For example, in a virtual conference room where the sound space is a speaker and the sound source is the speaker, the volume changes intermittently over a short period of time. That is, the spoken and silent parts alternate. Conversely, in a concert hall where the sound space is a concert hall and the sound source is a performer, the volume is maintained for a certain duration. Furthermore, in a battlefield where the sound space is an explosive device and the sound source is an explosive device, the volume of the explosion sound only increases momentarily, then remains silent or low.
[0208] In this way, the volume information of the sound source includes not only the magnitude of the sound, but also information about changes in the magnitude of the sound. This information can also be used to represent the properties of the sound data.
[0209] Transition information can also be represented by time-series data showing frequency characteristics. Transition information can also be represented by data showing the duration of the sound interval. Transition information can also be represented by time-series data showing the duration of both the sound and silent intervals. Transition information can also be represented by multiple sets of time-series data listing the duration during which the amplitude of the sound signal can be considered stationary (approximately constant) and the amplitude values of the signal during that period.
[0210] Transition information can also be represented by data that allows the frequency characteristics of a sound signal to be considered as a stationary duration. Transition information can also be represented by listing multiple sets of data, in time series, of the duration during which the frequency characteristics of the sound signal can be considered stationary and the frequency characteristics during that period. Transition information can also be represented, for example, by data showing the approximate shape of a spectrogram.
[0211] Alternatively, the volume used as the reference for the aforementioned frequency characteristics can also be the aforementioned reference volume. Information about the reference volume and information representing the properties of the sound data can be used for calculating the volume of direct or reflected sounds perceived by the listener, or for selecting whether or not to allow the listener to perceive them (also known as decision processing). Other examples and methods of utilizing information representing the properties of the sound data will be described later.
[0212] Furthermore, the reflected sound in this embodiment is an example of an indirect sound. An indirect sound can be a reflected sound, a diffracted sound, or the like. Additionally, the direct sound in this embodiment is an example of a defined sound, different from the aforementioned indirect sound. A defined sound can be a direct sound or a high-order ambisonic (HOA). Alternatively, representative sounds representing multiple sounds can be created by bundling multiple sound signals (performing mixing, etc.), and these representative sounds can be used as defined sounds. In this case, the representative sound can be called a representative sound.
[0213] In this embodiment, a reflected tone, as an example of an indirect tone, and a direct tone, as an example of a prescribed tone, are used for explanation. However, the same treatment is performed even if an indirect tone is used instead of a reflected tone and a prescribed tone is used instead of a direct tone.
[0214] Information related to the orientation of a sound source object (orientation information) is typically represented by yaw, pitch, and roll. Alternatively, the roll rotation can be omitted, and the orientation information of the sound source object can be represented by azimuth (yaw) and pitch (pitch). The orientation information of the sound source object can also change over time, and any changes are transmitted to the rendering unit 1300.
[0215] Information relevant to the listener relates to their position and orientation within the sound space. Position information is represented by the XYZ axes of Euclidean space, but it doesn't necessarily have to be three-dimensional; it can also be two-dimensional. Orientation information is typically represented by yaw, pitch, and roll. Alternatively, the roll can be omitted, and the listener's orientation information can be represented by azimuth (yaw) and pitch (pitch).
[0216] The listener's location and orientation information can also change over time, and if changes occur, they will be transmitted to the rendering unit 1300.
[0217] The sensor information includes information such as the amount of rotation or displacement detected by the sensor 1405 worn by the listener, as well as the listener's position and orientation. The sensor information is transmitted to the rendering unit 1300, which updates the listener's position and orientation information based on the sensor information. The sensor information may also include, for example, position information obtained by the portable terminal using GPS, a camera, or LiDAR for self-position estimation.
[0218] Alternatively, instead of sensor 1405, information obtained from an external source via a communication module can be used as sensor information for detection. Information indicating the temperature of the sound signal processing device 1001 and the remaining battery level can also be obtained from sensor 1405. Furthermore, the computing resources (CPU capacity, memory resources, or PC performance, etc.) of the sound signal processing device 1001 or the sound prompting device 1002 can be obtained in real time.
[0219] The analysis unit 1301 analyzes the sound signal contained in the input signal and the spatial information received from the spatial information management units 1201 and 1211, and calculates the information required for generating direct sound and reflected sound in the reproduction unit 1303, as well as the information required for determining (selecting) whether to generate reflected sound.
[0220] The information required for the generation of direct and reflected tones includes values related to the path taken to reach the listening position, the time taken to reach the position, and the volume upon arrival for each direct and reflected tone. These values represent, for example, the path taken to reach the listening position, the time taken to reach the position, and the volume upon arrival.
[0221] The information required for selecting the output reflected tone is information representing the relationship between the direct tone and the reflected tone, such as values related to the time difference between the direct tone and the reflected tone, and values related to the volume ratio of the direct tone and the reflected tone at the listening position. For example, the values related to the time difference between the direct tone and the reflected tone, and the values related to the volume ratio of the direct tone and the reflected tone at the listening position, are respectively values representing the time difference between the direct tone and the reflected tone, and values representing the volume ratio of the direct tone and the reflected tone at the listening position.
[0222] Furthermore, when volume is expressed in decibels on a logarithmic axis (in the case of representing volume in the decibel range), the volume ratio of two signals is naturally represented by the difference in decibel values. Specifically, the volume ratio of two signals can be the difference between the amplitude values of each signal expressed in the decibel range. This value can also be calculated based on energy values or power values, etc. Moreover, this difference can be referred to as the gain difference or simply the gain difference in the decibel range.
[0223] That is, the volume ratio in this disclosure is essentially the ratio of signal amplitudes, so it can also be expressed as Soundvolume ratio, Volume ratio, Amplitude ratio, Soundlevel ratio, Sound intensity ratio, or Gain ratio, etc. Furthermore, when the unit of volume is decibels, the volume ratio in this disclosure can of course also be referred to as volume difference.
[0224] In this disclosure, "volume ratio" typically refers to the gain difference between two sounds expressed in decibels. In examples of embodiments, the threshold data is also typically defined by the gain difference expressed in decibels. However, the volume ratio is not limited to the gain difference in decibels. When using a volume ratio expressed outside the decibel range, the threshold data defined in the decibel range can be converted to the units of the calculated volume ratio for use. Alternatively, the threshold data defined in each unit can be stored in memory in advance.
[0225] That is, even if a ratio such as energy value or power value is used instead of volume ratio, it is obvious that the algorithm in this disclosure can be applied to the solution of the problem of this disclosure.
[0226] The time difference between the arrival of the direct tone and the reflected tone is, for example, the time difference between the arrival time of the direct tone (arrival moment) and the arrival time of the reflected tone (arrival moment). Alternatively, for simplicity, the time difference between the arrival of the direct tone and the reflected tone is sometimes recorded as the time difference between the direct tone and the reflected tone. The time difference between the direct tone and the reflected tone can also be the time difference between the moments when the direct tone and the reflected tone arrive at the listening position, the difference in time required for the direct tone and the reflected tone to arrive at the listening position, or the time difference between the moment the direct tone ends and the moment the reflected tone arrives at the listening position. The methods for calculating these values will be described later.
[0227] The determination unit 1302 performs at least one of a first determination process using a first threshold and a second determination process using a second threshold. In this embodiment, the determination unit 1302 performs the second determination process on the reflected sound (i.e., the sound signal representing the reflected sound). The second determination process will be described in more detail below.
[0228] The determination unit 1302 uses the information calculated by the analysis unit 1301 and the threshold data representing the second threshold to determine whether the reproduction unit 1303 generates a reflected sound. In other words, the determination unit 1302 determines whether to select the reflected sound as the target reflected sound for generation. In other words, the determination unit 1302 selects which of the multiple reflected sounds generated by the reproduction unit 1303. Furthermore, hereinafter, the determination that the determination unit 1302 determines that the reproduction unit 1303 generates a reflected sound is sometimes described as the determination unit 1302 selecting a reflected sound, or the determination unit 1302 selecting to generate a reflected sound.
[0229] Threshold data, for example, is represented in a graph where the horizontal axis represents the time difference between the direct tone and the reflected tone, and the vertical axis represents the volume ratio of the direct tone to the reflected tone, as the boundary (threshold) between whether the reflected tone is perceived or not. Threshold data can also be represented by an approximation using the time difference between the direct tone and the reflected tone as a variable, or by an arrangement of values indexed by the time difference between the direct tone and the reflected tone, with corresponding thresholds. Furthermore, in Embodiment 1, sometimes the second threshold is simply described as a threshold.
[0230] For example, if the ratio of the volume of the direct sound to the volume of the reflected sound in the time difference between the arrival time of the direct sound and the arrival time of the reflected sound is greater than a threshold set based on reference threshold data, the determination unit 1302 selects to generate the reflected sound. Furthermore, the arrival volume refers to the volume of the sound when it reaches the listening position.
[0231] The time difference between the arrival time of the direct tone and the arrival time of the reflected tone is, in other words, the difference in time required for the direct tone and the reflected tone to reach the listening position respectively. Alternatively, the time difference between the end of the direct tone's emission and the arrival time of the reflected tone at the listening position can also be used as the time difference between the direct tone and the reflected tone. In this case, different threshold data can be used than the threshold data set based on the time difference between the arrival times of the direct tone and the reflected tone, or a common threshold data can be used.
[0232] The threshold data can be obtained from the memory 1404 of the audio signal processing device 1001, or from an external storage device via the communication module. The method for storing the threshold data and the method for setting the threshold will be described later.
[0233] The reproduction unit 1303 synthesizes the direct sound signal with the reflected sound signal selected and generated by the determination unit 1302.
[0234] Specifically, the reproduction unit 1303 processes the input sound signal to generate a direct tone based on the information about the arrival time and volume of the direct tone calculated by the analysis unit 1301. Furthermore, the reproduction unit 1303 processes the input sound signal to generate a reflected tone based on the information about the arrival time and volume of the reflected tone selected by the determination unit 1302. Then, the reproduction unit 1303 combines the generated direct tone and reflected tone and outputs them.
[0235] [Example of actions in the rendering department] Figure 8 This is a flowchart illustrating an example of the operation of the sound signal processing device 1001. Figure 8 The text indicates that the processing is mainly performed by the rendering unit 1300 of the sound signal processing device 1001.
[0236] In the analysis and processing of input signals ( Figure 8 In step S101, the analysis unit 1301 analyzes the input signal input to the sound signal processing device 1001 and detects direct tones and reflected tones that can be generated in the sound space. The reflected tones detected here are candidate reflected tones selected by the determination unit 1302 as the final reflected tones to be generated by the reproduction unit 1303. Furthermore, the analysis unit 1301 analyzes the input signal and calculates the information needed for generating direct tones and reflected tones, as well as the information needed for selecting the target reflected tones.
[0237] First, the characteristics of direct and reflected tones are calculated. Specifically, the arrival time and volume of the direct and reflected tones when they reach the listener are calculated. If multiple objects exist in the sound space as reflection targets, the characteristics of the reflected tones are calculated for each object separately.
[0238] The direct arrival time (td) is calculated based on the direct arrival path (pd). The direct arrival path (pd) is the path connecting the location information S(xs, ys, zs) of the sound source object with the location information A1(xa, ya, za) of the listener. The direct arrival time (td) is the value obtained by dividing the length of the path connecting the location information S(xs, ys, zs) and the location information A1(xa, ya, za) by the speed of sound (approximately 340 m / s).
[0239] For example, the path length (X) is calculated using (xs-xa)^2 + (ys-ya)^2 + (zs-za)^2)^0.5. Volume decreases inversely with distance. Therefore, given the volume as N in the location information S(xs, ys, zs) of the sound source object and the unit distance as U, the volume (ld) when the direct sound arrives is calculated using ld = N. Find U / X.
[0240] The volume N at the sound source location can also be the reference volume described earlier.
[0241] The arrival time (tr) of the reflected sound is calculated based on the arrival path (pr). The arrival path (pr) is the path that connects the position of the sound image of the reflected sound with the position information A1 (xa, ya, za).
[0242] Furthermore, the location of the sound image of reflected sound can be derived using methods such as the "mirror method" or the "ray tracing method," or any other method. The mirror method assumes that the reflected wave on the wall of a room has a mirror image at a position symmetrical to the sound source relative to the wall, and simulates the sound image by assuming that sound waves are emitted from that mirror image location. The ray tracing method simulates the image (sound image) observed at a certain point by tracing waves that propagate in straight lines, such as light rays or sound rays.
[0243] Figure 9 It is a diagram that shows the relative positional relationship between the listener and the obstacle object. Figure 10 This is a diagram showing the relatively close positional relationship between the listener and the obstacle object. That is, Figure 9 and Figure 10 These examples illustrate sound images formed at positions symmetrical to the sound source location, separated by a wall. By determining the position of the sound image of the reflected sound on the x, y, and z axes based on this relationship, the arrival time of the reflected sound can be calculated in the same way as the arrival time of the direct sound.
[0244] The arrival time (tr) of the reflected sound is obtained by dividing the length (Y) of the path connecting the position of the sound image of the reflected sound to the position information A1 (xa, ya, za) by the speed of sound (approximately 340 m / s). Volume decreases inversely with distance. Therefore, when the volume at the sound source is N, the unit distance is U, and the attenuation rate of the volume in the reflection is G, the volume (lr) when the reflected sound arrives decreases by lr = N. G Find U / Y.
[0245] As explained above, the attenuation rate G can be represented by a real number greater than or equal to 0 and less than 1, or by a negative decibel value. In this case, the overall volume attenuation of the signal corresponds to the amount of G. Furthermore, the attenuation rate can also be set for each of the multiple frequency bands. In this case, the analysis unit 1301 applies the specified attenuation rate for each frequency component of the signal. Additionally, to reduce computational complexity, the analysis unit 1301 can use representative values or average values of multiple attenuation rates from multiple frequency bands as the overall attenuation rate, thereby correspondingly attenuating the overall volume of the signal.
[0246] Next, the analysis unit 1301 calculates the volume ratio (L) required in the selection of the reflected sound of the generated object, namely the ratio of the volume (ld) when the direct sound arrives to the volume (lr) when the reflected sound arrives, and the time difference (T) between the direct sound and the reflected sound.
[0247] The ratio of the volume (ld) of the direct tone to the volume (lr) mentioned above, i.e., the volume ratio (L), is, for example, the value obtained by dividing the volume (lr) of the reflected tone by the volume (ld) of the direct tone, given by L = (N G U / Y) / (N) U / X) = G X / Y is calculated. Since the calculated value is a volume ratio, the values of N and U can be any pre-set values.
[0248] The time difference (T) between the direct tone and the reflected tone can also be the time difference between the direct tone and the reflected tone when they arrive at the listening position. For example, the time difference (T) between the direct tone and the reflected tone when they arrive at the listening position can be calculated by T = tr - td.
[0249] Furthermore, the time difference (T) can also be the difference between the arrival times of the direct tone and the reflected tone at the listening position. Alternatively, the time difference (T) can also be the time difference between the end of the direct tone's speech and the arrival time of the reflected tone at the listening position. In other words, the time difference (T) can also be the time difference between the end of the direct tone and the beginning of the reflected tone at the listening position.
[0250] Next, in the selective processing of reflected sounds ( Figure 8 In step S102, the determination unit 1302 selects whether the reproduction unit 1303 generates the reflected sound calculated by the analysis unit 1301. In other words, the determination unit 1302 determines whether to select the reflected sound as the target reflected sound for generation. When multiple reflected sounds exist, the determination unit 1302 selects whether to generate each reflected sound. The result of the determination unit 1302's selection of whether to generate each reflected sound can be either selecting more than one target reflected sound from multiple reflected sounds, or selecting none of the target reflected sounds.
[0251] Furthermore, the determination unit 1302 is not limited to generation processing; it can also select reflected sounds that are the application targets of other processing. For example, the determination unit 1302 can also select reflected sounds that are the application targets of binaural processing. In addition, the determination unit 1302 generally selects only one or more reflected sounds that are the processing targets. However, the determination unit 1302 can also select only one or more reflected sounds that are not the processing targets. Furthermore, processing can be applied to one or more reflected sounds that are not selected.
[0252] For example, the selection of reflected tones is based on the volume ratio (L) and time difference (T) calculated by the analysis unit 1301. By performing selection processing based on the time difference (T) between the direct tone and the reflected tone, it is possible to more appropriately select reflected tones that have a greater impact on the listener's perception compared to the case where selection processing is based solely on the volume difference between the direct tone and the reflected tone.
[0253] Specifically, the decision to generate a reflected tone is made, for example, by comparing the volume ratio of the direct tone to the reflected tone, corresponding to the time difference between the direct tone and the reflected tone, with a pre-set threshold. The threshold is set with reference to threshold data. The threshold data is an indicator representing the boundary at which the reflected tone of a direct tone is perceived by the listener, defined by the ratio of the volume (Id) when the direct tone arrives to the volume (lr) when the reflected tone arrives.
[0254] Furthermore, a threshold corresponds to a value expressed as a numerical value set in relation to a time difference (T). Threshold data corresponds to the relationship between the time difference (T) and the threshold, and to tabular data or formulas used to determine or calculate the threshold under the time difference (T). The form and type of threshold data are not limited to tabular data or formulas.
[0255] Figure 11 This is a graph showing the relationship between the time difference of the direct tone and the reflected tone and the threshold. For example, you can also refer to... Figure 11 The threshold data shown is for the volume ratio preset according to each value of the time difference between the direct tone and the reflected tone. Alternatively, one can refer to the data from... Figure 11 The threshold data shown is obtained through interpolation or extrapolation.
[0256] Furthermore, the threshold for the volume ratio under the time difference (T) calculated by the analysis unit 1301 is determined based on the threshold data. The determination unit 1302 then decides whether to select the reflected sound as the target reflected sound based on whether the volume ratio (L) of the direct sound to the reflected sound calculated by the analysis unit 1301 is higher than this threshold.
[0257] By using threshold data of volume ratios preset according to each value of the time difference between direct and reflected tones, selection processing that takes into account backmasking or priority effects can be achieved. Detailed explanations of the types, formats, storage methods, and setting methods of the threshold data will be provided later.
[0258] Next, in the generation and processing of direct and reflected tones ( Figure 8 In S103, the reproduction unit 1303 generates a direct sound signal and a reflected sound signal selected by the determination unit 1302 as the target reflected sound, and synthesizes them.
[0259] The direct tone sound signal is generated by applying the direct tone arrival time (td) and the direct tone arrival volume (ld) calculated by the analysis unit 1301 to the sound data of the sound source object contained in the input information. Specifically, the sound data is delayed by the direct tone arrival time (td) and multiplied by the direct tone arrival volume (ld). The sound data delay process is a process of shifting the position of the sound data back and forth on the time axis. For example, the sound data delay process disclosed in Patent Document 2 without degrading the sound quality can also be applied.
[0260] The sound signal of the reflected sound is generated in the same way as the direct sound by applying the sound data of the sound source object to the arrival time (tr) of the reflected sound and the volume (lr) of the reflected sound when it arrives, which is calculated by the analysis unit 1301.
[0261] However, the volume (lr) of the reflected sound during its arrival differs from the volume of the direct sound. It is affected by the attenuation rate G applied to the volume of the reflected sound. G can be an attenuation rate applied across the entire frequency band. Alternatively, it can be a reflection rate specified for each defined frequency band to reflect the bias in the frequency components generated by reflection. In this case, the application of the volume (lr) of the reflected sound can also be implemented as a frequency equalizer process, multiplying the volume by the attenuation rate for each frequency band.
[0262] In the example above, the path lengths of the direct and reflected tone candidates as they reach the listener are calculated. Then, the arrival time and volume are calculated based on each path length. Finally, the reflected tone candidates are selected based on their time difference and volume ratio.
[0263] Alternatively, as another example, selection can be performed based on the path lengths of the direct and reflected sounds as they reach the listener, omitting the calculations of arrival time and volume, as well as time difference and volume ratio. In this case, a threshold corresponding to the path length difference can be pre-set for the path length ratio. Furthermore, selection can be performed based on whether the calculated path length ratio is above the threshold corresponding to the calculated path length difference. Thus, selection can be performed based on the path length difference corresponding to the time difference while reducing computational load.
[0264] In addition to the path length difference, the value of a parameter representing the speed of sound propagation, or the value of a parameter that affects the speed of sound propagation, can also be used.
[0265] (Select the detailed processing option) The details of the selection process for whether or not to generate reflected sound are explained.
[0266] The selection of the reflected sound is performed by comparing a threshold value for the volume ratio of the direct sound to the reflected sound at a given time difference (T), i.e., a volume ratio threshold, with a volume ratio (L) calculated by the analysis unit 1301. For example, among the volume ratio thresholds preset for each value of the time difference between the direct sound and the reflected sound, the threshold value for the volume ratio at the time difference (T) between the direct sound and the reflected sound calculated by the analysis unit 1301 is referenced. Furthermore, whether the reflected sound is selected as the target reflected sound is determined based on whether the volume ratio (L) calculated by the analysis unit 1301 is higher than the threshold value.
[0267] The time difference (T) can be, for example, the difference in the time when the direct tone and the reflected tone arrive at the listening position, the time difference in the time required for the direct tone and the reflected tone to arrive at the listening position, or the time difference between the time when the direct tone ends and the time when the reflected tone arrives at the listening position. Here, the end time of the direct tone can also be obtained, for example, by adding the duration of the direct tone to the time of its arrival.
[0268] Regarding threshold data, it can also be determined through auditory nerve action or cognitive function of the brain. More specifically, it can be determined based on the minimum time difference between two sounds that the listener can perceive and detect, through the prioritization effect, the time-dependent masking phenomenon, or a combination thereof (described later). Specific values can be derived from known research findings on time-dependent masking effects, prioritization effects, or echo detection limits, or they can be determined through audiovisual experiments applied to this virtual space.
[0269] Figure 12A , Figure 12B and Figure 12C This is a diagram illustrating an example of how threshold data is set. For example... Figure 12A , Figure 12B and Figure 12C As shown in the figure, the threshold data is represented by the boundary (threshold) at which the reflected sound is perceived or not in the curve with the horizontal axis representing the time difference between the direct sound and the reflected sound and the vertical axis representing the volume ratio between the direct sound and the reflected sound.
[0270] Threshold data can also be approximated by using the time difference between the direct tone and the reflected tone as a variable. Furthermore, threshold data can also be used as... Figure 11 The indexes of the time differences between the direct and reflected tones, and the corresponding thresholds, are stored in the region of memory 1404.
[0271] In addition, Figure 12CIn Example 4, when the height of the line parallel to the horizontal axis (the minimum audible limit) is used as the threshold, the volume ratio (L) of the direct tone to the reflected tone is not compared to the threshold; rather, the volume of the reflected tone itself is compared to the threshold. This is because the threshold represents the volume boundary of whether a sound can be perceived by the listener, and it is the threshold used to determine sounds with volumes lower than this threshold as unreproducible sounds. That is, the threshold corresponding to the minimum audible limit is not a threshold for the ratio of the volume of the reflected tone to the volume of the direct tone.
[0272] When the minimum audible limit is used as the threshold, the threshold is constant and independent of the time difference (T), so the time difference (T) can be disregarded.
[0273] Additionally, in the parsing process ( Figure 8 In the case where multiple reflected sounds are generated in S101, selection processing can be performed on all reflected sounds, or selection processing can be performed only on reflected sounds with high evaluation values based on the evaluation values derived from each reflected sound using a pre-set evaluation method. Here, the evaluation value of a reflected sound corresponds to its perceived importance. Furthermore, a high evaluation value corresponds to a large evaluation value, and these interpretations can be interchanged.
[0274] The judgment unit 1302 may calculate the evaluation value of the reflected sound by, for example, a pre-set evaluation method corresponding to the volume of the sound source, the visuality of the sound source, the locality of the sound source, the visuality of the reflecting object (obstacle object), or the geometric relationship between the direct sound and the reflected sound.
[0275] Specifically, a higher volume of the sound source can result in a higher evaluation value. Furthermore, to ensure consistency between visual and acoustic localization, a higher evaluation value can also be achieved when the sound source object or its reflection (obstacle) is visible to the listener, or when the sound source object has high localization.
[0276] Furthermore, the opening of the angles of arrival of the direct and reflected tones, as well as the difference in their arrival times, significantly impact spatial perception. Therefore, a higher evaluation value can be achieved when the openings of the angles of arrival of the direct and reflected tones are large, or when the difference in their arrival times is significant.
[0277] The volume information of the audio source can also represent the base volume set for each content, the time transition of the volume, or both.
[0278] For example, in a virtual conference room where the virtual space is a virtual meeting room and the direct audio is conversational sound, the volume changes intermittently over a short period of time. That is, the audible and silent parts alternate. Conversely, in a concert hall where the virtual space is a concert hall and the direct audio is a musical performance, the volume is maintained for a certain duration. Finally, in a battlefield where the virtual space is a battlefield and the direct audio is an explosion, the volume increases only momentarily, then remains silent or low.
[0279] In this way, the volume information of the sound source not only includes information about the reference volume that corresponds to the volume setting when the sound is radiated into the virtual space, but also information about the changes in the volume of the sound.
[0280] Transition information can also be represented by time-series data showing frequency characteristics. Transition information can also be represented by data showing the duration of the sound interval. Transition information can also be represented by time-series data showing the duration of both the sound and silent intervals. Transition information can also be represented by multiple sets of time-series data listing the duration during which the amplitude of the sound signal can be considered stationary (approximately constant) and the amplitude values of the signal during that period.
[0281] Information about the transition can also be represented by data that shows the duration of a stationary frequency response of a sound signal. Alternatively, it can be represented by a time series listing multiple sets of data showing the duration of a stationary frequency response of a sound signal and the frequency response during that period.
[0282] Furthermore, there has been a long-standing and widespread practice of using temporal shifts in the frequency characteristics of signals for audio processing in virtual space (Patent Document 1, etc.). Given this prior art, the aforementioned group could also be a group of time lengths with constant frequency characteristics and their frequency characteristics.
[0283] Geometric relationships can also refer to the relationships between the positions of a sound source, a listener, and a reflecting object within a virtual space. Through these relationships, the path lengths of both direct and reflected sounds can be calculated geometrically. Therefore, by utilizing the inverse relationship between volume and distance, the reference volume of the reflected sound relative to the reference volume of the direct sound can be calculated.
[0284] In calculating the reference volume of reflected sound, the reflection coefficient of the reflecting object can also be used. Alternatively, a typical value commonly used can be used as the reflection coefficient. On the other hand, in special cases such as when the reflecting object is covered by sound-absorbing material, a specially assigned reflection coefficient can be used as the reflection coefficient of the reflecting object.
[0285] Reflected sound can also be evaluated based on its volume. The volume of the reflected sound can be determined based on the geometric relationship between the direct sound and the reflected sound as described above, as well as the index assigned to the reflecting object. The volume can also be compared to a pre-set threshold to evaluate the reflected sound.
[0286] Furthermore, information representing the temporal change in the volume of the sound source can also be reflected in the evaluation. For example, if the information representing the temporal change in the volume of the sound source represents the duration of the sound interval, the evaluation value of the reflected sound can remain unchanged when the time is within the sound interval. On the other hand, if the time is outside the sound interval, even if the reference volume of the reflected sound exceeds a threshold, the evaluation value of the reflected sound can be reduced or set to zero.
[0287] Alternatively, information representing the temporal shift in the volume of a sound source can also be data that lists multiple groups of amplitude values of the signal over a duration during which the amplitude of the sound signal is considered approximately constant, in a time series. In this case, the processing of the reflected sound can also be evaluated by changing the reference volume of the reflected sound in conjunction with the changes in the amplitude values in the data.
[0288] Alternatively, the volume information representing the direct tone can be obtained by using both the reference volume information and the volume information that changes over time. For example, after calculating an evaluation value based on the reference volume information, the evaluation value can be corrected using the volume information that changes over time.
[0289] In evaluating reflected sounds, all of the above methods can be performed, or only some of them can be performed. For example, reflected sounds can be evaluated using multiple evaluation methods, or they can be evaluated using only one evaluation method.
[0290] When evaluating reflected sounds using multiple evaluation methods, the decision to select a reflected sound can be based on the combined evaluation value determined by the multiple evaluation methods, or it can be based on the individual evaluation values of the multiple evaluation methods.
[0291] The sound signal processing device 1001 may select a sound if, when deciding whether to select a reflected sound based on each of multiple evaluation methods, all evaluation results based on multiple evaluation methods indicate that sound should be selected. Alternatively, the sound signal processing device 1001 may select a sound if one of the evaluation results based on multiple evaluation methods indicates that sound should be selected.
[0292] Alternatively, priorities can be set for the first to third evaluation methods. Furthermore, the sound signal processing device 1001 can also ultimately determine that sound is not selected if the first evaluation method determines that sound is not selected, regardless of the determination results in the second and third evaluation methods. Additionally, the sound signal processing device 1001 can also ultimately determine that sound is selected if one of the second and third evaluation methods determines that sound is not selected, but the other determines that sound is selected.
[0293] Furthermore, selection processing and evaluation processing can be performed independently, or only one of them can be performed. Alternatively, evaluation processing can be performed only on reflected sounds that were determined to be selected in the selection processing, and the evaluation processing can then determine whether to select the reflected sound again. Or, evaluation processing can be performed only on reflected sounds that were determined not to be selected in the selection processing, and the evaluation processing can then determine whether to select the reflected sound again.
[0294] The selection process described above can be interpreted as selecting reflected sounds based on the properties of the direct sound. For example, in selecting reflected sounds based on the properties of the direct sound, a threshold used in the selection of reflected sounds is set or adjusted according to the properties of the direct sound. Alternatively, an evaluation value used in the selection of reflected sounds can be calculated based on one or more of the following: the volume of the sound source, the visibility of the sound source, the localization of the sound source, the visibility of the reflecting object (obstacle object), and the geometric relationship between the direct sound and the reflected sound.
[0295] Furthermore, the process of selecting reflected tones based on the properties of direct tones is not limited to setting or adjusting thresholds based on the properties of direct tones, or calculating evaluation values used in selecting reflected tones of the processed object; other processes may also be performed. Moreover, when performing processes such as setting or adjusting thresholds based on the properties of direct tones, or calculating evaluation values used in selecting reflected tones of the processed object, a portion of the process may be modified, or new processes may be added.
[0296] In addition, setting thresholds can also include adjusting thresholds and changing thresholds.
[0297] [How to set the threshold] The threshold data used in the selection process can be set by referring to known values of echo detection limits based on priority effects or masking thresholds based on backmasking effects.
[0298] The priority effect refers to the phenomenon where, when two sounds are heard from different locations, the listener perceives which sound source was heard first in time. If two short sounds merge and sound like one, the overall sound's perceived location (locality) is largely determined by the location of the initial sound. The echo detection limit, a phenomenon occurring through the priority effect, is the minimum time difference that allows a listener to perceive the discrepancy between two sounds.
[0299] exist Figure 12C In Example 2, the horizontal axis corresponds to the arrival time of the reflected sound (echo), specifically, the delay time from the arrival time of the direct sound to the arrival time of the reflected sound. The vertical axis corresponds to the volume ratio of the detectable reflected sound to the direct sound, specifically, the threshold for whether the reflected sound arriving with the delay time can be detected.
[0300] Figure 13 This is a diagram illustrating an example of how a threshold is set. Figure 13 The horizontal axis in the figure corresponds to the arrival time of the reflected sound, specifically the time difference (T) between the direct sound and the reflected sound. Figure 13 The vertical axis corresponds to the volume of the reflected sound. Specifically, Figure 13 The vertical axis can correspond to either the volume of the reflected sound (volume ratio) which is set relatively to the volume of the direct tone, or the volume of the reflected sound which is determined absolutely without depending on the volume of the direct tone.
[0301] For example, in such Figure 9 When the listener is relatively far from the obstacle, the reflected sound arrives later, such as... Figure 13 As shown in C, the threshold is set low. As a result, in Figure 9 In such cases, reflected sounds are generated. On the other hand, in cases such as Figure 10 When the listener is relatively close to the obstacle, the arrival time of the reflected sound is... Figure 9 The condition occurs earlier, such as Figure 13 As shown in B, the threshold is set high. As a result, in Figure 10 In this case, no reflected sound is generated.
[0302] In addition, threshold data can also be stored in memory 1404 and retrieved from memory 1404 for use in selection processing.
[0303] Figure 14 This is a flowchart illustrating an example of the selection process. First, the determination unit 1302 specifies the reflected sound detected by the analysis unit 1301 (S201). Next, the determination unit 1302 detects the volume ratio (L) of the direct sound and the reflected sound, as well as the time difference (T) between the direct sound and the reflected sound (S202 and S203).
[0304] The time difference (T) can be, for example, the time difference between the arrival time of the direct tone and the reflected tone at the listening position, the time difference between the arrival time of the direct tone and the arrival time of the reflected tone, or the time difference between the end of the direct tone's sound and the arrival time of the reflected tone at the listening position. An example based on the time difference between the arrival time of the direct tone and the arrival time of the reflected tone is given here.
[0305] Specifically, the determination unit 1302 calculates the difference between the length of the direct sound path and the length of the reflected sound path based on the position information of the sound source object and the listener, as well as the position and shape information of the obstacle object. Furthermore, the determination unit 1302 detects the time difference (T) between the time when the direct sound arrives at the listener's position and the time when the reflected sound arrives at the listener's position by dividing this length by the speed of sound.
[0306] The volume reaching the listener decreases proportionally (inversely proportional to distance) to the volume of the sound source. Therefore, the volume of the direct tone is obtained by dividing the volume of the sound source by the length of the direct tone's path. The volume of the reflected tone is obtained by dividing the volume of the sound source by the length of the reflected tone's path and then multiplying by the attenuation rate assigned to the virtual obstacle object. The determination unit 1302 detects the volume ratio by calculating the ratio of their volumes.
[0307] Furthermore, the determination unit 1302 uses threshold data to determine a threshold corresponding to the time difference (T) (S204). Next, the determination unit 1302 determines whether the detected volume ratio (L) is above the threshold (S205).
[0308] If the volume ratio (L) is above the threshold ("Yes" in S205), the determination unit 1302 selects the reflected sound as the reflected sound of the generation object (S206). If the volume ratio (L) is below the threshold ("No" in S205), the determination unit 1302 does not select the reflected sound as the reflected sound of the generation object (S207). That is, in this case, the determination unit 1302 determines the reflected sound as a reflected sound outside the generation object.
[0309] Then, the determination unit 1302 determines whether there is an unspecified reflected sound (S208). If there is an unspecified reflected sound ("Yes" in S208), the determination unit 1302 repeats the above process (S201 to S207). If there is no unspecified reflected sound ("No" in S208), the determination unit 1302 ends the process.
[0310] This selection process can be performed on all reflected sounds generated in the parsing process, or only on reflected sounds with high evaluation values.
[0311] [Details on the threshold storage method] The threshold data for this embodiment is stored in the memory 1404 of the sound signal processing apparatus 1001. The stored threshold data can be of any form and type. When multiple forms and types of thresholds are stored, it is possible to determine which form and type of threshold to use for the selection processing of reflected sounds during the selection process. The method for determining which threshold data to use for the selection process will be described later.
[0312] Furthermore, threshold data of multiple forms and types can be combined and stored. The combined threshold data can also be read from the spatial information management units 1201 and 1211 and set as the threshold used in the selection process. Alternatively, threshold data stored in the memory 1404 can be stored in the spatial information management units 1201 and 1211.
[0313] Threshold data can also be stored, for example, as thresholds at various time differences, to depict... Figure 12C The threshold lines shown in [Example 1] and [Example 2] are examples of this.
[0314] In addition, threshold data can also be used as follows Figure 11 The diagram shows a table data storage system that establishes a correspondence between the threshold and the time difference (T). That is, the threshold data can also be stored as table data indexed by the time difference (T). Of course, Figure 11 The threshold shown is an example; the threshold is not limited to... Figure 11 Examples are provided. Alternatively, instead of storing the threshold itself, one can approximate the threshold with a function that takes the time difference (T) as a variable, and store the coefficients of that function. Furthermore, multiple approximations can be combined and stored.
[0315] For example, the time difference (T) can be set as timeDiff, and the threshold can be set as gainThresh, and the threshold data can be represented by the following formula.
[0316] [Mathematical Expression 1] The threshold is defined solely by the time range in which the priority effect occurs. When the time difference is outside this time range (a value less than 1 ms or greater than 40 ms in the above formula), a judgment based on gainThresh may not be performed, and the judgment may be made solely by the threshold representing the minimum volume reproduced in the virtual space, as described later.
[0317] Through experiments, the inventors have determined that, within the timeframe in which the preference effect occurs, the threshold is preferably approximated by an upwardly convex function. The above formula is an example of an approximation generated based on this experiment.
[0318] The memory 1404 may also store information related to the relational formulas representing the relationship between time difference (T) and threshold. That is, it may also store formulas that take time difference (T) as a variable. The threshold for each time difference (T) may also be approximated by a straight line or curve, and parameters representing the geometric shape of the straight line or curve may be stored. For example, if the geometric shape is a straight line, the starting point and slope of the straight line may also be stored.
[0319] Furthermore, the type and format of threshold data can be set and stored according to each property of the direct tone. Additionally, parameters used to adjust the threshold based on the property of the direct tone and for selection processing can be stored. The process of adjusting the threshold based on the property of the direct tone and for selection processing will be described later as a variation of the threshold setting method.
[0320] As an example of storing multiple threshold data combinations, it can also be like... Figure 12C As shown in [Example 3], the larger of the masking threshold and the echo detection limit threshold is stored for each time difference (T). Alternatively, as... Figure 12C As shown in [Example 4], the value of the larger of the minimum volume and the echo detection limit threshold that are reproduced in the virtual space is stored for each time difference (T).
[0321] The combination of multiple types of threshold data is not limited to this. For example, information on the maximum value can also be stored in multiple threshold data for each time difference (T).
[0322] Furthermore, in the above, the information related to the threshold has time items as a one-dimensional index. The information related to the threshold can also have a two-dimensional or three-dimensional index that also includes variables related to the direction of arrival.
[0323] Figure 15 It is a graph showing the relationship between the direction of the direct tone, the direction of the reflected tone, the time difference, and the threshold. For example, as... Figure 15 As shown, thresholds can also be pre-calculated based on the relationship between the direction of the direct tone (θ), the direction of the reflected tone (γ), the time difference (T), and the volume ratio (L).
[0324] The direction of the direct tone (θ) corresponds to the angle relative to the direction in which the direct tone arrives at the listener. The direction of the reflected tone (γ) corresponds to the angle relative to the direction in which the reflected tone arrives at the listener. Here, the direction the listener is facing is set to 0 degrees. The time difference (T) corresponds to the difference between the arrival time of the direct tone and the arrival time of the reflected tone towards the listening position. The volume ratio (L) corresponds to the volume ratio of the volume of the direct tone arrival to the volume of the reflected tone arrival.
[0325] certainly, Figure 15 The threshold shown is an example; the threshold is not limited to... Figure 15Examples. Furthermore, in Figure 15 The example primarily illustrates the threshold when the angle (θ) of the direction of arrival of the direct tone is 0 degrees. However, the threshold for cases where the direction of arrival of the direct tone (θ) is other than 0 degrees is also stored in memory 1404.
[0326] Furthermore, in the above, the threshold is stored as an arrangement where the direction of the direct tone (θ) (more specifically, the angle (θ) of the direction of arrival of the direct tone) and the direction of the reflected tone (γ) (more specifically, the angle (γ) of the direction of arrival of the reflected tone) are treated as independent variables or indices. However, the angle (θ) of the direction of arrival of the direct tone and the angle (γ) of the direction of arrival of the reflected tone may also not be used as independent variables.
[0327] For example, the angle difference between the direction of arrival of the direct tone (θ) and the direction of arrival of the reflected tone (γ) can also be used. This angle difference corresponds to the angle formed by the direction of arrival of the direct tone and the direction of arrival of the reflected tone, and can also be expressed as the arrival angles of the direct tone and the reflected tone.
[0328] Figure 16 This is a graph representing the relationship between angular difference, time difference, and threshold. For example, it can also be like... Figure 16 The example shown stores a pre-calculated threshold, using the angle difference (Φ) between the angle of arrival of the direct sound (θ) and the angle of arrival of the reflected sound (γ) as a variable. Of course, Figure 16 The threshold shown is an example; the threshold is not limited to... Figure 16 Examples.
[0329] exist Figure 16 In the example, the number of variables used in deriving the threshold can be reduced. Therefore, the number of thresholds stored in memory 1404 can be reduced. Consequently, the amount of data stored in memory 1404 can be reduced.
[0330] Furthermore, when using the angle difference (Φ) between the angle (θ) of the direction of arrival of the direct tone and the angle (γ) of the direction of arrival of the reflected tone, the threshold data can also be stored in a two-dimensional arrangement. Additionally, in the selection process, a three-dimensional arrangement can be used to calculate the difference between the angle (θ) of the direction of arrival of the direct tone and the angle (γ) of the direction of arrival of the reflected tone.
[0331] The method of selecting reflected sounds using a threshold corresponding to the direction of arrival will be described later.
[0332] [First variation of the threshold setting method] exist Figure 12A , Figure 12B and Figure 12CIn the example, multiple forms and types of thresholds can also be stored in the spatial information management units 1201 and 1211. Furthermore, it can be determined which form and type of threshold among the multiple forms and types will be used for the selection processing of reflected sounds. Specifically, it can also be done as follows: Figure 12C As shown in Example 3, the highest threshold is used when the time difference (T) corresponds to the arrival time of the reflected sound.
[0333] Alternatively, as shown in Example 4, a masking threshold, an echo detection limit threshold, and a threshold representing the minimum volume reproduced in the virtual space can be stored. Furthermore, the highest threshold can be used for the time difference (T) corresponding to the arrival time of the reflected sound.
[0334] [Second variation of the threshold setting method] As another example of a method for setting a threshold, a method for setting a threshold based on the properties of a direct tone will be explained.
[0335] Figure 17 It means Figure 7 A block diagram of another configuration example of the rendering unit 1300 shown. Figure 17 The rendering unit 1300 and Figure 7 The rendering unit 1300 differs from the one in that it includes a threshold adjustment unit 1304. The description other than the threshold adjustment unit 1304 is the same as that in... Figure 7 The content described in the previous section is the same, so it is omitted.
[0336] The threshold adjustment unit 1304 selects a threshold to be used by the determination unit 1302 from the threshold data based on information representing the nature of the sound signal. Alternatively, the threshold adjustment unit 1304 may adjust the threshold included in the threshold data based on information representing the nature of the sound signal.
[0337] Information indicating the nature of the sound signal can also be included in the input signal. Furthermore, the threshold adjustment unit 1304 can also obtain information indicating the nature of the sound signal from the input signal. Alternatively, the analysis unit 1301 can analyze the sound signal contained in the received input signal, derive the nature of the sound signal, and output information indicating the nature of the sound signal to the threshold adjustment unit 1304.
[0338] Information representing the nature of a sound signal can be obtained either before rendering begins or at any time during rendering.
[0339] Furthermore, the threshold adjustment unit 1304 may not be included in the sound signal processing device 1001, or it may function as a threshold adjustment unit 1304 in another communication device. In this case, the analysis unit 1301 or the determination unit 1302 may obtain information representing the nature of the sound signal, threshold data corresponding to the nature, or information for adjusting the threshold data according to the nature from other communication devices via the communication IF 1403.
[0340] Figure 18 This is a flowchart representing another example of the selection process. Figure 19 This is another flowchart illustrating the selection process. Figure 18 and Figure 19 In this context, a threshold is set based on the properties of the direct sound. Specifically, in... Figure 18 In this process, the threshold adjustment unit 1304 determines the threshold from the threshold data based on the time difference (T) and the properties of the sound signal. Figure 19 In the process, the threshold adjustment unit 1304 adjusts the threshold determined from the threshold data based on the time difference (T) based on the properties of the sound signal.
[0341] The actions in each example are explained below. Additionally, regarding... Figure 14 The examples commonly use omitting explanations.
[0342] First, let me explain in Figure 18 The following is an example of the processing. Here, threshold data is pre-stored in memory 1404 according to each property of the direct tone. Thus, multiple threshold data corresponding to multiple properties are pre-stored in memory 1404. Furthermore, the threshold adjustment unit 1304 determines the threshold data to be used in the selection processing of the reflected tone from the multiple threshold data.
[0343] For example, the threshold adjustment unit 1304 obtains the properties of the direct tone based on the input signal (S211). The threshold adjustment unit 1304 may also obtain the properties of the direct tone that are associated with the input signal. Next, the threshold adjustment unit 1304 determines a threshold corresponding to the time difference (T) and the properties of the direct tone (S212).
[0344] In addition, such as Figure 19 As shown, the threshold adjustment unit 1304 can also adjust the threshold determined by the determination unit 1302 based on the properties of the direct tone (S221).
[0345] In any case, the input signal may include information representing the nature of the sound signal, information for adjusting the threshold according to the nature of the sound signal, or both. The threshold adjustment unit 1304 may also use one or both of these to adjust the threshold.
[0346] Furthermore, information indicating the nature of the sound signal, information used to adjust the threshold, or both, can be transmitted via an input signal different from the input signal containing the sound signal. In this case, information relating to an input signal different from the input signal can also be included in the input signal containing the sound signal, or the information relating to the input signal different from the input signal can be stored in memory 1404 along with information about the threshold.
[0347] exist Figure 18 and Figure 19 In the example, the threshold used in selecting the reflected tone is set based on the properties of the direct tone, i.e., the properties of the sound signal. This can be achieved as follows: Figure 18 Using pre-defined threshold data for each property, it is also possible to... Figure 19 In this way, the threshold can be adjusted based on the properties of the sound signal. Alternatively, the parameters of the threshold data can also be adjusted based on the properties of the sound signal.
[0348] Furthermore, the actions performed by the threshold adjustment unit 1304 can also be performed by the analysis unit 1301 or the determination unit 1302. For example, the analysis unit 1301 may obtain the properties of the sound signal. Alternatively, the determination unit 1302 may set the threshold based on the properties of the sound signal.
[0349] Next, the relationship between the properties of the sound signal and the threshold will be explained.
[0350] Two short sounds arriving at a listener's ear consecutively, if the time interval between them is sufficiently short, are perceived as a single sound. This phenomenon is called the priority effect. It is known that the priority effect occurs only for discontinuous, i.e., transient sounds (Non-Patent Document 1). Therefore, when the sound signal represents a stationary tone, the echo detection limit can be set lower compared to when the sound signal represents a non-stationary tone.
[0351] That is, based on the characteristics of this priority effect, for example, when the direct tone is a stable sound, the threshold is set to be relatively small. Alternatively, the higher the stability, the smaller the threshold can be set.
[0352] An example of processing when the sound signal is stationary will be explained. First, the threshold adjustment unit 1304 or the analysis unit 1301 determines the stationarity based on the amount of change in the frequency components of the sound signal over time. For example, if the amount of change is small, the stationarity is determined to be high. Conversely, if the amount of change is large, the stationarity is determined to be low. The result of the determination can be used to set a flag representing the level of stationarity, or a parameter representing stationarity can be set based on the amount of change.
[0353] Next, the threshold adjustment unit 1304 may adjust the threshold data or threshold based on information indicating the stability of the sound signal, such as a flag or parameter, and set the adjusted threshold data or threshold as the threshold data or threshold used in the determination unit 1302.
[0354] Alternatively, parameters for setting threshold data based on information representing the stability of the direct tone can be pre-stored in the memory 1404. In this case, the threshold adjustment unit 1304 can also determine the stability of the sound signal and set the threshold data used in selecting the reflected tone based on the information and parameters representing the stability.
[0355] Alternatively, multiple parameters of the threshold data may be pre-stored in the memory 1404 corresponding to multiple patterns of the direct tone's stability. In this case, the threshold adjustment unit 1304 may also determine the stability of the sound signal, select parameters of the threshold data based on the pattern of the direct tone's stability, and set the threshold data used in the selection of reflected tones based on the parameters of the threshold data.
[0356] In addition, the stability of a sound signal can be determined based on the amount of change in the frequency components of the sound signal each time a sound signal is input.
[0357] Alternatively, the stationarity of the sound signal can be determined based on information representing stationarity that has been pre-associated with the sound signal. That is, information representing the stationarity of the sound signal can be pre-associated with the sound signal and stored in the memory 1404. The parsing unit 1301 can also obtain the information representing stationarity associated with the sound signal each time an input sound signal is received. Furthermore, the threshold adjustment unit 1304 can adjust the threshold based on the information representing stationarity associated with the sound signal.
[0358] As another example of setting a threshold based on the properties of the sound signal, the application range of the echo detection limit can be set shorter when the sound signal represents a shorter sound (such as a click) compared to when the sound signal represents a longer sound. This processing is based on the characteristics of the priority effect.
[0359] It is known that, through the priority effect, two short sounds arriving consecutively at a listener's ear are perceived as a single sound if the time interval between them is sufficiently short. The upper limit of this time interval depends on the length of the sound. For example, the upper limit of this time interval is approximately 5 ms for a click sound, and sometimes as high as 40 ms for complex sounds such as human voices or music (Non-Patent Document 1).
[0360] Based on this priority effect, for example, in the case of sounds with shorter direct durations, a shorter duration threshold is set. Furthermore, the shorter the direct duration, the shorter the duration threshold is set.
[0361] Setting a shorter time threshold means setting a threshold corresponding to the echo detection limit based on the priority effect characteristics within a range where the time difference (T) between the direct tone and the reflected tone is small. Outside this range, no threshold corresponding to the echo detection limit based on the priority effect characteristics is set. That is, outside this range, the threshold is small. Therefore, setting a shorter time threshold for shorter sounds corresponds to setting a smaller threshold for shorter sounds.
[0362] As another example of setting a threshold based on the nature of the direct sound, the threshold can be set lower when the direct sound is an intermittent sound (speech, etc.) compared to when the direct sound is a continuous sound (music, etc.).
[0363] For example, when the direct sound corresponds to speech, there are repeated vocal and non-vocal parts, and as a masking effect, only the aftermasking effect occurs in the non-vocal part. On the other hand, when the direct sound is a continuous sound like musical content, both the aftermasking effect and the simultaneous masking effect based on the sound produced at that time occur. Therefore, the comprehensive masking effect is higher in the case of music, etc., than in the case of speech, etc.
[0364] Based on the masking effect characteristics described above, the threshold can be set higher in the case of music, etc., compared to the case of speech, etc. Conversely, the threshold can be set lower in the case of speech, etc., compared to the case of music, etc. That is, the threshold can be set lower when there are many interruptions in the direct sound.
[0365] As mentioned above, information indicating the properties of a direct tone can also include information about its smoothness, discontinuity, and duration. Furthermore, information indicating the properties of a direct tone can be any combination of these characteristics. Additionally, information indicating the properties of a direct tone can be information about the temporal variation of any one of these characteristics, or information about the temporal variation of any combination thereof. In other words, information indicating the properties of a direct tone can also be information about the temporal variation of the direct tone.
[0366] For example, as shown in the explanation of stationarity determination, information representing the properties of a direct tone can also be time-series data of frequency characteristics. Here, frequency characteristics can also be represented in conventional forms such as gain values for each frequency band, Fourier series of the signal over the time axis, or LPC coefficients or cepstral coefficients used to calculate the frequency envelope.
[0367] Furthermore, information representing the properties of a direct tone can also be presented as information representing the discontinuity of the direct tone, listing multiple sets of information (the approximate shape of the amplitude envelope) of the signal's amplitude stability over a time series. Here, the amplitude value can also be expressed as a ratio relative to a reference volume.
[0368] Furthermore, the information representing the properties of a direct tone can also be information related to the frequency characteristics of the direct tone. For example, the information representing the properties of a direct tone can also be information representing the stability of the frequency characteristics of the direct tone. Specifically, the information representing the properties of a direct tone can also be information that lists multiple groups of frequency characteristics of the signal during a period of small frequency variation in a time series (approximate shape of a spectrum). Here, the volume used as a reference for the aforementioned frequency characteristics can also be the aforementioned reference volume.
[0369] For example, information indicating the temporal variation of a direct tone is information representing the envelope of the direct tone. Information indicating the temporal variation of a direct tone can also be found in... Figure 12C The “minimum audible limit” described in [Example 4] is used when the threshold is a threshold. The signal compared with the minimum audible limit is the volume of the reflected sound.
[0370] The volume of the reflected sound is obtained through geometric calculations based on the positions of the sound source, the listener, and the reflecting object. Specifically, a reference volume of the reflected sound is obtained relative to a reference volume of the sound source. By using information about changes in the volume of the sound source as information representing the properties of the direct sound, the reference volume of the reflected sound is increased or decreased, allowing for an accurate determination of the volume of the reflected sound at any given moment. This is because changes in the volume of the sound source are reflected in changes in the volume of the reflected sound.
[0371] After adjusting the volume of the reflected sound, by comparing the volume of the reflected sound with the threshold, it is possible to more accurately and appropriately select the reflected sound that is audibly desired.
[0372] Of course, the same result can be obtained by adjusting the threshold based on the reciprocal of the change in the volume of the sound source, without adjusting the reference volume of the reflected sound, and then comparing the adjusted threshold with the reference volume of the reflected sound. That is, the reference volume of the reflected sound can be adjusted using information about the change in the volume of the sound source, and the threshold can also be adjusted using the same information. The adjustment of the reference volume of the reflected sound and the adjustment of the threshold are mutually corresponding.
[0373] Depending on the surface composition of the object reflecting the sound, the reflectivity of the sound (and the attenuation rate of the reflected sound) varies in each frequency band. Therefore, as described later, the reflectivity (attenuation rate) of the sound can also be associated with each frequency band for the object reflecting the sound. Based on this reflectivity information and the information from the spectrogram, it is possible to more accurately determine whether to select the reflected sound. For example, the following process can be performed.
[0374] Specifically, for example, information from the spectrogram indicates that frequency components in the high-frequency band are more dominant than those in the low-frequency band within a certain time interval. Additionally, for example, information from the reflectivity of sound indicates that the reflectivity of frequency components in the high-frequency band is extremely low compared to that in the low-frequency band.
[0375] In this case, even if the signal amplitude of the sound source is large on the time axis, the volume of the reflected sound is reduced by multiplying the frequency components represented by the information in the spectrogram with the attenuation rate of each frequency band represented by the information in the reflectivity, and the reflected sound may not be selected.
[0376] As mentioned above, information representing the properties of a direct tone can also be information representing the temporal variation of the direct tone. For example, information representing the properties of a direct tone can also represent values obtained by analyzing the direct tone over a predetermined time period.
[0377] Specifically, information representing the properties of a direct tone can also be obtained by calculating the average energy or average amplitude of the direct tone for each predetermined time length. Alternatively, information representing the properties of a direct tone can also be obtained by calculating the energy or average amplitude of the direct tone for each short time analysis length and then taking a weighted average of the energy or average amplitude for each long time analysis length that is longer than the short time analysis length.
[0378] More specifically, for example, information representing the temporal variation of the direct tone can also be obtained by calculating the energy or average amplitude of the direct tone for each pre-defined short time length (e.g., 5 ms, hereinafter referred to as an analysis frame). Alternatively, information representing the temporal variation of the direct tone can also be represented by a weighted average of the energy or average amplitude calculated over the past N-1 analysis frames.
[0379] Assuming the energy of the nth analysis frame is represented by E(n), the information I(n) representing the properties of the direct tone is obtained according to the following formula.
[0380] [Mathematical Expression 2] Here, the parameter a(i) represents the weighting coefficient. Typically, a(i) is set such that a(i) ≥ 0 and the sum of a(i) is 1. However, the method of setting a(i) is not limited to this.
[0381] Furthermore, information I(n) representing the properties of the direct tone is calculated every 5ms since the direct tone was received. That is, the temporal variation of information I(n) representing the properties of the direct tone can be calculated with low latency. Therefore, this method is suitable for applications requiring real-time performance.
[0382] Alternatively, the information I(n) representing the properties of direct sounds can be obtained using the following formula.
[0383] [Mathematical Expression 3] Here, the parameter b(i) represents the weighting coefficient. Typically, b(i) is set such that b(i) ≥ 0 and the sum of b(i) is 1. However, the method of setting b(i) is not limited to this.
[0384] In this formula, information I(n) representing the properties of the direct tone is recursively obtained. Therefore, the average energy over a long time period can be calculated with minimal computation.
[0385] Equations 1 and 2 above can be viewed as filters where E(n) is the input signal and I(n) is the output signal. In this case, Equation 1 is a filter for the moving average (MA) model, and Equation 2 is a filter for the autoregressive (AR) model, both exhibiting the characteristics of low-pass filters. Alternatively, a filter for the ARMA model, which combines both, can also be used.
[0386] Furthermore, the method for deriving information representing the temporal variation of the direct tone is not limited to the aforementioned calculation formula or filter; other known methods can also be used. As mentioned above, the information representing the temporal variation of the direct tone represents a value obtained by analyzing the direct tone over a predetermined time length. The direct tone can also be analyzed from a perspective other than average energy.
[0387] Furthermore, as mentioned above, the information representing the properties of a direct tone can also be information related to the frequency characteristics of the direct tone. This information related to the frequency characteristics of the direct tone can also be information calculated using those characteristics. For example, information related to the frequency characteristics of the direct tone can be obtained by averaging the low-frequency components of the direct tone over a predetermined analysis length to obtain the average energy of the low-frequency components.
[0388] Specifically, the low-frequency components of the direct tone are determined by applying a low-pass filter to the direct tone contained in the analysis frame length. Based on the energy or average amplitude of this low-frequency component, information representing the properties of the direct tone is derived in the same manner as in Equation 1 above.
[0389] Assume the energy of the low-frequency components in the nth analysis frame is determined by E LIn the case of (n), the information I(n) representing the nature of the direct sound is obtained according to the following formula.
[0390] [Mathematical Expression 4] Here, the parameter c(i) represents the weighting coefficient. Typically, c(i) is set such that c(i) ≥ 0 and the sum of c(i) is 1. However, the method of setting c(i) is not limited to this.
[0391] Furthermore, information I(n) representing the properties of the direct tone is calculated every 5ms since the direct tone was received. That is, the temporal variation of information I(n) representing the properties of the direct tone can be calculated with low latency. Therefore, this method is suitable for applications requiring real-time performance.
[0392] In addition, similar to Equation 2, the information I(n) representing the properties of direct sounds can also be obtained according to the following formula.
[0393] [Mathematical Expression 5] Here, the parameter d(i) represents the weight coefficient. Typically, d(i) is set such that d(i) ≥ 0 and the sum of d(i) is 1. However, the method of setting d(i) is not limited to this.
[0394] In this formula, information I(n) representing the properties of the direct tone is recursively obtained. Therefore, the average energy over a long time period can be calculated with minimal computation.
[0395] Equations 3 and 4 above can be considered as filters with E(n) as the input signal and I(n) as the output signal. In this case, Equation 3 is a filter for the moving average (MA) model, and Equation 4 is a filter for the autoregressive (AR) model, both of which have the characteristics of low-pass filters. Alternatively, a filter for the ARMA model, which combines the two, can also be used.
[0396] In the above method for determining the low-frequency components of a direct tone, a filter with low-pass characteristics is used; however, the method for determining the low-frequency components of a direct tone is not limited to this. Furthermore, the method for deriving information representing the time variation of the direct tone is not limited to the aforementioned calculation formula or filter; other known methods can also be used. For example, the spectrum of a direct tone can be calculated by performing a frequency transformation on the direct tone. Furthermore, the energy or average amplitude of the low-frequency components of the spectrum can be calculated.
[0397] Furthermore, in the above, MA or AR models are used to derive information representing the temporal variation of direct tones. The coefficients of these models can be pre-set fixed values or time-varying values.
[0398] In addition, the relationship between the analysis frame length and the information update thread generation interval can also be as follows.
[0399] For example, when the analysis frame duration is TA (msec) and the information update thread generation interval is TU (msec), the value of N in Equations (1) and (3) above in the MA filter can also be approximately equal to the value given by TU / TA. Additionally, b(i) and d(i) (1≤i<N) in Equations (2) and (4) above in the AR filter can also be approximately equal to the filter's time constant, which is approximately equal to TU (msec).
[0400] The reason for this setting is that the filter is expected to converge during the information update interval.
[0401] On the other hand, in the above-described configuration, if the value of the information representing the temporal variation of the direct tone changes too drastically, I(n) can be pre-calculated. Furthermore, the pre-calculated I(n) can be applied to the selection processing of reflected tones. For example, in the processing of the frame at time t, I(t+tau) can be used. Here, tau is a value determined based on the convergence characteristics of the filter. In cases of slow convergence, the value of tau is larger compared to cases of fast convergence.
[0402] Furthermore, auditory masking (frequency masking) information calculated based on the direct tone can also be used as information representing the characteristics of the direct tone. The auditory masking information represents a threshold value for the amplitude in the frequency domain that is masked by the direct tone. It is also possible to compare the amplitude value of reflected tones in the same frequency domain with the threshold value, without selecting reflected tones with amplitude values smaller than the threshold. The amplitude value of reflected tones in the frequency domain can also be obtained by the analysis unit 1301 as information representing the characteristics of the reflected tones.
[0403] By setting the threshold used in selecting reflected sounds based on the properties of the direct sound, the reflected sounds that are audibly required can be appropriately selected, and auditory characteristics can be effectively reflected in the stereo sound reproduction system 1000. The processing of detecting the properties of the direct sound, determining the threshold based on the properties, and adjusting the threshold based on the properties can be performed either during the rendering process or before the rendering process begins.
[0404] For example, these processes can occur during virtual space creation (when the software is created), at the start of virtual space processing (when the software starts or rendering begins), or at timed intervals in information update threads that occur periodically during virtual space processing. Furthermore, virtual space creation can be timed to build the virtual space before the start of sound processing, or it can occur when virtual space information (spatial information) is acquired, or it can occur when the software acquires the information.
[0405] Here, in the information update thread, processing is performed to update the spatial information managed by the Spatial Information Management Departments 1201 and 1211.
[0406] The information update thread is responsible for tasks such as updating the position and orientation of the listener's avatar configured in the virtual space based on the position and orientation of the VR goggles worn by the listener, or updating the position of moving objects in the virtual space. This processing is provided within a processing thread that starts at a relatively low frequency of around tens of Hz.
[0407] In such low-frequency processing threads, the processing of information representing the properties of direct tones can still be performed. This is because the frequency of changes in the properties of direct tones is lower than the frequency of audio processing frames used for audio output. Therefore, the computational load of this processing can be relatively reduced. Furthermore, updating information at an unnecessarily fast frequency carries the risk of generating impulse noise. This risk can be avoided by updating information at a low frequency.
[0408] [Third variation of the threshold setting method] As another example of a method for setting the threshold, the threshold can also be set based on the computing resources (CPU power, memory resources, PC performance, or remaining battery power, etc.) used to process the reproduction of the virtual space. More specifically, the sensor 1405 of the sound signal processing device 1001 detects the amount of computing resources and sets a higher threshold when the amount of computing resources is low. As a result, the volume of more reflected sounds is lower than the threshold, thus reducing the reflected sounds that are processed by both ears and reducing the amount of computing power.
[0409] Alternatively, in cases where signal processing is performed by battery-powered devices such as smartphones or VR headsets, it is desirable to prioritize long processing times and conserve computing resources. In such cases, the threshold can be set high without checking the amount or remaining amount of computing resources.
[0410] [Fourth variation of the threshold setting method] As another example of a method for setting a threshold, the virtual space administrator or listener may set the threshold by having a threshold setting unit (not shown) in the sound signal processing device 1001 or the sound prompting device 1002.
[0411] For example, the listener wearing the sound prompt device 1002 could choose between an "energy-saving mode" (fewer reflected sounds and less computation) and a "high-performance mode" (more reflected sounds and more computation). Alternatively, the administrator of the stereo sound reproduction system 1000 or the producer of the stereo sound content could choose the mode. Furthermore, instead of a mode, a threshold or threshold data could be directly selected.
[0412] [The first variation of the rendering department's actions] Figure 20 This is a flowchart illustrating a first variation of the operation of the sound signal processing device 1001. Figure 20 The diagram shows the processing mainly performed by the rendering unit 1300 of the sound signal processing device 1001. In this modified example, volume compensation processing is added to the operation of the rendering unit 1300.
[0413] For example, the analysis unit 1301 acquires data (input signal) (S301). Next, the analysis unit 1301 analyzes the data (S302). Next, the determination unit 1302 determines whether to select reflected sounds based on the analysis results (S303). Next, the reproduction unit 1303 performs volume compensation processing based on the unselected reflected sounds (S304). Next, the reproduction unit 1303 performs audio processing on both direct and reflected sounds (S305). Finally, the reproduction unit 1303 outputs both direct and reflected sounds as audio (S306).
[0414] In the processes described above (S301 to S306), the processes other than the volume compensation process (S304) are common to the other examples described above, so their descriptions are omitted.
[0415] Volume compensation processing is performed for reflected sounds that were not selected in the selection process. For example, by not selecting reflected sounds in the selection process, a lack of volume perception occurs. Volume compensation processing can suppress the unpleasantness that accompanies this lack of volume perception. Two methods are disclosed as examples of methods for compensating for volume perception. Either method can be used.
[0416] First, the method of compensating for the sense of volume by increasing the volume of the direct tone will be explained. The reproduction unit 1303 generates the direct tone by increasing its volume by an amount corresponding to the volume of the unselected reflected tone. As a result, the sense of volume lost due to the lack of reflected tone generation is compensated.
[0417] When the reproduction unit 1303 increases the volume, it can also increase the volume according to the frequency characteristics of the reflected sound, one frequency component at a time. To enable this process, a predetermined attenuation rate for the volume of the reflected sound can be assigned to each frequency band. Thus, the frequency characteristics of the reflected sound can be derived.
[0418] Next, a method for compensating for the sense of volume by synthesizing reflected sounds into direct sounds will be explained. In this method, the reproduction unit 1303 adds unselected reflected sounds to the direct sounds to generate a direct sound, thereby compensating for the sense of volume caused by the lack of generated reflected sounds. The generated direct sound reflects the volume (amplitude), frequency, and delay of the unselected reflected sounds.
[0419] In the case of increasing the volume of the direct tone, the computational workload of the compensation process is very small, but only the volume is compensated. In the case of synthesizing the reflected tone into the direct tone, the computational workload of the compensation process is larger compared to the method of increasing the volume of the direct tone, but the characteristics of the reflected tone are compensated more accurately.
[0420] In both cases, no reflected tones are generated, only direct tones, thus reducing the overall computational load. In particular, the computational load required for binaural processing, which includes convolutional HRTF processing, is reduced, resulting in a significant reduction in overall computational load. This is because the computational load required for binaural processing is far greater than that required for the aforementioned compensation processing.
[0421] In addition, if the reason for not selecting the reflected sound is that the volume of the reflected sound is lower than the masking threshold, since the sense of volume will not be lost, the reflected sound can be removed without compensation.
[0422] [Second variation of the rendering department's actions] Figure 21 This is a flowchart illustrating a second variation of the operation of the sound signal processing device 1001. Figure 21 The text primarily describes the processing performed by the rendering unit 1300 of the sound signal processing device 1001. In this modified example, the operation of the rendering unit 1300 includes left and right volume difference adjustment processing.
[0423] For example, the analysis unit 1301 analyzes the input signal (S401). Next, the analysis unit 1301 detects the direction of sound arrival (S402). Next, the determination unit 1302 adjusts the difference in volume between the sound perceived by the left and right ears (S403). Furthermore, the determination unit 1302 adjusts the difference in arrival time (delay) between the sound perceived by the left and right ears (S404). Based on the adjusted sound information, the determination unit 1302 determines whether to select the reflected sound (S405).
[0424] In the above-described processes (S401 to S405), the processes other than the left and right volume difference adjustment process (S403) and the delay adjustment process (S404) are common to the other examples described above, so their descriptions are omitted.
[0425] Figure 22 This is a diagram illustrating the configuration of avatars, sound source objects, and obstacle objects. For example, in the case where the listener's facing direction is 0 degrees, such as... Figure 22 As shown, if the polarity (e.g., positive or negative) of the direction of arrival of the direct tone (θ) and the direction of arrival of the reflected tone (γ) (the direction of the reflected tone (γ)) is different, the volume difference generated between the two ears is corrected.
[0426] Specifically, when the polarities of θ and γ are different, the ear that primarily (first) perceives the sound in the direct tone and the reflected tone is different. In this case, the determination unit 1302 performs a left-right volume difference adjustment process (S403) to adjust the volume of the direct tone according to the position of the ear that primarily perceives the reflected tone. For example, the determination unit 1302 attenuates the volume of the direct tone when it reaches the listener by multiplying the volume by (1.0 - 0.3sin(θ)) (0 ≤ θ ≤ 180°).
[0427] The determination unit 1302 calculates the volume ratio of the corrected direct tone volume to the reflected tone volume as described above, and compares the calculated volume ratio with a threshold to determine whether to select the reflected tone. As a result, the volume difference between the two ears is corrected, the volume of the direct tone affecting the reflected tone is derived more accurately, and the determination of whether to select the reflected tone is made more accurately.
[0428] In addition to adjusting the left and right volume difference (S403), the determination unit 1302 can also perform a delay adjustment process (S404) to match the position of the ear that perceives the reflected sound, thus delaying the arrival time of the direct sound. Specifically, the determination unit 1302 can also delay the arrival time of the direct sound by adding (a(sinθ+θ) / c)ms (where a is the radius of the head and c is the speed of sound) to the arrival time of the direct sound.
[0429] [The third variation of the rendering department's actions] The method for setting a threshold corresponding to the direction of arrival is explained.
[0430] Figure 23 This is another flowchart illustrating the selection process. Regarding... Figure 14 Examples of this type of writing commonly involve omitting explanatory notes. Figure 23 In the example, the determination unit 1302 uses a threshold corresponding to the direction of arrival to select the reflected sound.
[0431] Specifically, the determination unit 1302 calculates the direct sound arrival path (pd), the reflected sound arrival path (pr), and the orientation information D1 of the avatar, using the orientation of the avatar as a reference, the direct sound arrival direction (θ) and the reflected sound arrival direction (γ) (the direction of the reflected sound (γ)). That is, the determination unit 1302 detects the direct sound arrival direction (θ) and the reflected sound arrival direction (γ) (S231). The orientation of the avatar corresponds to the orientation of the listener. The orientation information D1 of the avatar may also be included in the input signal.
[0432] The determination unit 1302 uses three indices, including the direct direction of arrival (θ) and the reflected direction of arrival (γ), as well as the time difference (T), according to... Figure 15 The three-dimensional arrangement shown determines the threshold used in the selection process (S232).
[0433] As an example, to illustrate in such Figure 22 The method shown describes the threshold setting method used in the selection process when an avatar, a sound source object, and an obstacle object are configured.
[0434] The position information of the avatar, sound source object, and obstacle object, as well as the orientation information D1 of the avatar, are obtained from the input signal. Using this position information and orientation information D1, the direction of the direct sound (θ) and the direction of the sound image of the reflected sound are calculated when the orientation of the avatar is set to 0 degrees. Figure 22 In this case, the direction of the direct sound (θ) is about 20 degrees, and the direction of the sound image of the reflected sound (γ) is about 265 degrees (-95 degrees).
[0435] Next, refer to Figure 15 The threshold data shown is stored in a three-dimensional arrangement. The threshold is determined from the arrangement region corresponding to the values of the two directions (θ) and (γ) and the value of the time difference (T) calculated by the analysis unit 1301. Even if there is no index corresponding to the calculated values of (θ), (γ), and (T), the threshold corresponding to the nearest index can be determined.
[0436] Alternatively, the threshold can be determined by interpolation, extrapolation, or other processing based on one or more thresholds corresponding to one or more indices close to the calculated values of (θ), (γ), and (T). For example, the threshold corresponding to (20 degrees, 265 degrees, T) can be determined based on four thresholds corresponding to the four indices (0 degrees, 225 degrees, T), (0 degrees, 270 degrees, T), (45 degrees, 225 degrees, T), and (45 degrees, 270 degrees, T).
[0437] The selection process based on the difference between the angle (θ) of the direction of arrival of the direct sound and the angle (γ) of the direction of arrival of the reflected sound is explained.
[0438] For example, it can also be pre-made and set up as follows: Figure 16 The threshold data shown is obtained by arranging the angle difference (Φ) between the direction of arrival of the direct tone (θ) and the direction of arrival of the reflected tone (γ) and the time difference (T) as a two-dimensional index. In this case, the angle difference (Φ) and the time difference (T) are referenced in the selection process. Alternatively, the angle difference (Φ) between the angle of arrival of the direct tone (θ) and the angle of arrival of the reflected tone (γ) can be calculated in the selection process, and the calculated angle difference (Φ) can be used to determine the threshold.
[0439] Alternatively, a threshold data can be set to be used as an index for arranging the combination of the angle difference (Φ), the direction of arrival of the direct sound (θ), and the time difference (T), or the combination of the angle difference (Φ), the direction of arrival of the reflected sound (γ), and the time difference (T).
[0440] Alternatively, it can be set as follows: Figure 15 The threshold data shown is obtained by arranging the values of (θ), (γ), and (T) as a three-dimensional index.
[0441] [The fourth variation of the rendering department's actions] The processing performed by the analysis unit 1301, the determination unit 1302 and the reproduction unit 1303 described above can also be performed as pipeline processing as described in Patent Document 3.
[0442] Figure 24 This is a block diagram illustrating a configuration example for pipeline processing in the rendering unit 1300.
[0443] Figure 24 The rendering unit 1300 includes a reverberation processing unit 1311, an initial reflection processing unit 1312, a distance attenuation processing unit 1313, a determination unit 1314, a generation unit 1315, and a binocular processing unit 1316. These multiple components can also be... Figure 7 The rendering unit 1300 shown can be composed of multiple constituent elements, or it can be made of... Figure 5 It constitutes at least a portion of the multiple constituent elements of the sound signal processing device 1001 shown.
[0444] Pipeline processing refers to dividing the processing used to impart sound effects into multiple processes and executing these processes sequentially. Each process may perform signal processing on the audio signal or generate parameters used in signal processing.
[0445] The rendering unit 1300 can also perform reverberation processing, initial reflection processing, distance attenuation processing, and binaural processing as a pipeline process. However, these are just examples; pipeline processing may include other processing methods, or it may exclude some of them. For example, pipeline processing may also include diffraction processing and occlusion processing. Furthermore, reverberation processing, for example, can be omitted if it is not needed.
[0446] Furthermore, each process can be represented as a stage. Additionally, the results of each process, and the generated sound signals such as reflected sounds, can be represented as rendering elements. The multiple stages in the pipeline processing and their order are not limited to... Figure 24 The example shown.
[0447] Here, the parameters used in the selection process (arrival path, arrival time, and volume ratio related to direct and reflected tones) can also be calculated in one of the multiple stages used to generate the render item. That is, the parameters used in the selection of reflected tones are calculated in a part of the pipeline processing used to generate the render item. Alternatively, not all stages may be performed by the rendering unit 1300. For example, some stages may be omitted, or they may be performed outside of the rendering unit 1300.
[0448] The reverberation processing, initial reflection processing, distance attenuation processing, selection processing, generation processing, and binaural processing that may be included as stages in pipeline processing are described. Within each stage, metadata contained in the input signal can also be parsed to calculate the parameters used in generating the reflected tones.
[0449] In reverberation processing, the reverberation processing unit 1311 generates parameters used in the generation of a sound signal representing reverberant sound. Reverberant sound refers to the sound that arrives at the listener as reverberation after the direct tone. As an example, reverberant sound is the sound that arrives at the listener after a relatively late stage (e.g., from the arrival of the direct tone to about one hundred and several tens of ms) following the arrival of the initial reflected sound, as described later. It undergoes more reflections (e.g., dozens of times) than the initial reflected sound.
[0450] The reverberation processing unit 1311 refers to the sound signal and spatial information contained in the input signal and calculates the reverberation sound using a pre-prepared function that is used to generate the reverberation sound.
[0451] The reverberation processing unit 1311 can also apply known reverberation generation methods to the sound signal contained in the input signal to generate reverberation. An example of a known reverberation generation method is the Schroeder method, but known reverberation generation methods are not limited to the Schroeder method. Furthermore, the reverberation processing unit 1311 uses the shape and acoustic characteristics of the sound reproduction space represented by spatial information in the application of known reverberation generation methods. Therefore, the reverberation processing unit 1311 can calculate the parameters used to generate the reverberation.
[0452] In the initial reflection processing, the initial reflection processing unit 1312 calculates parameters for generating the initial reflection tone based on spatial information. The initial reflection tone is the reflection tone that reaches the listener after more than one reflection in a relatively early stage (e.g., about tens of milliseconds from the arrival of the direct tone) after the direct tone reaches the listener from the sound source object.
[0453] The initial reflection processing unit 1312 calculates, for example, the path of the reflected sound from the sound source object to the listener via the reflection object, by referring to the sound signal and metadata. For example, the shape of the three-dimensional sound field (space), the size of the three-dimensional sound field, the position of the reflecting object such as the structure, and the reflectivity of the reflecting object can also be used in the path calculation.
[0454] Furthermore, the initial reflection processing unit 1312 may also calculate the path of the direct tone. This path information may also be used as a parameter for the initial reflection processing unit 1312 to generate the initial reflected tone, or as a parameter for the determination unit 1314 to select the reflected tone.
[0455] In distance attenuation processing, the distance attenuation processing unit 1313 calculates the volume of the direct tone and the reflected tone reaching the listener based on the path lengths of the direct tone and the reflected tone. The volume of the direct tone and the reflected tone reaching the listener is attenuated proportionally to the distance of the path to the listener (inversely proportional to the distance) relative to the volume of the sound source. Therefore, the distance attenuation processing unit 1313 can calculate the volume of the direct tone by dividing the volume of the sound source by the path length of the direct tone, and can calculate the volume of the reflected tone by dividing the volume of the sound source by the path length of the reflected tone.
[0456] In the selection process, the determination unit 1314 selects the object to be reflected based on parameters calculated prior to the selection process. A selection method of this disclosure may also be used in selecting the object to be reflected.
[0457] Selection processing can be performed on all reflected sounds, or, as described above, on only reflected sounds with high evaluation values. That is, reflected sounds with low evaluation values are automatically disqualified without any selection processing. For example, reflected sounds with very low volume can also be considered as having low evaluation values and are therefore disqualified.
[0458] Furthermore, for example, selection processing can be applied to all reflected sounds. Also, the evaluation value of the selected reflected sounds in the selection process can be determined, and reflected sounds with low evaluation values can be re-selected as not selected.
[0459] The selection and evaluation processes can be executed independently or in combination. When the selection and evaluation processes are executed in combination, one of the two processes can be executed first.
[0460] In the generation process, the generation unit 1315 generates direct tones and reflected tones. For example, the generation unit 1315 generates a direct tone based on the sound signal contained in the input signal, according to the arrival time and volume of the direct tone. Furthermore, regarding the reflected tone selected in the selection process, the generation unit 1315 generates a reflected tone based on the sound signal contained in the input signal, according to the arrival time and volume of the reflected tone.
[0461] In binaural processing, the binaural processing unit 1316 performs signal processing to make the direct sound signal perceived as sound arriving at the listener from the direction of the sound source object. Furthermore, the binaural processing unit 1316 performs signal processing to make the reflected sound selected by the determination unit 1314 perceived as sound arriving at the listener from the reflecting object.
[0462] For example, the binaural processing unit 1316 performs HRIR DB processing based on the position and orientation of the listener in the sound space, so that the sound reaches the listener from the position of the sound source object or the position of the obstacle object.
[0463] Additionally, HRIR (Head-Related Impulse Responses) describes the response characteristics when a single impulse is generated. Specifically, HRIR is the response characteristic obtained by transforming the head-related transfer function from its frequency domain representation to its time domain representation using a Fourier transform. This head-related transfer function represents the changes in sound produced by surrounding objects, including the auricle, head, and shoulders, as a transfer function. The HRIR DB is a database containing such information.
[0464] Furthermore, the position and orientation of the listener in the sound space can be, for example, the position and orientation of a virtual listener in a virtual sound space. Alternatively, the position and orientation of the virtual listener in the virtual sound space can change in accordance with the movement of the listener's head. Furthermore, the position and orientation of the virtual listener in the virtual sound space can also be determined based on information obtained from sensor 1405.
[0465] The program, spatial information, HRIR DB, threshold data or other parameters used in the above processing are obtained from the memory 1404 of the sound signal processing device 1001 or from outside the sound signal processing device 1001.
[0466] Furthermore, pipeline processing may also include other processing. Additionally, the rendering unit 1300 may include a processing unit (not shown) for performing other processing included in pipeline processing. For example, the rendering unit 1300 may also include a diffraction processing unit and an occlusion processing unit.
[0467] The diffraction processing unit performs processing to generate a sound signal representing a sound containing diffracted tones, which are caused by an obstacle object in a three-dimensional sound field (space) located between the listener and the sound source object. A diffracted tone is a sound that reaches the listener from the sound source object by bypassing an obstacle object when such an obstacle object exists between the sound source object and the listener.
[0468] The diffraction processing unit, for example, refers to the sound signal and metadata to calculate the path of the diffracted sound from the sound source object, bypassing the obstacle object, to reach the listener, and generates the diffracted sound based on the path. In the path calculation, the positions of the sound source object, the listener, and the obstacle object in the three-dimensional sound field (space), as well as the shape and size of the obstacle object, can also be used.
[0469] When a sound source object exists on the opposite side of an obstacle object, the occlusion processing unit generates a sound signal of the sound leaking from the sound source object through the obstacle object based on spatial information and information such as the material of the obstacle object.
[0470] [Example of a sound source object] In the above, the positional information assigned to the sound source object represents the position of the sound source object as a "point" in the virtual space. That is, in the above, the sound source is defined as a "point sound source".
[0471] On the other hand, a sound source in virtual space can also be defined as an object with length, size, and shape, that is, a non-point sound source that extends spatially. In this case, the distance between the listener and the sound source and the direction of sound arrival are uncertain. Therefore, reflected sounds caused by such sound sources do not need to be analyzed by the analysis unit 1301, or are limited to being selected by the determination unit 1302 regardless of the analysis result. As a result, sound quality degradation that may occur due to failure to select reflected sounds can be avoided.
[0472] Alternatively, a representative point, such as the object's center of gravity, can be determined, and the processing disclosed herein can be applied assuming that sound originates from that point. In this case, the threshold can also be adjusted based on the spatial extension information of the sound source.
[0473] Examples of direct and reflected sounds For example, a direct sound is a sound that is not reflected by a reflecting object, while a reflected sound is a sound that is reflected by a reflecting object. A direct sound can also be a sound that reaches the listener from the sound source without being reflected by a reflecting object, and a reflected sound can also be a sound that reaches the listener from the sound source after being reflected by a reflecting object.
[0474] Furthermore, direct tone and reflected tone are not limited to the sound that reaches the listener; they can also be the sound before it reaches the listener. For example, direct tone can also be the sound output from the sound source, or in other words, the sound of the sound source.
[0475] Figure 25 This is a diagram illustrating the transmission and diffraction of sound. For example... Figure 25 As shown, sometimes the direct sound does not reach the listener because of an obstacle object between the sound source and the listener. In this case, the sound emitted from the sound source, passing through the obstacle object, and reaching the listener can also be considered as the direct sound. Furthermore, the sound emitted from the sound source, diffracted through the obstacle object, and reaching the listener can also be considered as the reflected sound.
[0476] Furthermore, the two sounds compared in the selection process are not limited to the direct and reflected sounds of a sound emitted from a single sound source. For example, sound selection can also be performed by comparing two reflected sounds of a sound emitted from a single sound source. In this case, the direct sound in this disclosure can be replaced with the sound that arrives at the listener first, and the reflected sound in this disclosure can be replaced with the sound that arrives at the listener later.
[0477] [Example of bitstream construction] Bitstreams may contain, for example, audio signals and metadata. Audio signals are audio data representing sound, including information related to the frequency and intensity of the sound. Furthermore, metadata contains spatial information related to the space of the sound field, i.e., the sound space.
[0478] For example, spatial information is information relating to the space in which a listener is located when receiving sound based on a sound signal. Specifically, spatial information is information related to a predetermined location (location position) used to position a sound image at a specific location in sound space (e.g., a three-dimensional sound field), that is, information used to enable the listener to perceive sound arriving from a direction corresponding to that predetermined location. Spatial information may include, for example, information about the sound source and location information indicating the listener's position.
[0479] Sound source object information refers to the information about the sound source object that generates sound based on the sound signal. That is, sound source object information is information related to the object (sound source object) that reproduces the sound signal, and it is information related to a virtual sound source object configured in a virtual sound space. Here, the virtual sound space can also correspond to the real space where the object generating the sound is configured, and the sound source object in the virtual sound space can also correspond to the object generating sound in the real space.
[0480] Sound source object information can also represent the location of a sound source object configured in the sound space, the orientation of the sound source object, the directionality of the sound emitted by the sound source object, whether the sound source object is a living being, and whether the sound source object is a moving object. For example, a sound signal can be associated with one or more sound source objects represented by the sound source object information.
[0481] Bitstreams, for example, have a data structure consisting of metadata (control information) and sound signals.
[0482] The audio signal and metadata can be contained in a single bitstream or in multiple separate bitstreams. Furthermore, the audio signal and metadata can be contained in a single file or in multiple separate files.
[0483] Bitstreams can exist either per audio source or per playback time. When bitstreams exist per playback time, multiple bitstreams can be processed in parallel simultaneously.
[0484] Metadata can be assigned to each bitstream individually, or it can be assigned to multiple bitstreams together as information to control them. In this case, multiple bitstreams can also share the metadata. Alternatively, metadata can be assigned at each playback time.
[0485] In the presence of multiple bitstreams or multiple files, information indicating associated bitstreams or associated files may be included in more than one bitstream or more than one file. Alternatively, information indicating associated bitstreams or associated files may be included in each of the individual bitstreams or each of the individual files.
[0486] Here, associated bitstreams or associated files refer, for example, to bitstreams or files that may be used simultaneously during audio processing. Additionally, it may also include bitstreams or files that together contain information representing associated bitstreams or associated files.
[0487] Here, the information representing the associated bitstream or file can be, for example, an identifier representing the associated bitstream or file. Alternatively, the information representing the associated bitstream or file can be, for example, the filename, URL (Uniform Resource Locator), or URI (Uniform Resource Identifier).
[0488] In this case, the acquisition unit can also determine and acquire the associated bitstream or associated file based on information representing the associated bitstream or associated file. Alternatively, information representing the associated bitstream or associated file can be included in the bitstream or file, and also in other bitstreams or other files.
[0489] Here, the file containing information representing the associated bitstream or associated file can also be a control file such as a declaration file for content distribution.
[0490] In addition, all or part of the metadata can be obtained from outside the audio signal bitstream. For example, metadata for controlling the audio and metadata for controlling the video can be obtained from outside the bitstream, or metadata for both can be obtained from outside the bitstream.
[0491] Furthermore, metadata for controlling the image may also be included in the bitstream acquired by the stereo sound reproduction system 1000. In this case, the stereo sound reproduction system 1000 may also output the metadata for controlling the image to a display device that displays the image or a stereo image reproduction device that reproduces the stereo image.
[0492] [Example of information contained in metadata] Metadata can also be information used in the description of a scene represented by a sound space. Here, a scene is a term that refers to the collection of all elements of a sound space, including three-dimensional images and sound events, modeled by a sound signal reproduction system using metadata.
[0493] That is, metadata includes not only information used to control audio processing, but also information used to control video processing. Metadata can contain only one of the information used to control audio processing or the information used to control video processing, or it can contain both.
[0494] The stereo sound reproduction system 1000 processes sound signals using metadata contained in the bitstream and interactive listener location information acquired through appending, to generate virtual sound effects. These sound effects can include initial reflection processing, obstacle removal, diffraction processing, masking, and reverberation processing, as well as other sound processing using metadata. For example, sound effects such as distance attenuation, localization, or Doppler effects can be added.
[0495] In addition, information can be attached to the metadata to toggle the on / off of all or some of the additional sound effects, or priority information for multiple processing of sound effects.
[0496] In addition, as an example, metadata includes information related to the sound space, including sound source objects and obstacle objects, and information related to the positioning location used to locate the sound image in a specified position within the sound space (i.e., to make the listener perceive the sound coming from a specified direction).
[0497] Here, an obstacle object is an object that may affect the listener's perception of sound by blocking or reflecting it before the sound emitted by the sound source reaches the listener. Besides stationary objects, obstacle objects can also include moving bodies such as animals or machines. Animals can also be people.
[0498] Furthermore, when multiple sound source objects exist in the sound space, for any given sound source object, the other sound source objects may become obstacle objects. That is, objects that do not emit sound, such as building materials or inanimate objects (i.e., non-sound-emitting objects), as well as sound source objects that emit sound, can all become obstacle objects.
[0499] The metadata contains all or part of the information representing the shape of the sound space, the shape and location of obstacle objects in the sound space, the shape and location of sound source objects in the sound space, and the location and orientation of the listener in the sound space.
[0500] The sound space can be either enclosed or open. Furthermore, the metadata can also include information about the reflectivity of obstacles within the sound space that can reflect sound. For example, the floor, walls, or ceiling that form the boundary of the sound space can also be considered obstacles.
[0501] Reflectivity is the energy ratio of reflected sound to incident sound, and it can be set for each frequency band of the sound. Of course, reflectivity can also be set uniformly regardless of the frequency band of the sound. In addition, when the sound space is an open space, parameters such as attenuation rate, diffraction tone, and initial reflection tone can be set uniformly, for example.
[0502] Metadata can also include information beyond reflectivity as parameters relating to obstacle or sound source objects. For example, metadata can also include information about the material of an object as parameters relating to both the sound source and the non-sound-producing object. Specifically, metadata can also include information such as diffusivity, transmissivity, and sound absorption.
[0503] Information related to a sound source object may also include volume, radiation characteristics (directivity), reproduction conditions, the number and type of sound sources in an object, and information representing the sound source region within the object. Reproduction conditions may, for example, specify whether the sound is a continuously flowing sound or an event-triggered sound. The sound source region within an object can be set based on the relative position of the listener and the object, or it can be set using the object as a reference.
[0504] For example, when the sound source area is set according to the relative position of the listener and the object, from the listener's perspective, the listener can perceive sound A from the right side of the object and sound B from the left side of the object.
[0505] Furthermore, when using an object as a reference to define a sound source region, it is possible to fix which area of the object emits which sound. For example, when the listener views the object from the front, the listener can perceive high frequencies from the right side of the object and low frequencies from the left side. And when the listener views the object from the back, the listener can perceive low frequencies from the right side of the object and high frequencies from the left side.
[0506] Spatial metadata can also include the time up to the initial reflection, reverberation time, and the ratio of direct to diffuse sound. When the ratio of direct to diffuse sound is zero, the listener can perceive only the direct sound.
[0507] (Implementation Method 2) Hereinafter, Embodiment 2 will be described. The description will focus on the differences from Embodiment 1, and the description of the commonalities will be omitted or simplified.
[0508] [The Composition of the Rendering Department] First, the configuration of the rendering unit 2300 in this embodiment will be explained. Figure 26 This is a block diagram showing an example of the configuration of the rendering unit 2300 in this embodiment.
[0509] The rendering unit 2300 includes a resolution unit 2301, a determination unit 2302, and a reproduction unit 2303. Furthermore, as described above, the audio signal processing apparatus of this embodiment is an example of a decoding apparatus, which includes a decoder, and the decoder includes the rendering unit 2300. That is, it can be said that the audio signal processing apparatus of this embodiment includes a resolution unit 2301, a determination unit 2302, and a reproduction unit 2303. The rendering unit 2300 performs additional audio processing on the audio data contained in the input signal and outputs it.
[0510] Similar to Implementation 1, the input signal consists of, for example, spatial information, sensor information, and sound data. The spatial information also includes physical information such as the reflection coefficient, transmission coefficient, and diffraction coefficient of the non-sound-emitting object (obstacle object).
[0511] The analysis unit 2301 can perform all or part of the processing performed by the analysis unit 1301 in Embodiment 1. Furthermore, the analysis unit 2301 generates sound signals representing reflected sounds and sound signals representing direct sounds, and stores the generated sound signals representing reflected sounds and sound signals representing direct sounds.
[0512] Furthermore, in this embodiment, reflected sounds, as an example of indirect sounds, and direct sounds, as an example of defined sounds, are primarily used. That is, in this embodiment, reflected sounds can be used as indirect sounds, and direct sounds can be used as defined sounds. Furthermore, while reflected sounds and direct sounds are primarily used in the explanation here, the same treatment applies even when indirect sounds are used instead of reflected sounds, or when defined sounds are used instead of direct sounds. Additionally, as an example, an indirect sound is a reflected sound or a diffracted sound, etc., and a defined sound is a sound different from an indirect sound; as an example, it is a direct sound, a HOA sound, or a representative sound.
[0513] The parsing unit 2301 includes a propagation path detection unit 2301a and a memory 2301b.
[0514] The propagation path detection unit 2301a generates sound signals representing reflected sounds and sound signals representing direct sounds based on spatial information and sound data.
[0515] More specifically, the propagation path detection unit 2301a generates sound signals representing reflected sound and sound signals representing direct sound based on the location information of the sound source object, the location information of the non-sound-emitting object (obstacle object), the location information and physical information of the listener, and sound data contained in the spatial information.
[0516] That is, the propagation path detection unit 2301a generates a sound signal generated in the virtual space based on spatial information and sound data, assigns attribute information representing the attributes of the generated sound signal to the generated sound signal, and generates a sound signal containing attribute information. The attribute is information indicating whether the sound represented by the sound signal is a direct tone (defined tone) or a reflected tone (indirect tone).
[0517] A sound that reaches the listener's head directly from a sound source is a direct sound. A sound that reaches the listener's head after being output from a sound source and reflected or diffracted by a non-sound-producing object is an indirect sound (reflected sound or diffracted sound).
[0518] Furthermore, the direct sound associated with the indirect sound refers to a direct sound originating from the same sound source as the indirect sound. The indirect sound associated with the direct sound refers to an indirect sound originating from the same sound source as the direct sound.
[0519] A sound signal with the attribute of reflected tone (indirect tone) contains information representing the sound signal of the direct tone associated with that reflected tone (indirect tone).
[0520] The propagation path detection unit 2301a stores the generated sound signal in the memory 2301b. Figure 26 Sound signals A, B, C, and D stored in memory 2301b are shown, along with the properties of each sound signal. Furthermore, these sound signals can also be stored in a storage device other than memory 2301b.
[0521] Furthermore, similar to Embodiment 1, the propagation path detection unit 2301a can calculate values related to the path to the listening position, the time taken to reach the position, and the volume at the time of arrival for both direct and reflected sounds. Similarly, the propagation path detection unit 2301a can calculate information representing the relationship between direct and reflected sounds, such as values related to the time difference between the arrival of direct and reflected sounds (the time difference between direct and reflected sounds) and values related to the volume ratio of direct and reflected sounds at the listening position.
[0522] Sound signals with the attribute of reflected sound and sound signals with the attribute of direct sound can also contain information indicating the volume of the sound represented by the sound signal at the listening position. Sound signals with the attribute of reflected sound can also contain information indicating the volume (lr) when the reflected sound arrives, which is the volume of the reflected sound represented by the sound signal at the listening position. Sound signals with the attribute of direct sound can also contain information indicating the volume (ld) when the direct sound arrives, which is the volume of the direct sound represented by the sound signal at the listening position.
[0523] The determination unit 2302 can perform all or part of the processing performed by the determination unit 1302 in Embodiment 1. In addition, the determination unit 2302 determines whether the reproduction unit 2303 outputs (reproduces) an output signal based on the sound signal produced by the analysis unit 2301 (more specifically, the propagation path detection unit 2301a).
[0524] The determination unit 2302 has a classification unit 2302a, a first determination unit 2302b, and a second determination unit 2302c.
[0525] The classification unit 2302a acquires an audio signal containing attribute information generated by the propagation path detection unit 2301a and stored in the memory 2301b. The classification unit 2302a outputs the audio signal to the first determination unit 2302b or the second determination unit 2302c according to the attribute determined by the attribute information contained in the acquired audio signal.
[0526] If the attribute determined by the attribute information contained in the acquired sound signal is information representing a direct tone (standard tone), the classification unit 2302a outputs the sound signal to the first determination unit 2302b. If the attribute determined by the attribute information contained in the acquired sound signal is information representing a reflected tone (indirect tone), the classification unit 2302a outputs the sound signal to the second determination unit 2302c.
[0527] The first determination unit 2302b performs a first determination process to determine whether the sound signal output from the classification unit 2302a satisfies the first condition. If the sound signal satisfies the first condition, the first determination unit 2302b outputs the sound signal to the reproduction unit 2303.
[0528] Furthermore, the first determination unit 2302b does not perform the second determination process. That is, if the attribute determined by the attribute information contained in the acquired sound signal is information representing a specified tone, the second determination process is not performed. In this way, by not performing the second determination process, the amount of computation and computational load can be reduced.
[0529] The second determination unit 2302c performs a first determination process to determine whether the sound signal output from the classification unit 2302a satisfies the first condition, and a second determination process to determine whether it satisfies a second condition different from the first condition. If the sound signal satisfies both the first and second conditions, the second determination unit 2302c outputs the sound signal to the reproduction unit 2303.
[0530] The reproduction unit 2303 can perform all or part of the processing performed by the reproduction unit 1303 in Embodiment 1. Furthermore, the reproduction unit 2303 acquires the sound signal output from the determination unit 2302 and outputs an output signal based on the acquired sound signal.
[0531] The reproduction unit 2303 has a first reproduction unit 2303a and a second reproduction unit 2303b. The first reproduction unit 2303a outputs an output signal (first output signal) based on a sound signal with attributes different from an indirect tone. The second reproduction unit 2303b outputs an output signal (second output signal) based on a sound signal with attributes of an indirect tone.
[0532] That is, the first reproduction unit 2303a acquires the sound signal output from the first determination unit 2302b and outputs a first output signal. The second reproduction unit 2303b acquires the sound signal output from the second determination unit 2302c and outputs a second output signal.
[0533] The first reproduction unit 2303a generates and outputs a first output signal by performing binaural filtering on the acquired sound signal. The binaural filtering is implemented, for example, by processing the acquired sound signal using a head-related transfer function.
[0534] The second reproduction unit 2303b generates and outputs a second output signal by performing binaural filtering and diffusion filtering on the acquired sound signal. The diffusion filtering process, for example, improves the realism of the indirect tones represented by the acquired sound signal by diffusing the indirect tones it represents. Furthermore, the diffusion filtering process uses a filter that simulates the audible strength of the diffusion of sound represented by the acquired sound signal (i.e., simulates the audible strength of the diffusion of sound perceived by the listener). In the diffusion filtering process, a finite-pulse filter and / or an infinite-pulse filter are used.
[0535] In addition, the first reproduction unit 2303a outputs a first output signal without performing diffusion filtering on the acquired sound signal.
[0536] With the above-described structure in the reproduction unit 2303, since the sound signal with the attribute of indirect tone is processed by diffusion filtering and then output, the listener can hear indirect tone with higher sound quality. Furthermore, since the sound signal with the attribute of a specific tone is not processed by diffusion filtering, the computational load and computational complexity are further reduced.
[0537] Hereinafter, an example of the operation of the sound signal processing method performed by the sound signal processing apparatus (more specifically, the rendering unit 2300) of this embodiment will be described.
[0538] [Example of actions in the rendering department] Figure 27 This is a flowchart illustrating an example of the operation of the sound signal processing apparatus in this embodiment. Figure 27 The main focus is on the processing performed by the rendering unit 2300 of the sound signal processing apparatus of this embodiment.
[0539] First, the analysis unit 2301 performs analysis processing on the input signal (S501). More specifically, the propagation path detection unit 2301a of the analysis unit 2301 generates a sound signal representing reflected sound and a sound signal representing direct sound based on spatial information and sound data. The propagation path detection unit 2301a stores the generated sound signals in the memory 2301b.
[0540] The analysis unit 2301 analyzes the input signal and calculates the following values for the direct tone and the reflected tone: values for the path to the listening position, the time taken to reach the position, and the volume at the time of arrival; and values for the time difference between the direct tone and the reflected tone and the volume ratio between the direct tone and the reflected tone at the listening position.
[0541] First, the propagation path detection unit 2301a calculates the characteristics of the direct tone and the reflected tone of the produced sound signal. Specifically, it calculates the arrival time and volume of the direct tone and the reflected tone when they reach the listener (listening position). Furthermore, the method for calculating these arrival times and volume can be the method shown in Embodiment 1.
[0542] Next, the propagation path detection unit 2301a calculates the volume ratio (L) of the direct tone arrival volume (ld) to the reflected tone arrival volume (lr), and the time difference (T) between the direct tone and the reflected tone (the time difference (T) between their arrival times). The direct tone arrival volume (ld) refers to the volume of the direct tone, as an example of a defined tone, when it arrives at the listener's location in the virtual space, i.e., the listening position; in other words, it is the volume of the defined tone (direct tone) at the listening position. The reflected tone arrival volume (lr) refers to the volume of the reflected tone, as an example of an indirect tone, when it arrives at the listening position; in other words, it is the volume of the indirect tone (reflected tone) at the listening position. That is, the volume ratio (L) is the volume ratio between the direct tone and the reflected tone (indirect tone) at the listening position. Furthermore, the calculation method for these volume ratios (L) and the time difference (T) can be implemented using the method shown in Embodiment 1.
[0543] The determination unit 2302 determines the sound signal (S502). More specifically, the determination unit 2302 determines whether the sound signal generated by the propagation path detection unit 2301a is output by the reproduction unit 2303. That is, the determination unit 2302 performs determination processing (selection processing).
[0544] First, the classification unit 2302a acquires a sound signal containing attribute information generated by the propagation path detection unit 2301a and stored in the memory 2301b. The classification unit 2302a outputs a sound signal with the attribute of direct tone (defined tone) to the first determination unit 2302b, and outputs a sound signal with the attribute of reflected tone (indirect tone) to the second determination unit 2302c.
[0545] In addition, the classification unit 2302a also obtains the volume (ld) of the direct sound when it arrives, the volume (lr) of the reflected sound when it arrives, the volume ratio (L) and the time difference (T) calculated by the propagation path detection unit 2301a.
[0546] Next, the first determination unit 2302b performs a first determination process on the sound signal output from the classification unit 2302a. Additionally, the second determination unit 2302c performs both the first and second determination processes on the sound signal output from the classification unit 2302a. That is, if the attribute determined from the attribute information contained in the acquired sound signal is information representing reflected sound (indirect sound), the determination unit 2302 (the second determination unit 2302c) performs both the first and second determination processes.
[0547] Here, the first determination process and the second determination process will be explained. First, the second determination process will be explained. Furthermore, in this embodiment, the second determination process is performed only on sound signals whose attribute is reflected sound (indirect sound).
[0548] The second determination process is a process of determining whether the acquired sound signal satisfies the second condition. In the second determination process, if the volume ratio of the direct tone (prescribed tone) and the reflected tone (indirect tone) when they arrive at the listening position is above a second threshold determined based on the time difference between the direct tone and the reflected tone involved in the reflected tone, the acquired sound signal is determined to satisfy the second condition. That is, the second determination process is equivalent to the process performed by the determination unit 1302 in Embodiment 1. In addition, as described above, this time difference is calculated by the propagation path detection unit 2301a, and the second determination unit 2302c obtains the calculated time difference and performs the second determination process.
[0549] As mentioned above, reflected sounds and the direct sounds involved in reflected sounds are sounds from the same sound source.
[0550] The time difference between the direct tone and the reflected tone is as described in Embodiment 1, for example, the time difference between the arrival time of the direct tone (arrival time) and the arrival time of the reflected tone (arrival time), but is not limited thereto. The volume ratio of the direct tone to the reflected tone when the direct tone arrives at the listening position is equivalent to the volume ratio (L) of the direct tone when it arrives (ld) to the reflected tone when it arrives (lr) in Embodiment 1.
[0551] In addition, the volume ratio (L) is calculated using the same method as in Implementation 1.
[0552] The second threshold is a value determined based on the time difference between the direct tone and the reflected tone (indirect tone) related to the reflected tone (indirect tone). In other words, it is a value dependent on the time difference and is the value shown by the threshold data in Implementation 1. The threshold data, for example, is represented in a graph with the time difference between the direct tone and the reflected tone on the horizontal axis and the volume ratio of the direct tone and the reflected tone on the vertical axis, as the threshold (second threshold) for whether the reflected tone is perceived or not.
[0553] More specifically, the threshold data representing the second threshold is Figures 11-13 The data shown is as follows.
[0554] In the second determination process, if the volume ratio of the direct tone to the reflected tone in the acquired sound signal is greater than or equal to the second threshold, the acquired sound signal is determined to satisfy the second condition.
[0555] Next, the first determination process will be explained. In this embodiment, the first determination process is the processing of sound signals with the attribute of indirect tone (reflected tone) and sound signals with the attribute of representing information of a specified tone (direct tone).
[0556] The first determination process is to determine whether the acquired sound signal satisfies the first condition. In the first determination process, if the amplitude value of the acquired sound signal is greater than or equal to the first threshold, the acquired sound signal is determined to satisfy the first condition. That is, the amplitude value of the sound signal is equivalent to the volume of the sound (reflected sound (indirect sound) or direct sound (prescribed sound)) represented by the sound signal. Therefore, in the first determination process, if the volume of the sound represented by the acquired sound signal is greater than or equal to a certain threshold, the acquired sound signal is determined to satisfy the first condition.
[0557] The first threshold differs from the second threshold; it is a constant value that does not depend on the time difference between the direct and reflected (indirect) tones involved in the reflected sound. The first threshold relates to the volume of the acquired sound signal, in other words, it relates to the amplitude value. Furthermore, the first threshold represents the volume boundary that can be perceived by the listener; it is a threshold used to determine sounds with volumes lower than this threshold as sounds that are not reproduced. Figure 28 This is a graph representing the threshold data of the first threshold in this embodiment. For example, the first threshold is -70dB. Figure 28 The amplitude of the sound signal B shown is above the first threshold, therefore it is determined that the sound signal B satisfies the first condition. Additionally, Figure 28The amplitude value of the sound signal A shown is less than the first threshold, therefore it is determined that the sound signal A does not meet the first condition.
[0558] In addition, the first threshold can also be set (determined) by the administrator or listener of the virtual space, for example, using the threshold setting unit described above.
[0559] The first threshold used in the first determination section 2302b and the first threshold used in the second determination section 2302c can be the same value or different values.
[0560] As described above, the determination unit 2302 performs the first determination process and the second determination process.
[0561] Then, the first determination unit 2302b performs a first determination process on the sound signal output from the classification unit 2302a, and outputs the sound signal to the reproduction unit 2303 if the sound signal meets the first condition.
[0562] The second determination unit 2302c performs a first determination process to determine whether the sound signal output from the classification unit 2302a satisfies the first condition, and a second determination process to determine whether it satisfies a second condition different from the first condition. Note that in this embodiment, the second determination process is performed after the first determination process.
[0563] Therefore, firstly, the second determination unit 2302c performs a first determination process to determine whether the sound signal output from the classification unit 2302a satisfies the first condition. If the sound signal satisfies the first condition, the second determination process is performed on the sound signal. Then, if the sound signal satisfies the second condition, the second determination unit 2302c outputs the sound signal to the reproduction unit 2303.
[0564] Then, the reproduction unit 2303 acquires the sound signal output from the determination unit 2302 and outputs an output signal based on the sound signal (S503).
[0565] More specifically, the first reproduction unit 2303a acquires the sound signal output from the first determination unit 2302b, performs binaural filtering on the sound signal, and generates and outputs a first output signal. The second reproduction unit 2303b acquires the sound signal output from the second determination unit 2302c, performs binaural filtering and diffusion filtering on the sound signal, and generates and outputs a second output signal. That is, the reproduction unit 2303 (more specifically, the second reproduction unit 2303b) outputs an output signal (second output signal) based on the acquired sound signal when the sound signal satisfies the first condition and the second condition.
[0566] Furthermore, if the sound signal does not meet the first condition, the first determination unit 2302b will not output the sound signal to the reproduction unit 2303 (first reproduction unit 2303a). In this case, the first reproduction unit 2303a will not output an output signal based on the sound signal, thus reducing the computational load.
[0567] Furthermore, the second determination unit 2302c does not output the sound signal to the reproduction unit 2303 (second reproduction unit 2303b) if the sound signal does not meet the first condition, and it also does not output the sound signal to the reproduction unit 2303 (second reproduction unit 2303b) if the sound signal does not meet the second condition. In this case, the second reproduction unit 2303b does not output an output signal based on the sound signal, thus reducing the computational load and computational complexity.
[0568] Thus, the sound signal processing method of this embodiment is a sound signal processing method executed by a sound signal processing apparatus (rendering unit 2300), and includes an acquisition step, a determination step, and a reproduction step.
[0569] The acquisition step acquires a sound signal containing attribute information that determines the properties of the sound signal. In the determination step, if the attribute determined by the attribute information contained in the acquired sound signal is information representing an indirect tone, a first determination process is performed to determine whether the acquired sound signal satisfies a first condition, and a second determination process is performed to determine whether the acquired sound signal satisfies a second condition different from the first condition. In the reproduction step, if the acquired sound signal satisfies both the first and second conditions, an output signal based on the acquired sound signal is output.
[0570] That is, a first determination process and a second determination process are performed on a sound signal with the attribute of indirect tone (reflected tone), and an output signal based on the sound signal obtained under the condition that the sound signal satisfies the first and second conditions is output. In other words, it appropriately determines whether to output an output signal based on the sound signal with the attribute of indirect tone (reflected tone). When no output signal is output, the computational complexity and computational load are reduced. That is, a sound signal processing method that can appropriately reduce the computational complexity and computational load is implemented.
[0571] Furthermore, in this embodiment, in the first determination process, if the amplitude value of the acquired sound signal is above a first threshold, it is determined that the acquired sound signal satisfies the first condition. In the second determination process, if the volume ratio of the direct tone and the indirect tone when the indirect tone arrives at the listener's location (i.e., the listening position) is above a second threshold determined based on the arrival time difference between the direct tone and the indirect tone, it is determined that the acquired sound signal satisfies the second condition.
[0572] Therefore, in the first determination process, if the amplitude value is above a first threshold, the sound signal is determined to satisfy the first condition; in the second determination process, if the volume ratio is above a second threshold, the sound signal is determined to satisfy the second condition. An output signal based on the sound signal obtained under the condition of satisfying this second condition is output. That is, it more appropriately determines whether to output an output signal based on the sound signal. In other words, it enables a sound signal processing method that can more appropriately reduce computational complexity and computational load.
[0573] Furthermore, in this embodiment, a first determination unit 2302b and a second determination unit 2302c are provided, but the embodiment is not limited thereto. For example, the first determination unit 2302b may be omitted, and the second determination unit 2302c may be provided instead. In this case, the second determination unit 2302c acquires both a sound signal with the attribute of a specified tone (direct tone) and a sound signal with the attribute of an indirect tone (reflected tone), and performs a first determination process on both of them. When it is determined in the first determination process that the sound signal with the attribute of a specified tone (direct tone) satisfies the first condition, the sound signal is output to the first reproduction unit 2303a, and the first reproduction unit 2303a outputs a first output signal based on the output sound signal. When it is determined in the first determination process that the sound signal with the attribute of an indirect tone (reflected tone) satisfies the first condition, the sound signal is further subjected to a second determination process. When the sound signal is determined to meet the second condition in the second determination process, the sound signal is output to the second reproduction unit 2303b, and the second reproduction unit 2303b outputs a second output signal based on the output sound signal. It can be said that a rendering unit that does not have such a first determination unit 2302b but has a second determination unit 2302c performs essentially the same processing as the rendering unit 2300 of this embodiment.
[0574] (Implementation Method 3) Hereinafter, Embodiment 3 will be described. The description will focus on the differences from Embodiment 2, and the description of the commonalities will be omitted or simplified.
[0575] [The Structure of the Rendering Department] First, the configuration of the rendering unit 3300 in this embodiment will be explained. Figure 29 This is a block diagram showing an example of the configuration of the rendering unit 3300 in this embodiment.
[0576] The rendering unit 3300 includes a resolution unit 2301, a determination unit 3302, and a reproduction unit 3303.
[0577] Except for the fact that it has a second determination unit 3302c instead of a second determination unit 2302c, the determination unit 3302 has the same configuration as the determination unit 2302 in Embodiment 2.
[0578] Except for having a gain setting unit 3303c, the reproduction unit 3303 has the same configuration as the reproduction unit 2303 in Embodiment 2.
[0579] The gain setting unit 3303c sets (determines) the gain in the diffusion filtering process of the second reproduction unit 2303b. That is, the gain setting unit 3303c sets (determines) the amplification rate of the amplitude of the sound signal in the diffusion filtering process. In this embodiment, the second reproduction unit 2303b uses the gain determined by the gain setting unit 3303c to perform diffusion filtering.
[0580] For example, the gain setting unit 3303c obtains the gain set (determined) by the administrator or listener of the virtual space, and sets (determined) the obtained gain as the gain in the diffusion filtering process.
[0581] The gain setting unit 3303c outputs the determined gain to the determination unit 3302 (more specifically, the second determination unit 3302c).
[0582] In this embodiment, if the attribute determined by the attribute information contained in the acquired sound signal is an indirect tone, the classification unit 2302a also outputs the sound signal to the second determination unit 3302c.
[0583] The second determination unit 3302c performs a first determination process to determine whether the sound signal output from the classification unit 2302a satisfies the first condition. In this embodiment, the second determination unit 3302c does not perform a second determination process.
[0584] Before performing the first determination process, the second determination unit 3302c acquires the gain output by the gain setting unit 3303c. The second determination unit 3302c adds the acquired gain to the sound signal output from the classification unit 2302a, and performs the first determination process on the sound signal with the added gain. If the sound signal satisfies the first condition, the second determination unit 3302c outputs the sound signal to the reproduction unit 3303 (more specifically, the second reproduction unit 2303b).
[0585] (Implementation Method 4) Hereinafter, Embodiment 4 will be described. The description will focus on the differences from Embodiment 2, and the description of the commonalities will be omitted or simplified.
[0586] [The Structure of the Rendering Department] First, the configuration of the rendering unit 4300 in this embodiment will be explained. Figure 30 This is a block diagram showing an example of the configuration of the rendering unit 4300 in this embodiment.
[0587] The rendering unit 4300 includes a resolution unit 2301, a rendering pipeline unit 4304, a reproduction unit 4303, a first gain accumulation unit 4305, and a second gain accumulation unit 4306.
[0588] Furthermore, in this embodiment, reflected sounds, as an example of indirect sounds, and direct sounds, as an example of standard sounds, are mainly used for explanation. However, the same treatment is applied even when indirect sounds are used instead of reflected sounds and standard sounds are used instead of direct sounds.
[0589] In this embodiment, the parsing unit 2301 outputs the sound signal stored in the memory 2301b to the rendering pipeline unit 4304.
[0590] The rendering pipeline unit 4304 acquires the sound signal output from the parsing unit 2301. The rendering pipeline unit 4304 acquires, for example, a sound signal with the attribute of indirect tone (reflected tone) (i.e., a sound signal representing indirect tone (reflected tone)).
[0591] Furthermore, the analysis unit 2301 outputs the reference volume contained in the spatial information of the input signal, more specifically, the reference volume of the indirect tone (reflected tone) represented by the sound signal, to the first gain accumulation unit 4305, and the first gain accumulation unit 4305 acquires the reference volume. The rendering pipeline unit 4304 performs one or more first processes (here, multiple first processes), a first determination process, a second determination process, and one or more second processes (here, multiple second processes) on the acquired sound signal.
[0592] The rendering pipeline 4304 has one or more first processing units 4304a (here, multiple first processing units 4304a), a decision unit 4302, and one or more second processing units 4304b (here, multiple second processing units 4304b).
[0593] The plurality of first processing units 4304a include first processing unit 4304a1, first processing unit 4304a2, first processing unit 4304a3, and first processing unit 4304a4. Each of the plurality of first processing units 4304a performs a first processing on the acquired sound signal. In this embodiment, the first processing is the process of determining the amount by which the amplitude of the sound signal is amplified, after processing the sound signal based on its physical characteristics. However, the first processing is not limited to this and may not be the process of determining the amount by which the amplitude of the sound signal is amplified.
[0594] The first processes performed by the multiple first processing units 4304a can also be different from each other. As an example, the first processing units 4304a1 to 4304a4 each perform the following first processes.
[0595] When the first processing unit 4304a1 has processed the acquired sound signal as described above by the initial reflection processing unit 1312, it performs a first processing step to determine the amount of increase or decrease in the amplitude of the sound signal.
[0596] When the first processing unit 4304a2 has performed the aforementioned diffraction processing unit processing on the acquired sound signal, it performs a first processing step to determine the amount of increase or decrease in the amplitude amplification of the sound signal.
[0597] When the first processing unit 4304a3 performs the aforementioned distance attenuation processing unit 1313 processing on the acquired sound signal, it performs a first processing step to determine the amount of increase or decrease in the amplitude amplification of the sound signal.
[0598] When the first processing unit 4304a4 has processed the acquired sound signal by the reverberation processing unit 1311 as described above, it performs a first processing step to determine the amount of increase or decrease in the amplitude of the sound signal.
[0599] In addition, as another example, the first processing unit 4304a may also perform the first processing to determine the amount of increase or decrease in the amplitude amplification of the sound signal when the acquired sound signal has been processed by transmission, direction, sound image localization or diffusion.
[0600] The first processing units 4304a1 to 4304a4 perform the first processing described above, and output the determined increase or decrease amount to the first gain accumulation unit 4305.
[0601] In this embodiment, after the first processing is performed by the multiple first processing units 4304a, the obtained sound signal is output to the determination unit 4302 for processing.
[0602] The determination unit 4302 performs the first determination process and the second determination process. Furthermore, the first determination process and the second determination process of this embodiment will be explained again after the first gain accumulation unit 4305 and the second gain accumulation unit 4306 have been described.
[0603] The plurality of second processing units 4304b includes second processing unit 4304b1 and second processing unit 4304b2. Each of the plurality of second processing units 4304b performs a second processing on the acquired sound signal. The second processing determines the amount by which the amplitude of the sound signal is amplified, after processing the sound signal based on the listener's auditory perception characteristics. Furthermore, the second processing can also be a sound quality adjustment function based on the listener's preferences or convenience, and is a processing that accompanies the increase or decrease of the signal amplitude. However, the second processing is not limited to this and may not be a process that determines the amount by which the amplitude of the sound signal is amplified.
[0604] The second processes performed by the multiple second processing units 4304b can also be different from each other. As an example, the second processing units 4304b1 and 4304b2 each perform the following second processes.
[0605] When the acquired sound signal has undergone diffusion filtering, the second processing unit 4304b1 performs a second processing step to determine the amount by which the amplitude of the sound signal is amplified. The diffusion filtering process, for example, improves the realism of the indirect tones by diffusing the acquired sound signal to enhance the indirect tones, and uses finite pulse filters and / or infinite pulse filters.
[0606] In the case of performing the aforementioned binaural filtering on the acquired sound signal, the second processing unit 4304b2 performs a second processing step to determine the amount of increase or decrease in the amplitude amplification of the sound signal.
[0607] Here, use Figures 31-34 The second processing performed by the second processing unit 4304b2 will be described. Here, the head-related transfer function used as an example for binaural filtering processing will be used for explanation.
[0608] Figures 31-34 These are graphs illustrating the energy of the head-related transfer function in this embodiment. Figures 31-34 The listener's listening position P is shown. Additionally, in Figures 31-34 The text shows the front, back, left, right, above, and below from the listener's perspective.
[0609] exist Figure 31 The diagram schematically illustrates a sphere centered at the listening position P, within which circles are shown along the listener's horizontal, sagittal, and coronal planes, respectively. Figure 32 In the image, a circle along the listener's horizontal plane is shown in thick lines. Figure 33 In the image, a circle along the sagittal plane of the listener is shown in thick lines. Figure 34In the image, a circle along the coronal plane of the listener is shown in thick lines.
[0610] exist Figures 31-34 In each of these, the direction of maximum energy of the head-related transfer function is indicated by a circle with a dot, and the direction of minimum energy of the head-related transfer function is indicated by a hollow circle. Furthermore, in... Figures 31-34 In each of the above, the maximum and minimum energy values of the head-related transfer function are shown.
[0611] like Figures 31-34 As shown, the energy of the head-related transfer function differs for each positioning orientation. Figures 31-34 In the text, the energy of the head-related transfer function is converted to dB and displayed as the energy of the transfer function with a transfer characteristic of 1 being 0 dB.
[0612] exist Figures 31-34 An example of the energy of the head-related transfer function is shown, but this energy varies significantly depending on the head-related transfer function used. In the second processing performed by the second processing unit 4304b2, a process is performed to determine the increase or decrease amount based on the head-related transfer function used.
[0613] The second processing units 4304b1 and 4304b2 perform the second processing described above, and output the determined increase or decrease amount to the second gain accumulation unit 4306.
[0614] Furthermore, the first gain accumulation unit 4305 and the second gain accumulation unit 4306 will be described.
[0615] The first gain accumulation unit 4305 calculates the first increase / decrease. The first increase / decrease is a value obtained by accumulating the increase / decrease determined by each of the multiple first processing units (more specifically, each of the multiple first processing units 4304a), which is the increase / decrease amount that amplifies the amplitude of the acquired sound signal. More specifically, the first gain accumulation unit 4305 calculates the first increase / decrease by accumulating this increase / decrease over a reference volume. That is, the first increase / decrease is a value obtained by accumulating the increase / decrease determined by each of the multiple first processing units 4304a and a reference volume.
[0616] As described above, multiple first processing units 4304a determine multiple (in this case, four) increments / decreases for the acquired sound signal. Each of the multiple first processing units 4304a outputs its determined increment / decrease to a first gain accumulator 4305. The first gain accumulator 4305 acquires the output increments / decreases and calculates a first increment / decrease by accumulating the multiple (four) increments / decreases with a reference volume. The first gain accumulator 4305 outputs the calculated first increment / decrease to a second gain accumulator 4306, which acquires the output first increment / decrease.
[0617] The second gain accumulator 4306 calculates a second increase / decrease. The second increase / decrease is a value obtained by accumulating the increase / decrease determined by each of the multiple second processing units (more specifically, each of the multiple second processing units 4304b), whereby the increase / decrease is the amount by which the amplitude of the acquired sound signal is amplified. More specifically, the second gain accumulator 4306 accumulates this increase / decrease, and further accumulates the first increase / decrease output from the first gain accumulator 4305 to calculate the second increase / decrease. That is, the second increase / decrease is a value obtained by accumulating the increase / decrease determined by each of the multiple second processing units 4304b and the first increase / decrease.
[0618] As described above, the multiple second processing units 4304b determine multiple (in this case, two) increments / decreases for the acquired sound signal. Each of the multiple second processing units 4304b outputs a parameter representing the determined increment / decrease to the second gain accumulation unit 4306. The second gain accumulation unit 4306 acquires the output multiple parameters and calculates the second increment / decrease by accumulating the increment / decrease indicated by the multiple parameters with the first increment / decrease.
[0619] Alternatively, the multiple second processing units 4304b may each output an increase or decrease amount determined without using the above parameters to the second gain accumulation unit 4306. In this case, the second gain accumulation unit 4306 acquires the multiple increase or decrease amounts output, and calculates the second increase or decrease amount by accumulating the multiple (two) increase or decrease amounts with the first increase or decrease amount.
[0620] Furthermore, as described above, the second processing can also be a sound quality adjustment function performed through the listener's selection, accompanied by an increase or decrease in the signal amplitude. For example, the second increase or decrease amount can also include an increase or decrease in amplitude based on the sound quality adjustment function, set by the listener's selection in the virtual space. In this case, the listener can, for example, make a selection by operating the operation receiving unit, which accepts the selection, and the second gain accumulation unit 4306 calculates the second increase or decrease amount based on the selection accepted by the operation receiving unit. The listener can, for example, listen to a sound with a preferred sound quality by making a selection that they like. In addition, the second increase or decrease amount can also include an increase or decrease in amplitude based on the sound quality adjustment function, set by the administrator of the virtual space instead of the listener in the virtual space.
[0621] The first and second determination processes of this embodiment will be explained again.
[0622] First, let's explain the first judgment process.
[0623] The determination unit 4302 performs a first determination process on the sound signal obtained by the rendering pipeline unit 4304. In the first determination process, if the first increment / decrement (G1) calculated by the first gain accumulation unit 4305 is greater than or equal to the first threshold (T1), it is determined that the obtained sound signal satisfies the first condition.
[0624] The first threshold in this embodiment is a fixed value, which is related to the volume of the sound signal, or in other words, related to the amplitude value. That is, the first threshold in this embodiment is the same as the first threshold described in Embodiment 2.
[0625] Furthermore, in the first determination process of Embodiment 2, if the amplitude value of the acquired sound signal is greater than or equal to a first threshold, it is determined that the acquired sound signal satisfies the first condition. That is, the first determination process of this embodiment is the same as the first determination process of Embodiment 2, except that the amplitude value of the acquired sound signal is changed to a first increment / decrement (G1).
[0626] Furthermore, the second determination process will be explained.
[0627] The determination unit 4302 performs a second determination process on the sound signal acquired by the rendering pipeline unit 4304. In the second determination process, if the value corresponding to the second increment / decrement calculated by the second gain accumulation unit 4306 is greater than or equal to the second threshold, the acquired sound signal is determined to satisfy the second condition.
[0628] The second threshold in this embodiment is a value different from the first threshold, and is equivalent to the second threshold described in Embodiment 2. That is, the second threshold is a value determined based on the arrival time difference between the direct tone involved in the indirect tone (reflected tone) shown by the acquired sound signal and the indirect tone (reflected tone).
[0629] The second increment / decrement calculated by the second gain accumulator 4306 is equivalent to the volume (Ir) when the reflected sound arrives as shown in Embodiments 1 and 2, that is, it represents the volume when the reflected sound shown by the acquired sound signal arrives at the listening position.
[0630] The value corresponding to the second increase / decrease is the ratio (volume ratio) of the volume of the direct tone involved in the indirect tone (reflected tone) shown by the acquired sound signal to the second increase / decrease calculated by the second gain accumulation unit 4306. That is, the value corresponding to the second increase / decrease is equivalent to the ratio of the volume (ld) when the direct tone arrives to the volume (lr) when the reflected tone arrives, as shown in Embodiment 2, i.e., the volume ratio (L).
[0631] Furthermore, in the second determination process of Embodiment 2, when the volume ratio (L) of the direct tone to the indirect tone at the listening position is a second threshold or higher, it is determined that the acquired sound signal satisfies the second condition. That is, except that the volume ratio of the direct tone to the indirect tone changes to a value corresponding to the second increment or decrement, the second determination process of this embodiment is the same as the second determination process of Embodiment 2.
[0632] Furthermore, the second increase / decrease amount and the corresponding value of the second increase / decrease amount in this embodiment can be calculated using the same method as the volume (Ir) and volume ratio (L) when the reflected sound arrives as described in Embodiment 2.
[0633] In this embodiment, if the first determination process determines that the acquired sound signal does not meet the first condition, the second determination process is not performed. Conversely, if the first determination process determines that the acquired sound signal meets the first condition, the second determination process is performed.
[0634] Note that, as described above, the second increase / decrease is calculated before the second determination process. The second processing performed by the second processing unit 4304b utilizes processing based on the listener's auditory perception characteristics, and therefore determines the increase / decrease independently of the sound signal (i.e., statically). Therefore, the second increase / decrease can be calculated before the second determination process.
[0635] Then, if the second determination process determines that the acquired sound signal meets the second condition, the determination unit 4302 outputs the acquired sound signal to a plurality of second processing units 4304b, and the rendering pipeline unit 4304 outputs the acquired sound signal to the reproduction unit 4303. In this case, the rendering pipeline unit 4304 may also output the second increment / decrement calculated by the second gain accumulation unit 4306 to the reproduction unit 4303. Then, processing in the reproduction unit 4303 is performed.
[0636] The reproduction unit 4303 acquires the audio signal output from the rendering pipeline unit 4304 and the second increment / decrement. The reproduction unit 4303 outputs an output signal based on the acquired audio signal. More specifically, the reproduction unit 4303 may also multiply the acquired audio signal by the acquired second increment / decrement to generate and output the output signal. Thus, in this embodiment, the reproduction unit 4303 outputs the output signal when both the first condition and the second condition are satisfied.
[0637] Hereinafter, an example of the operation of the sound signal processing method performed by the sound signal processing apparatus (more specifically, the rendering unit 4300) of this embodiment will be described.
[0638] [Example of actions in the rendering department] Figure 35This is a flowchart illustrating an example of the operation of the sound signal processing apparatus in this embodiment. Figure 35 The main focus is on the processing performed by the rendering unit 4300 included in the sound signal processing apparatus of this embodiment.
[0639] First, the analysis unit 2301 performs analysis processing on the input signal (S601).
[0640] Furthermore, the rendering pipeline unit 4304 acquires the sound signal output from the analysis unit 2301 (S602). Note that the rendering pipeline unit 4304 acquires, for example, a sound signal representing an indirect tone.
[0641] Then, the multiple first processing units 4304a each determine the amount by which the amplitude of the acquired sound signal is amplified (S603). The multiple first processing units 4304a each output the determined amount of amplification to the first gain accumulation unit 4305. In addition, the resolution unit 2301 outputs the reference volume of the indirect tone (reflected tone) to the first gain accumulation unit 4305.
[0642] The first gain accumulator 4305 acquires multiple (4) increment / decrement values and a reference volume, and accumulates the acquired increment / decrement values and the acquired reference volume to calculate a first increment / decrement value (S604). Then, the first gain accumulator 4305 outputs the calculated first increment / decrement value to the determination unit 4302. In addition, the first gain accumulator 4305 outputs the calculated first increment / decrement value to the second gain accumulator 4306.
[0643] Next, the multiple second processing units 4304b each determine the amount by which the amplitude of the acquired sound signal is amplified (S605). The multiple second processing units 4304b each output the determined amount of amplification to the second gain accumulation unit 4306.
[0644] The second gain accumulation unit 4306 acquires the multiple (two) increments and decrements output and the calculated first increment and decrement, and accumulates the acquired multiple increments and decrements and the acquired first increment and decrement to calculate the second increment and decrement (S606). Then, the second gain accumulation unit 4306 outputs the calculated second increment and decrement to the determination unit 4302.
[0645] The determination unit 4302 performs the first determination process (S607). That is, the determination unit 4302 determines whether the acquired sound signal meets the first condition by determining whether the calculated first increase or decrease amount is greater than or equal to the first threshold.
[0646] The determination unit 4302 performs a second determination process (S608). That is, the determination unit 4302 determines whether the acquired sound signal satisfies the second condition by determining whether the value corresponding to the calculated second increase or decrease is above the second threshold.
[0647] Furthermore, in this embodiment, if the acquired sound signal satisfies the first condition, the second determination process is performed.
[0648] Furthermore, if the acquired sound signal satisfies the second condition, the reproduction unit 4303 outputs an output signal based on the acquired sound signal (S609).
[0649] Furthermore, if the sound signal does not meet the first condition, the sound signal is not output to the reproduction unit 4303; if the sound signal does not meet the second condition, the sound signal is not output to the reproduction unit 4303. In this case, the reproduction unit 4303 does not output an output signal based on the sound signal, thus reducing the computational load.
[0650] Note that in this embodiment, the rendering pipeline 4304 may have the classification section 2302a described in Embodiment 2. The classification section 2302a may be provided, for example, between a plurality of first processing sections 4304a and the determination section 4302, but is not limited thereto.
[0651] As described above, the audio signal processing method of this embodiment is an audio signal processing method executed by an audio signal processing apparatus, including an acquisition step, a rendering pipeline step, a reproduction step, a first gain accumulation step, and a second gain accumulation step.
[0652] In the acquisition step, a sound signal representing an indirect tone is acquired. The rendering pipeline step performs one or more first processing, first decision processing, second decision processing, and one or more second processing different from the first processing on the acquired sound signal.
[0653] The first gain accumulation step includes: a reproduction step, outputting an output signal based on the acquired sound signal; and a step of accumulating and calculating a first increase / decrease amount determined by one or more first processes, wherein the increase / decrease amount is the increase / decrease amount that amplifies the amplitude of the acquired sound signal. The second gain accumulation step accumulates and calculates a second increase / decrease amount determined by one or more second processes, wherein the increase / decrease amount is the increase / decrease amount that amplifies the amplitude of the acquired sound signal. In the first determination process, if the calculated first increase / decrease amount is above a first threshold, the acquired sound signal is determined to satisfy the first condition. If the acquired sound signal is determined not to satisfy the first condition, the second determination process is not performed. If the acquired sound signal is determined to satisfy the first condition, the second determination process is performed. In the second determination process, if the value corresponding to the calculated second increase / decrease amount is above a second threshold that is different from the first threshold, the acquired sound signal is determined to satisfy the second condition, and if the acquired sound signal is determined to satisfy the second condition, the reproduction step is performed.
[0654] Therefore, the sound signal representing an indirect tone is subjected to a first determination process and a second determination process, and an output signal based on the sound signal obtained under the condition that the sound signal satisfies the first and second conditions is output. That is, it is appropriately determined whether to output an output signal based on the sound signal representing the indirect tone. When no output signal is output, the computational complexity and computational load are reduced. That is, a sound signal processing method that can appropriately reduce the computational complexity and computational load can be implemented.
[0655] In this embodiment, the second increment / decrement calculated in the second gain accumulation step includes the increment / decrement applied when performing diffusion filtering to improve the realism of the indirect tone by diffusing the indirect tone. This increment / decrement amplifies the amplitude of the acquired sound signal. The value is the ratio of the volume of the direct tone involved in the indirect tone to the second increment / decrement calculated in the second gain accumulation step, and the second threshold is determined based on the arrival time difference between the direct tone and the indirect tone.
[0656] Therefore, in the second determination process, when the aforementioned ratio is greater than or equal to the second threshold, it is determined that the sound signal representing the indirect tone satisfies the second condition, and an output signal based on the sound signal obtained under the condition of satisfying such the second condition is output. That is, it more appropriately determines whether to output an output signal based on the sound signal representing the indirect tone. In other words, it is possible to realize a sound signal processing method that can more appropriately reduce the computational load and computational complexity.
[0657] In this embodiment, the second increment / decrease calculated in the second gain accumulation step includes an increment / decrease in the amplitude of the sound quality adjustment function, set by the listener's selection. The above value is the ratio of the volume of the direct tone involved in the indirect tone to the second increment / decrease calculated in the second gain accumulation step, and the second threshold is determined based on the arrival time difference between the direct tone and the indirect tone.
[0658] Thus, listeners can listen to sounds with their preferred sound quality by making a selection that aligns with their own preferences.
[0659] In this embodiment, the first threshold is a value related to volume.
[0660] Therefore, it is possible to implement a sound signal processing method that uses a volume-related value as the first threshold.
[0661] In this embodiment, in the second gain accumulation step, parameters representing the increase or decrease amount determined by one or more second processes are obtained, and the second increase or decrease amount is calculated based on the obtained multiple parameters. The increase or decrease amount is the increase or decrease amount that amplifies the amplitude of the obtained sound signal.
[0662] Therefore, a second increment or decrement corresponding to the above parameters is calculated, thereby enabling a more appropriate determination of whether to output an output signal based on an indirect tone. In other words, a sound signal processing method that can more appropriately reduce computational complexity and workload can be implemented.
[0663] (Implementation Method 5) Hereinafter, Embodiment 5 will be described. The description will focus on the differences from Embodiment 4, and the description of the commonalities will be omitted or simplified.
[0664] [The Structure of the Rendering Department] First, the configuration of the rendering unit 5300 in this embodiment will be explained. Figure 36 This is a block diagram illustrating an example of the configuration of the rendering unit 5300 in this embodiment.
[0665] Except for having a modification unit 5307 and an invalidation unit 5308, the rendering unit 5300 has the same configuration as the rendering unit 4300 in Embodiment 4.
[0666] The modification unit 5307 performs the process of setting (determining) the first threshold used in the first determination process. For example, the modification unit 5307 outputs an instruction for setting the first threshold to the determination unit 4302. The determination unit 4302 receives the output instruction and sets the first threshold according to the instruction. Therefore, in the first determination process, if the calculated first increment or decrement is greater than or equal to the set first threshold, it is determined that the acquired sound signal satisfies the first condition.
[0667] The invalidation unit 5308 performs invalidation processing to invalidate the second determination process. By performing invalidation processing, the second determination process is invalidated and therefore not performed. For example, the invalidation unit 5308 outputs an instruction indicating invalidation processing to the determination unit 4302. The determination unit 4302, by receiving the output instruction, does not perform the second determination process.
[0668] In this embodiment, the case where the sound signal obtained in the first determination process is determined to satisfy the first condition and is invalidated will be described. In this case, the determination unit 4302 outputs the obtained sound signal to a plurality of second processing units 4304b, and the rendering pipeline unit 4304 outputs the obtained sound signal to the reproduction unit 4303. That is, in embodiment 4, the reproduction unit 4303 outputs an output signal when the output sound signal satisfies both the first and second conditions, but in this case, the reproduction unit 4303 outputs an output signal only when the output sound signal satisfies the first condition.
[0669] Furthermore, the first threshold set by the modification unit 5307 can have different values depending on whether invalidation processing is performed or not. More specifically, the first threshold set when invalidation processing is performed can be greater than the first threshold set when invalidation processing is not performed. Hereinafter, using... Figure 37 The setting of the first threshold and the impact of invalidation processing are explained.
[0670] Figure 37 This is a diagram showing a table used to illustrate the setting of the first threshold and the effect of the invalidation process in this embodiment.
[0671] exist Figure 37 The diagram illustrates the cases where the first threshold increases ("First Threshold: Large") and decreases ("First Threshold: Small") depending on the setting of the first threshold. Additionally, in... Figure 37 The document illustrates the cases where invalidation processing is performed (“Second determination processing is invalid”) and the cases where invalidation processing is not performed (“Second determination processing is valid”). Additionally, in... Figure 37 The document shows "processing complexity", "sound quality maintenance", and "sparse number".
[0672] "Processing complexity" refers to the complexity of processing in the rendering unit 5300. A low "processing complexity" (i.e., easy processing) is represented by "O", while a high "processing complexity" (i.e., difficult processing) is represented by "X".
[0673] "Sound quality maintenance" indicates whether the sound quality of the output signal from the reproduction unit 4303 is maintained. It is indicated by "O" when the sound quality is maintained and "X" when the sound quality is reduced.
[0674] "Simplified number" refers to the number of sound signals (more specifically, multiple sound signals) acquired by the rendering pipeline 4304 that are not output as output signals from the reproduction unit 4303. When the "simplified number" is high, that is, when the number of non-output sound signals is large and the computational load is reduced, it is represented as "0". When the "simplified number" is low, that is, when the number of non-output sound signals is small and the computational load is difficult to reduce, it is represented as "X".
[0675] like Figure 37 As shown, in the case of "First threshold: large" and "Second determination processing invalid", the "processing complexity" is O, the "sound quality maintenance" is X, and the "sparse number" is O. That is, in this case, the processing in the rendering unit 5300 is easy, and the amount of computation and computational load can be significantly reduced.
[0676] Similarly, when "the first threshold is large" and "the second judgment is valid", the "processing complexity" is X, the "sound quality maintenance" is X, and the "sparse number" is O. That is, in this case, the amount of computation and the computational load can be significantly reduced.
[0677] Similarly, in the case of "First threshold: small" and "Second judgment processing invalid", it means that "processing complexity" is O, "sound quality maintenance" is O, and "sparse number" is X. That is, in this case, high sound quality can be maintained.
[0678] Similarly, when "the first threshold is small" and "the second judgment is valid", the "processing complexity" is X, the "sound quality maintenance" is O, and the "sparse number" is O. That is, in this case, high sound quality can be maintained, and the amount of computation and computational load can be significantly reduced.
[0679] Furthermore, whether the modification unit 5307 performs the setting of the first threshold and whether the invalidation unit 5308 performs invalidation processing can also be set (determined) by the administrator or listener of the virtual space. Additionally, the space information may also include information indicating whether the modification unit 5307 performs the setting of the first threshold and whether the invalidation unit 5308 performs invalidation processing, and the decision is made according to this information. Thus, the administrator or listener of the virtual space can learn through trial and error. Figure 37 The most appropriate effect is selected from the various effects shown. Here, in order to effectively repeat the trial and error process, it is preferable that the setting areas are adjacent. That is, the storage area that stores the value of the first threshold set by the change unit 5307 and the storage area that stores the signal indicating the invalidation process can be adjacent areas. This not only makes it easy to visually grasp the setting value by being adjacent, but also creates the special effect of being able to set both values simultaneously (through a single memory access process) by being adjacent in the memory area. In other words, by linking and configuring the value of the first threshold and the signal indicating the invalidation process in the high-order bit field and low-order bit field of an area that can be accessed by a single address, the two data can be set simultaneously by writing the series of data configured in this way through a single memory access.
[0680] Furthermore, the processing of setting the first threshold by the modification unit 5307 and the invalidation processing performed by the invalidation unit 5308 can also be performed during the update of the aforementioned thread.
[0681] As described above, the sound signal processing method of this embodiment includes a step of setting a first threshold and a step of invalidation that invalidates the second determination process. In the first determination process, if the calculated first increment or decrement is greater than or equal to the set first threshold, it is determined that the acquired sound signal satisfies the first condition. When invalidation is performed, if it is determined that the acquired sound signal satisfies the first condition, a reproduction step is performed.
[0682] Therefore, by eliminating the second decision-making process through invalidation, the computational load and computational complexity required for the second decision-making process can be reduced. In other words, a sound signal processing method that can more appropriately reduce computational load and computational complexity can be implemented.
[0683] In the sound signal processing method of this embodiment, the first threshold set when invalidation processing is performed is greater than the first threshold set when invalidation processing is not performed.
[0684] Thus, a sound signal processing method can be implemented, which can set the size of the first threshold according to whether invalidation processing is performed.
[0685] (Replenish) Furthermore, the methods disclosed herein are not limited to specific implementation methods and can be implemented through various modifications.
[0686] For example, in an implementation, a process that is performed by a specific component may be performed by other components instead of that specific component. Furthermore, the order of multiple processes may be changed, or multiple processes may be performed in parallel.
[0687] Furthermore, the ordinal numbers 1, 2, etc., used in the description can be appropriately replaced, removed, or newly assigned. These ordinal numbers do not necessarily correspond to a meaningful order and can also be used for element identification.
[0688] Furthermore, for example, in comparisons of thresholds, "above the threshold" and "greater than the threshold" can be interchanged. Similarly, "below the threshold" and "less than the threshold" can be interchanged. Additionally, for example, "time" and "moment" can be interchanged.
[0689] Furthermore, in the process of selecting one or more target sounds from multiple sounds, if no sound meets the conditions, then none of the sounds may be selected as target sounds. That is, the process of selecting one or more target sounds from multiple sounds may also include the case of not selecting any target sounds.
[0690] Furthermore, such a representation, for example, of at least one of the first, second, and third elements, can correspond to the first, second, third elements, or any combination thereof.
[0691] Furthermore, for example, the embodiments described herein illustrate implementation as a sound signal processing apparatus, encoding apparatus, or decoding apparatus based on the methods known in this disclosure. However, the methods known in this disclosure are not limited thereto, and may also be implemented as software for performing sound signal processing, encoding, or decoding methods.
[0692] For example, the program for performing the aforementioned sound signal processing, encoding, or decoding methods can also be pre-stored in ROM. Furthermore, the CPU can also execute the actions according to this program.
[0693] Alternatively, the program used to perform the aforementioned sound signal processing, encoding, or decoding methods can be stored in a computer-readable recording medium. Furthermore, the computer can also record the program stored in the recording medium into its RAM and operate according to that program.
[0694] Furthermore, the aforementioned components can typically be implemented as integrated circuits, i.e., LSIs, which have input and output terminals. They can be formed as a single chip or as a chip incorporating all or some of the components of the implementation. Depending on the level of integration, LSIs can also be classified as ICs, system LSIs, super LSIs, or very large-scale LSIs.
[0695] Furthermore, it is not limited to LSIs; dedicated circuits or general-purpose processors can also be used. Additionally, FPGAs that can be programmed after LSI manufacturing, or reconfigurable processors capable of reconfiguring the connection or configuration of circuit units within the LSI, can also be used. Moreover, if advancements in semiconductor technology or other derived technologies lead to integrated circuit technologies that replace LSIs, then these technologies can certainly be used for the integration of constituent elements. This could include applications in biotechnology, among others.
[0696] Furthermore, an FPGA or CPU can download all or part of the software used to implement the sound signal processing, encoding, or decoding methods described in this disclosure via wireless or wired communication. Additionally, all or part of the software for updates can be downloaded via wireless or wired communication. Moreover, the digital signal processing described in this disclosure can be executed by storing the downloaded software in a memory using an FPGA or CPU and operating based on the stored software.
[0697] At this time, devices equipped with FPGAs or CPUs can also be connected to the signal processing device wirelessly or via wired connection, or connected to the signal processing server via a network. Furthermore, the device and the signal processing device or signal processing server can perform the audio signal processing methods, encoding methods, or decoding methods described in this disclosure.
[0698] For example, the audio signal processing apparatus, encoding apparatus, or decoding apparatus of this disclosure may also include an FPGA or a CPU. Furthermore, the audio signal processing apparatus, encoding apparatus, or decoding apparatus may also include an interface for obtaining software from an external source to operate the FPGA or CPU, and a memory for storing the obtained software. Moreover, the FPGA or CPU may execute the signal processing described in this disclosure based on the stored software operations.
[0699] Alternatively, the server may provide software related to the audio processing, encoding, or decoding processes disclosed herein. Furthermore, a terminal or device may operate as the audio signal processing, encoding, or decoding apparatus described in this disclosure by installing the software. Alternatively, the terminal or device may install the software by connecting to the server via a network.
[0700] Alternatively, a device other than the terminal or device may obtain the software installation data via a network connection to a server, and install the software on the terminal or device by providing the software installation data to the terminal or device through this other device. Another example of software could be VR software or AR software used to enable the terminal or device to execute the sound signal processing method described in the embodiments.
[0701] Furthermore, in the above embodiments, each component may be constructed using dedicated hardware, or implemented by executing software programs suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.
[0702] The above description describes apparatuses with one or more embodiments based on the embodiments, but the embodiments disclosed herein are not limited to the embodiments. As long as the spirit of this disclosure is not departed, various modifications that can be conceived by those skilled in the art to the embodiments can be made, and embodiments constructed by combining the constituent elements of different modifications are also included within the scope of one or more embodiments.
[0703] Industrial applicability This disclosure includes, for example, forms that can be used in a sound signal processing apparatus, an encoding apparatus, a decoding apparatus, or a terminal or device having any of these apparatuses.
[0704] Label Explanation 1000 Stereo Sound Reproduction System 1001 Sound signal processing device (audio processing device) 1002 Sound Prompt Device 1100, 1120, 1500 encoding devices Input data for 1101 and 1113 1102 Encoder 1103 Encoded Data 1104, 1114, 1404, 1503, 2301b memory 1110, 1130 Decoding Devices 1111 sound signal 1112, 1200, 1210 decoders 1121 Sending Department 1122 Send signal 1131 Receiving Department 1132 Received signal Spatial Information Management Department, 1201 & 1211 1202 Audio Data Decoder Rendering Department (models 1203, 1213, 1300, 2300, 3300, 4300, 5300) Analysis Departments 1301 and 2301 Judgment Departments 1302, 1314, 2302, 3302, and 4302 Reproduction section 1303, 2303, 3303, 4303 1304 Threshold Adjustment Section 1311 Reverb Processing Department 1312 Initial Reflection Processing Unit 1313 Distance Attenuation Processing Unit 1315 Production Department 1316 Binocular Processing Unit 1401 Speaker 1402 and 1501 processors 1403, 1502 Communication IF 1405 sensor 2301a Transmission Path Detection Department Classification section 2302a 2302b First Judgment Section 2302c, 3302c Second Judgment Section 2303a First Reproduction Section 2303b, Part 2 (Reproduction) 3303c Gain Setting Section 4304 Rendering Pipeline Department Processing Unit 1 (4304a, 4304a1, 4304a2, 4304a3, 4304a4) Processing Unit 2, 4304b, 4304b1, 4304b2 4305 First Gain Accumulator 4306 Second Gain Accumulator 5307 Change Department 5308 Invalidation Section
Claims
1. A sound signal processing method, executed by a sound signal processing device, wherein, include: The acquisition step involves acquiring a sound signal, wherein the sound signal contains attribute information that determines the properties of the sound signal; The determination step involves performing a first determination process to determine whether the acquired sound signal satisfies a first condition and a second determination process to determine whether the acquired sound signal satisfies a second condition that is different from the first condition, if the attribute determined by the attribute information contained in the acquired sound signal is information representing an indirect tone. as well as In the reproduction step, if the acquired sound signal satisfies the first condition and the second condition, an output signal based on the acquired sound signal is output.
2. The sound signal processing method according to claim 1, wherein, The indirect sound is a reflected sound.
3. The sound signal processing method according to claim 1 or 2, wherein, If the attribute determined by the attribute information contained in the acquired sound signal is information representing a prescribed tone different from the indirect tone, then... The second determination process is not performed in the determination step.
4. The sound signal processing method according to claim 3, wherein, The specified pitch is the direct pitch.
5. The sound signal processing method according to any one of claims 1 to 4, wherein, In the first determination process, if the amplitude value of the acquired sound signal is above a first threshold, it is determined that the acquired sound signal satisfies the first condition. In the second determination process, if the volume ratio of the direct tone to the indirect tone when the indirect tone and the direct tone involved in the indirect tone arrive at the listener's location (i.e., the listening location) is above a second threshold determined based on the arrival time difference between the direct tone and the indirect tone, the obtained sound signal is determined to satisfy the second condition.
6. The sound signal processing method according to any one of claims 1 to 5, wherein, The reproduction step includes: The first reproduction step involves outputting a first output signal based on the sound signal whose attribute is a predetermined tone different from the indirect tone; and The second reproduction step involves outputting a second output signal based on the sound signal whose attribute is the indirect tone. In the second reproduction step, the acquired sound signal is subjected to diffusion filtering to output the second output signal. The diffusion filtering process improves the realism of the indirect sound by diffusing the indirect sound. In the first reproduction step, the diffusion filtering process is not applied to the acquired sound signal, and the first output signal is output.
7. The sound signal processing method according to any one of claims 1 to 6, wherein, After the first determination process is performed, the second determination process is performed.
8. A sound signal processing method, executed by a sound signal processing device, wherein, include: The acquisition step involves acquiring the sound signal representing the indirect tone; The rendering pipeline step involves performing one or more first processing steps, one first determination processing step, one second determination processing step, and one or more second processing steps that are different from the one or more first processing steps on the acquired sound signal. The reproduction step outputs an output signal based on the acquired sound signal; The first gain accumulation step calculates the first gain / loss by accumulating the gain / loss determined by the one or more first processes, wherein the gain / loss determined by the one or more first processes is the gain / loss that amplifies the amplitude of the acquired sound signal. as well as The second gain accumulation step calculates a second gain / loss by accumulating the gain / loss amounts determined by the one or more second processes. These gain / loss amounts, determined by the one or more second processes, are the gain / loss amounts that amplify the amplitude of the acquired sound signal. In the first determination process, if the calculated first increase or decrease is greater than or equal to a first threshold, the acquired sound signal is determined to satisfy the first condition. If it is determined that the obtained sound signal does not meet the first condition, the second determination process will not be performed. If the acquired sound signal is determined to meet the first condition, the second determination process is performed. In the second determination process, if the value corresponding to the calculated second increase / decrease is above a second threshold that is different from the first threshold, it is determined that the obtained sound signal satisfies the second condition. If it is determined that the acquired sound signal satisfies the second condition, the reproduction step is performed.
9. The sound signal processing method according to claim 8, wherein, The second increment / decrement calculated in the second gain accumulation step includes the increment / decrement under the condition of diffusion filtering, which improves the realism of the indirect tone by diffusing it. The increment / decrement under the condition of diffusion filtering is the increment / decrement that amplifies the amplitude of the acquired sound signal. The value is the ratio of the volume of the direct tone involved in the indirect tone to the second increment / decrease calculated in the second gain accumulation step. The second threshold is determined based on the time difference between the arrival of the direct tone and the indirect tone.
10. The sound signal processing method according to claim 8, wherein, The second increase / decrease calculated in the second gain accumulation step includes an increase / decrease in amplitude based on the sound quality adjustment function, set by the listener's selection. The value is the ratio of the volume of the direct tone involved in the indirect tone to the second increment / decrease calculated in the second gain accumulation step. The second threshold is determined based on the time difference between the arrival of the direct tone and the indirect tone.
11. The sound signal processing method according to any one of claims 8 to 10, wherein, The first threshold is a value related to volume.
12. The sound signal processing method according to any one of claims 8 to 11, wherein, In the second gain accumulation step, The parameters representing the increments or decrements determined by the one or more second processes are obtained, whereby the increments or decrements determined by the one or more second processes are increments or decrements that amplify the amplitude of the obtained sound signal. The second increase or decrease is calculated based on the obtained parameters.
13. The sound signal processing method according to any one of claims 8 to 12, wherein, include: The change step involves setting the first threshold. as well as The invalidation step involves performing an invalidation process that invalidates the second determination process. In the first determination process, if the calculated first increase or decrease amount is greater than or equal to the set first threshold, it is determined that the obtained sound signal satisfies the first condition. If the invalidation process has been performed, and the obtained sound signal is determined to meet the first condition, the reproduction step is performed.
14. The sound signal processing method according to claim 13, wherein, The storage area for storing the value of the first threshold set in the change step and the storage area for storing the signal indicating the implementation of the invalidation process are adjacent areas.
15. The sound signal processing method according to claim 13, wherein, The first threshold set when the invalidation process has been performed is greater than the first threshold set when the invalidation process has not been performed.
16. A computer program for causing a computer to execute the sound signal processing method according to any one of claims 1 to 15.
17. A sound signal processing device, wherein, have: The acquisition unit acquires a sound signal, the sound signal containing attribute information that determines the attributes of the sound signal; The determination unit performs a first determination process to determine whether the acquired sound signal satisfies a first condition and a second determination process to determine whether the acquired sound signal satisfies a second condition different from the first condition when the attribute determined by the attribute information contained in the acquired sound signal is information representing an indirect tone. as well as The reproduction unit outputs an output signal based on the acquired sound signal if the acquired sound signal satisfies the first condition and the second condition.
Citation Information
Patent Citations
Signal processor
JP2019022049A
Apparatus and method for rendering a sound scene using pipeline stages
WO2021180938A1