Method for generating a reverberant audio signal - Patent Application 20070122997

The method efficiently generates reverberant audio signals that mimic virtual acoustic spaces by using symmetry groups and virtual object representations, addressing computational inefficiencies and adaptability issues in existing technologies, enabling real-time audio reproduction with accurate spatial characteristics.

JP7824277B2Active Publication Date: 2026-03-04リキッド·オキシゲン·(エルオーイクス)·ベー·フェー
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2023513215
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-08-28
Filing Date
2021-08-26
Publication Date
2026-03-04
Estimated Expiration
2041-08-26

AI Technical Summary

Technical Problem

Existing methods for generating reverberation in audio simulations are computationally expensive and unable to accurately reproduce the characteristics of an acoustic space, such as its shape, size, and materiality, while also lacking adaptability to real-time changes in virtual environments.

Method used

A method for generating reverberant audio signals using a virtual object representation, where virtual points are defined with symmetry groups and distances, allowing for efficient computation and adaptation of audio signals based on the virtual object's shape, size, and material properties, incorporating delay and feedback operations to mimic real acoustic spaces.

Benefits of technology

The method efficiently generates reverberant audio signals that accurately mimic the characteristics of virtual acoustic spaces, enabling real-time adaptability and computational efficiency, allowing for interactive and high-quality audio reproduction across various speaker configurations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007824277000028
    Figure 0007824277000028
  • Figure 0007824277000029
    Figure 0007824277000029
  • Figure 0007824277000030
    Figure 0007824277000030
Patent Text Reader

Abstract

A method for generating a reverberant audio signal associated with a virtual object is disclosed. The method includes storing a representation of the virtual object, the representation defining a plurality of virtual points constituting the virtual object, the virtual points having respective virtual positions relative to one another, and the virtual points belonging to the virtual point's symmetry group. Each symmetry group of the virtual point is associated with one or more sets of symmetry group distances. The one or more sets of symmetry group distances, each associated with a symmetry group, together form one or more further sets of distances. The method further includes receiving, storing, and / or generating an input audio signal and, for each virtual point, determining a virtual point audio signal component based on the input audio signal or a filtered version thereof. The method also includes combining the virtual point audio signal components to obtain a composite audio signal and, for each distinct distance in the one or more further sets of distances, determining one or more distance audio signals based on the composite audio signal. The method also includes determining a reverberant audio signal based on the one or more distance audio signals and the virtual point audio signal component.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to systems and methods for generating reverberant audio signals. [Background technology]

[0002] The problem of reproducing accurate and perceptually convincing reverberation is well known.

[0003] For example, when considering the requirements for acoustically simulating a concert hall, we consider that we only need the response of the acoustic space at one discrete listening point, i.e., at one listener's ear, due to one discrete point source of acoustic energy. The direct signal propagating from the source to the listener's ear can be simulated using a single delay line in series with an attenuation scaling or low-pass filter. Second, each sound ray that reaches the listening point via one or more reflections can be simulated using a delay line and some scaling factor or filter. More specifically, a tapped delay line can simulate many reflections. Each tap elicits one echo at the appropriate delay and gain, and each tap can be independently filtered to simulate air absorption and reflection loss. Since reverberation actually consists of many paths of sound propagation from each sound source to each listening point, in principle, a tapped delay line can accurately simulate any reverberant environment (Smith, 1993).

[0004] The simulation approach appears simple, as does its basic problem, because i) tapped delay lines are relatively computationally expensive compared to other techniques such as signal attenuation for signal path distribution, and ii) each tapped line only handles a "point-to-point" transfer function, i.e., the transfer function from one point source to one ear, whereas many point-to-point transfer functions are required, and each point-to-point transfer should change when the source and / or listener moves or something else in the room changes.

[0005] The problem is compounded when one considers that each echo may be perceived as coming from a specific angle of arrival in three-dimensional space. For example, at least some reverberant reflections should be spatialized, i.e., distributed across spatially configured loudspeaker channels or filtered to take into account the head-related transfer functions (HRTFs) of the ear pinnae, so that the reflections appear to come from their natural directions (Kendall and Martens, 1984). Therefore, if anything in the listening space changes, including the position of the sound source or the listener, the spatialization should also change.

[0006] For music, a typical reverberation time is about one second. Let's assume we choose a reverberation time of exactly one second. At an audio sampling rate of 50 kHz, each filter requires 50,000 multiplications and additions per sample, or 2.5 billion multiplications and additions per second. For example, dealing with the simulation of three sound sources and two listening points (two ears per listener), this amounts to 30 billion operations per second to recreate the reverberation. This computational load would require at least 10 Pentium CPUs clocked at 3 GHz, assuming the CPUs are doing nothing simultaneously and that both a multiplication and an addition can be initiated every clock cycle without wait states caused by the necessary memory accesses (Smith, 1993).

[0007] It can therefore be concluded that point-to-point transfer functions for reproducing sound reverberation are computationally very expensive. While some applications in the field of acoustic wave simulation, primarily used for scientific or measurement purposes, use full point-to-point transfer function derivative models such as those described in U.S. Patent No. 2015 / 78563, "Acoustic Wave Reproduction Systems" (Robertson, 2015), it is clear from the above that explicit methods based on physical modeling are computationally too expensive for most applications, and one must first take into account the perceptually important aspects of reverberation and how these can be accommodated by more efficient computational structures.

[0008] It is generally believed that the reverberation problem can be simplified without sacrificing perceptual quality. For example, the echo density is typically 2 where t is time. Thus, beyond a certain time, the amount of echoes becomes so large that they can be modeled as a uniformly sampled stochastic process without loss of perceptual fidelity. In particular, there is no need to explicitly calculate each echo for every sound sample. For smoothly decaying late reverberation, a suitable random process sampled at the audio sampling rate sounds perceptually equivalent. A required temporal density that is considered perceptually acceptable is 1000 echoes per second (Schroeder, 1961). However, the temporal density may need to be as high as 10,000 for impulsive sounds with large transients (Gardner, 1998).

[0009] Similarly, the number of resonant modes in any given frequency band is f, such that above a certain frequency the resonant modes are densely packed together to be perceptually equivalent to a statistically generated random frequency response. 2In particular, it is not necessary to explicitly implement more densely packed resonances; rather, the required modal density is equal to a regularly spaced frequency density across the frequency range, but not so regularly as this would produce audible periodicity in the time domain and upset the smoothness of the response.

[0010] The set criteria come somewhat close to exponentially decaying a white noise signal as the reverberation impulse response, which satisfies the criteria for smoothness in both the time and frequency domains (Moorer, 1979). However, because natural reverberation decays faster at high frequencies, the ideal reverberation impulse response is more like an exponentially decaying "colored" noise, where high-frequency energy decays faster than low-frequency energy.

[0011] Methods known in the art for recreating reverberation can be divided into two directions. On the one hand, there is the field of artificial reverberation, a technique introduced as the Schroeder all-pass section (Schroeder, 1961), typically consisting of elements comprising a delay line, a comb filter, and an all-pass filter, which to this day serves as the basis for most commercial devices for artificial reverberation and related effects. In many applications of artificial reverberation, such techniques are combined with feedback delay networks (FDNs), first proposed for use in artificial reverberation by Gerzon, who reasoned that while an individual bank of all-pass filters would result in poor quality reverberation, several such filters, when cross-coupled, could produce a measure of high quality reverberation (Gerzon, 1972).

[0012] The state of the art presents many applications that use typical elements of artificial reverberation, such as those described in US Patent No. 2013 / 216073 "Speaker and Room Virtualization using Headphones" (Lau, 2013) and WO 2016 / 130834 "Reverberation Generation for Headphone Virtualization" (Fielder et al., 2016). While such systems are generally optimized for effective reverberation processing in stereo sound systems, i.e., stereo speaker configurations or headphone virtualization with binaural audio rendering techniques, some applications also typically address the adaptation of proposed methods for multi-channel distribution using specific diffusion algorithms of stereo signals across multiple channels, such as those described in US Patent No. 2008 / 273708 "Early Reflection Method for Enhanced Externalization" (Sandgren et al., 2008).

[0013] While such techniques known in the art have been successful in effectively processing sound reverberations that meet perceptual criteria, and as such are pleasing to the ear as reverberation, and comprise useful tools in various audio and music production practices and user applications, such applications are unable to translate actual fundamental aspects of acoustic space, such as the shape of the space, manifested in particular resonant modes that vary strongly throughout the space based on the standing wave distribution arising from echoes and occurring with sound reflecting in a particular dimensional shape, the size of the room, which affects the reflection lengths and fundamental frequencies of the resonant modes, and the materiality of the space, manifested in the decay times of various frequency bands that dissipate faster or slower based on the absorption and reflectivity of the materials from which the space is constructed.

[0014] In fact, due to their fundamental technical design, known techniques for artificial reverberation in the art are essentially unable to achieve such qualitative aspects that constitute the nature of real acoustic spaces. This is because smoothness and density criteria are achieved using a carefully selected set of fixed delay times and fixed signal distributions that satisfy results for a limited set of all-pass sections, including delay lines, passing through the FDN network. As a result, all reverberation systems based on such designs have their own "abstract" nature, not representative of any particular acoustic space or situation, such as a space of a certain shape, size, or constructed from a specific material, and only a limited set of parameters can be altered to change the given properties, such as decay time, damping amount, and predelay, which are typical front-end user variables.

[0015] On the other hand, there is the field of convolution-based reverberation, which typically uses recorded impulse responses (IRs) from real acoustic spaces to perform a convolution of the recorded signal with an audio input signal, as described in EP 3026666 "Reverberant Sound Adding Apparatus, Reverberant Sound Adding Method and Reverberant Sound adding Program" (Shirakihara et al., 2015). Such techniques capture the specific characteristics of the acoustic space, including properties due to its shape, size, and material composition, and successfully incorporate these aspects into the audio signal resulting from the convolution operation.

[0016] Nevertheless, obtaining recorded data from a space raises many practical issues, including the need to access a specific space for the purpose of recording the IR, as well as considerations regarding the standardization of technical means and instruments for obtaining the desired IR. Although several libraries with convolutional reverberation from real spaces have been collected in recent years, the availability of different rooms and geometries means that the choice is still very limited.

[0017] Another problem arises from the enormous amount of data required to obtain a sufficiently high resolution for this technique, corresponding to X amount of directional signals from X amount of positions in space. Current standards specify X32 degrees at X24 horizontal positions as the quality criterion for obtaining convolution-based reverberation applicable to limited-resolution (horizontal) listening points and discrete point sound sources in a virtualized spatial model. This constitutes 786 pre-recorded audio signals that require convolution with the input signal at a sufficiently high sample rate (≥ 44.1 kHz p / s), which is still prohibitively expensive in terms of available memory and processing on most currently available CPU standards, thus limiting user applications using high-quality, real-time convolution-based reverberation.

[0018] Furthermore, the recorded data constitutes a specific, fixed set of data for one space that cannot be adapted. Therefore, transformation of the virtual space model by changing or adapting its size, shape, or other attributes (in real time) is not possible. This gives convolution-based approaches a major disadvantage over systems based on artificial reverberation, where the features are more widely adaptable by the user, at least with regard to aspects of frequency response such as decay time and damping, and the generated audio output is more easily adaptable to provide reproductions for a variety of output systems, i.e., speaker configurations.

[0019] Therefore, there is a need in the art for a method for generating reverberation that can accurately reproduce the characteristics of an acoustic space, such as its shape, size, and materiality, with the efficiency and (real-time) adaptability of the elements used in artificial reverberation. [Prior art documents] [Patent documents]

[0020] [Patent Document 1] US Patent No. 2015 / 78563 [Patent Document 2] US Patent No. 2013 / 216073 [Patent Document 3] WO2016 / 130834 [Patent Document 4] US Patent No. 2008 / 273708 [Patent Document 5] Dutch Patent Application No. 2024434 [Patent Document 6] Dutch Patent Application No. 2025950 Summary of the Invention [Means for solving the problem]

[0021] Accordingly, a method for generating a reverberant audio signal associated with a virtual object is disclosed, the method comprising storing a representation of the virtual object, the representation defining a plurality of virtual points that make up the virtual object, the virtual points having respective virtual positions relative to one another, the virtual points belonging to a symmetry group of the virtual points, the symmetry group of the virtual points being: defining, for each virtual point of the plurality of virtual points, a set of one or more virtual distances including a respective virtual distance between the virtual point in question and each other virtual point of the plurality of virtual points; for each set of one or more distances associated with the hypothetical point, removing distances that are integer multiples of any other distance in the problem set to obtain a further set of one or more distances associated with the hypothetical point; For each further set of one or more distances associated with the virtual point, determining distinct distances within the further set in question to form a virtual point-specific set of one or more distances associated with the virtual point; determining virtual points having the same respective virtual point-specific set of one or more distances to form a symmetry group of the virtual points, such that the symmetry group of a virtual point is associated with the same set of one or more symmetry group distances as the virtual point-specific set of that virtual point; It can be obtained by

[0022] The one or more sets of symmetry group distances, each set associated with a symmetry group, together form one or more further sets of distances. The method further includes receiving, storing, and / or generating an input audio signal and, for each virtual point, determining a virtual point audio signal component based on the input audio signal or a filtered version thereof. The method also includes combining the virtual point audio signal components to obtain a composite audio signal and, for each distinct distance in the one or more further sets of distances, determining one or more distance audio signals based on the composite audio signal. The method also includes determining a reverberant audio signal based on the one or more distance audio signals and the virtual point audio signal component.

[0023] This method allows spatial information of a virtual object, such as a room, to be incorporated into a generated reverberant audio signal. As a result of the method, the generated reverberant audio signal possesses the same characteristics as an audio signal recorded in a particular room with its particular acoustics, but the reverberation is actually caused by a virtual object represented by a virtual point. For example, the virtual object may be a virtual room with virtual walls, floor, and ceiling, which are made of particular materials. In such a case, a subject listening to the reverberant audio signal may perceive the audio signal as if they were actually standing in the virtual room.

[0024] The center point may be understood to be a point about which a rotation axis can be defined such that a virtual object, starting from an initial orientation and position, can rotate less than 360 degrees around the rotation axis and arrive at an orientation and position identical to the initial orientation and position. Such a rotation of a virtual object may also be referred to as an object-preserving rotation. For example, if the virtual object is a rectangular plane, a rotation of 90 degrees, less than 360 degrees, around such a rotation axis through the midpoint will place the rectangular virtual object in the same position and orientation as the initial position and orientation, so the rotation axis is perpendicular to the plane and the center point is the midpoint of the rectangle.

[0025] Any first and second virtual points of a virtual object that are positioned symmetrically relative to a center point of the virtual object may be understood as the first virtual point being at the second virtual point when an object-preserving rotation is performed.

[0026] Additionally or alternatively, any first and second virtual points of a virtual object that are positioned equidistant from a center point may be understood to be positioned symmetrically about the center point.

[0027] As used herein, combining two or more signals may include summing the signals.

[0028] In one embodiment, the method includes, for each symmetry group, determining one or more symmetry group audio signals based on the determined distance audio signals, and this embodiment also includes determining a reverberation audio signal based on the symmetry group audio signals and the virtual point audio signal components.

[0029] Since each symmetry group audio signal is determined based on a distance audio signal, in this embodiment the reverberation audio signal is also determined based on one or more distance audio signals.

[0030] In one embodiment, determining the virtual point audio signal components for each virtual point based on the input audio signal or a filtered version thereof comprises, for each virtual point, performing a virtual point-specific operation on the input audio signal or a modified, e.g., filtered, inverted, and / or attenuated or amplified, version thereof, wherein performing the virtual point-specific operation comprises performing a time delay operation that introduces a time delay approximately equal to the virtual distance between the virtual point in question and the virtual sound source divided by the speed of sound.

[0031] Preferably, the virtual representation also defines a virtual position of the virtual point relative to the virtual sound source and / or relative to the observer. The virtual distance between a virtual point of a virtual object and a virtual sound source may be defined as the virtual distance between the virtual point and the center of the virtual sound source. The generated reverberant audio signal may be understood to reflect an audio signal originating from this virtual sound source.

[0032] This embodiment is advantageous in that the shape of the virtual object is taken into account when generating the reverberant audio signal.

[0033] The speed of sound may be the speed of sound in a virtual medium defined between the virtual object, the virtual sound source, and the observer. For example, the virtual medium between the virtual object and the virtual sound source may be defined as air at a temperature of 20° C. and an average humidity of 50%. In such a case, the speed of sound should be approximately 343 m / s, which is the speed of sound in air at a temperature of 20° C. and an average humidity of 50%. A representation of the virtual object may define the virtual medium between the virtual object and the virtual sound source.

[0034] In one embodiment, determining one or more distance audio signals for each distinct distance in the one or more further set of distances includes determining a first distance audio signal and a second distance audio signal for each distinct distance in the one or more further set of distances, where determining the first distance audio signal for the distinct distance includes modifying the composite audio signal by performing a time delay operation to introduce a time delay and a signal feedback operation, and further determining the second distance audio signal for the distinct distance includes modifying the composite audio signal by performing a second time delay operation to introduce a second time delay, a second signal feedback operation, and a signal inversion operation.

[0035] It should be appreciated that preferably the first distance audio signal and the second distance audio signal differ only in that one is an inverted version of the other.

[0036] This embodiment is advantageous in that it increases the number of harmonics per distinct distance, as described with reference to FIG. 11, thereby increasing the modal density in the reverberant audio signal, and the harmonics per distinct distance are optimally spread, i.e., odd and even harmonics are distributed to symmetrically opposing virtual points that make up the virtual object.

[0037] In one embodiment, the first time delay introduced by the first time delay operation is equal to the discrete distance divided by the speed of sound.

[0038] It should be understood that the second time delay and the first time delay are identical in principle.

[0039] As indicated above, the speed of sound is preferably the speed of sound associated with the virtual medium between the virtual object and the virtual sound source, for example as defined by the representation of the virtual object.

[0040] In one embodiment, determining one or more symmetry group audio signals based on the distance audio signals for each symmetry group includes determining a first symmetry group audio signal and a second symmetry group audio signal for each symmetry group, selecting distance audio signals from all pairs of first and second distance audio signals, each pair determined for a respective distance from a set of one or more symmetry group distances associated with the symmetry group in question, and combining the selected distance audio signals to determine the first symmetry group audio signal and combining unselected distance audio signals from all pairs of first and second distance audio signals to determine the second symmetry group audio signal.

[0041] The selection of the signals to be used to determine the first symmetric group audio signal, and therefore the selection of the signals to be used to determine the second symmetric group audio signal, may be performed in any manner. This embodiment ensures that one of the first and second distance audio signals contributes to the first symmetric group audio signal and the other contributes to the second symmetric group audio signal.

[0042] In one embodiment, determining an audio signal based on a symmetric group audio signal and a hypothetical point audio signal component comprises combining the symmetric group audio signal and the hypothetical point audio signal component to determine the reverberant audio signal.

[0043] The reverberant audio signal preferably mimics first order reflections from the virtual object as well as the reverberant tail produced by the virtual object, the virtual point audio signal component causing the reverberant audio signal to mimic first order reflections from the virtual object, and the symmetry group audio signal causing the reverberant audio signal to include the reverberant tail associated with the virtual object.

[0044] In one embodiment, determining an audio signal based on a symmetry group audio signal and a virtual point audio signal component comprises combining the symmetry group audio signal and the virtual point audio signal component to determine the reverberant audio signal, wherein combining the symmetry group audio signal and the virtual point audio signal component to determine the audio signal comprises determining modified audio signal components, and further determining the modified audio signal components comprises adding a first or second symmetry group audio signal of the symmetry group in question to each virtual point audio signal component determined for the virtual point belonging to the symmetry group.

[0045] This embodiment provides an efficient method for combining symmetric group audio signals and virtual point audio signal components.

[0046] In principle, all virtual points, and therefore all audio signal components, belong to a single symmetry group for which two symmetry group audio signals are determined. Determining whether the first or second symmetry group audio signal is added to an audio signal component can be performed according to the principle that the first and second symmetry group audio signals are added alternately to the respective audio signal components such that half of the audio signal components associated with a symmetry group, approximately half if the number of virtual points in the symmetry group is odd, are combined with the first symmetry group audio signal for that symmetry group, and the other half of the audio signal components are combined with the second symmetry group audio signal for that symmetry group.

[0047] The step of determining the modified audio signal components may further include other operations to add resonance, depth, height, and distance characteristics to the audio signal components.

[0048] In one embodiment, the method comprises performing a method according to any of the previous claims for generating a further reverberant audio signal for a further virtual object, wherein the determined reverberant audio signal associated with the virtual object is used as the input audio signal.

[0049] This embodiment makes it possible to simulate one reverberant virtual object generating a reverberant audio signal incident on another virtual object, which makes it possible to generate reverberant audio signals for complex virtual systems that include multiple virtual objects, such as walls, ceilings, floors, and several differently oriented surfaces.

[0050] In one embodiment, the method further includes combining a reverberant audio signal associated with the virtual object with a further reverberant audio signal associated with a further virtual object, and optionally providing the combination to one or more speakers.

[0051] In one embodiment, the method further includes providing the determined audio signal to one or more speakers.

[0052] In one embodiment, the method includes wherein providing the modified audio signal components to one or more speakers includes providing the modified audio signal components to a panning system configured to distribute the modified audio signal components to a plurality of speakers.

[0053] Distributing the modified audio signal components may include determining several output audio signals, one for each speaker, based on the modified audio signal components.

[0054] In one embodiment, the method further comprises filtering the input audio signal before determining the virtual point audio signal components for each virtual point, wherein filtering the input audio signal comprises applying a multi-band filter comprising attenuating respective frequency bands in the input audio signal using respective attenuation coefficients, the respective attenuation coefficients being determined based on the material of the virtual object.

[0055] Preferably, the representation that defines the virtual object also defines what material the virtual object is made of. To illustrate, the representation may define that the virtual object is made of limestone. The material of the virtual object may have known specific absorption coefficients for each frequency band. The attenuation coefficients used for the multi-band filter may be determined based on these absorption coefficients.

[0056] In one embodiment, determining one or more distance audio signals for each distinct distance in the one or more further set of distances includes determining a distance audio signal for each distinct distance in the one or more further set of distances. Determining such distance audio signals includes modifying the composite audio signal by performing a time delay operation to introduce a time delay, a signal attenuation operation, a low pass filter operation, and a signal feedback operation. This embodiment also includes determining a density index for at least one symmetry group of the virtual point, preferably for each symmetry group of the virtual point. Determining the density index for the at least one symmetry group includes: determining, for each distance from a set of one or more symmetry group distances associated with at least one symmetry group, how many feedback operations have been performed per unit of time to determine a distance audio signal for the distance in question, e.g., by dividing said unit of time by a time delay introduced by a time delay operation performed to determine the distance audio signal in question, thus obtaining a respective number of feedback operations performed for each distance from a set of one or more symmetry group distances associated with at least one symmetry group; summing the number of each feedback operation performed to obtain a density index for the symmetry group of the imaginary point; Includes:

[0057] This embodiment also includes receiving a threshold value for the density index, determining that the determined density index is lower than the threshold value, and, based on this determination, modifying the stored representation by increasing the number of virtual points that make up the virtual object.

[0058] In principle, increasing the number of virtual points, sometimes referred to as increasing the resolution of the virtual points, results in a larger amount of distinct distances constituting one or more additional sets of distances, and a smaller time delay used to determine each distance audio signal. As a result, more feedback operations are performed per unit of time. Each feedback operation can be understood to represent an echo. Therefore, this embodiment can be said to ensure that sufficient echoes are generated.

[0059] In one embodiment, the low pass filter operation is: determining that the signal to be low-pass filtered is associated with a Nyquist frequency that is lower than a cutoff frequency associated with the low-pass filter operation; based on this determination, upsampling the signal to be filtered so that it is associated with a Nyquist frequency greater than or equal to said cutoff frequency; low-pass filtering the upsampled signal; Optionally, determining that the filtered signal is associated with a sample rate that is higher than the output sample rate, the output sample rate being a sample rate that can be output by the output system, and based on this determination, downsampling the filtered signal. Includes:

[0060] One aspect of the present disclosure relates to a computer comprising a computer-readable storage medium having computer-readable program code embodied therein and a processor, preferably a microprocessor, coupled to the computer-readable storage medium, wherein in response to executing the computer-readable program code, the processor is configured to perform any of the methods described herein.

[0061] One aspect of the present disclosure relates to a computer program or suite of computer programs including at least one software code portion, or a computer program product storing at least one software code portion, which, when executed on a computer system, is configured to perform any of the methods described herein.

[0062] One aspect of the present disclosure relates to a non-transitory computer-readable storage medium storing at least one software code portion, the software code portion being configured to perform any of the methods described herein when executed or processed by a computer.

[0063] As will be appreciated by those skilled in the art, aspects of the present invention may be embodied as a system, method, or computer program product. Accordingly, aspects of the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which may be generally referred to herein as a "circuit," "module," or "system." Functionality described in this disclosure may be implemented as an algorithm executed by a computer processor / microprocessor. Furthermore, aspects of the present invention may take the form of a computer program product embodied in one or more computer-readable medium(s), e.g., having computer-readable program code stored thereon.

[0064] Any combination of one or more computer-readable media may be utilized. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of computer-readable storage media may include, but are not limited to, the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of the present invention, a computer-readable storage medium may be any tangible medium that contains or can store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0065] A computer-readable storage medium may include a propagated data signal having computer-readable program code embodied therein, for example, in baseband or as part of a carrier wave. Such a propagated signal may take any of a variety of forms, including, but not limited to, electromagnetic, optical, or any suitable combination thereof. A computer-readable signal medium is not a computer-readable storage medium but may be any computer-readable medium that can communicate, propagate, or transfer a program for use by or in connection with an instruction execution system, apparatus, or device.

[0066] Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including, but not limited to, wireless, wired, fiber optic, cable, RF, etc., or any suitable combination of the foregoing. Computer program code for carrying out operations for aspects of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java™, Smalltalk, C++, and conventional procedural languages ​​such as the “C” programming language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).

[0067] Aspects of the present invention are described below with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor, particularly a microprocessor or central processing unit (CPU), of a general-purpose computer, special-purpose computer, or other programmable data processing device, to create a machine such that the instructions, executed via the processor of the computer, other programmable data processing device, or other device, create means for performing the functions / acts specified in the flowchart and / or block diagram blocks.

[0068] These computer program instructions may also be stored in a computer-readable medium that can instruct a computer, other programmable data processing apparatus, or other device to function in a particular manner to produce an article of manufacture containing instructions that implement the functions / acts specified in the flowchart and / or block diagram blocks, where the instructions stored in the computer-readable medium.

[0069] Computer program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device and cause the computer, other programmable apparatus, or other device to perform a series of operational steps to generate a computer-implemented process, such that the instructions executing on the computer or other programmable apparatus provide a process for performing the functions / operations specified in the flowchart and / or block diagram blocks.

[0070] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code, comprising one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions described in the blocks may occur in a different order than that described in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the flowchart diagrams and / or block diagrams, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or a combination of dedicated hardware and computer instructions.

[0071] Further provided are computer programs for performing the methods described herein, as well as non-transitory computer-readable storage media storing the computer programs, which may, for example, be downloaded (updated) into existing data processing systems or stored in these systems at the time of manufacture.

[0072] Elements and aspects discussed with or in connection with a particular embodiment may be combined with elements and aspects of other embodiments as appropriate, unless expressly stated otherwise. Embodiments of the present invention will be further described with reference to the accompanying drawings, which schematically illustrate embodiments in accordance with the present invention. It will be understood that the present invention is in no way limited to these particular embodiments.

[0073] Aspects of the present invention will now be described in more detail with reference to exemplary embodiments illustrated in the drawings. [Brief explanation of the drawings]

[0074] [Figure 1] FIG. 2 illustrates an observer receiving a reverberant audio signal according to an embodiment. [Figure 2] 1 is a flowchart illustrating a method for generating a reverberant audio signal according to one embodiment. [Figure 3] 3 is a flow chart illustrating inputs for each module described in FIG. 2 according to one embodiment. [Figure 4] FIG. 1 illustrates a method for generating a reverberant audio signal according to one embodiment. [Figure 5] FIG. 1 illustrates an embodiment for generating a reverberant audio signal. [Figure 6A] 1 is a detailed flowchart illustrating a method for generating a reverberant audio signal according to one embodiment. [Figure 6B] 1 is a detailed flowchart illustrating a method for generating a reverberant audio signal according to one embodiment. [Figure 6C] 1 is a detailed flowchart illustrating a method for generating a reverberant audio signal according to one embodiment. [Figure 7] FIG. 1 illustrates the generation of a reverberant audio signal for a mono sound system according to one embodiment. [Figure 8] FIG. 1 illustrates the generation of a reverberant audio signal for a stereo sound system according to one embodiment. [Figure 9] FIG. 1 shows symmetry group distances for six symmetry groups. [Figure 10] FIG. 10 illustrates a further set of one or more distances according to one embodiment. [Figure 11] FIG. 1 illustrates the symmetry groups of several different virtual objects. [Figure 12] FIG. 10 is a flow process diagram for a “value filter” operation to determine a time delay according to one embodiment. [Figure 13] FIG. 1 illustrates a flow process for a “sample rate interpolation” operation according to one embodiment. [Figure 14] FIG. 2 is a diagram illustrating a user interface according to an embodiment. [Figure 15] 1 is a block diagram illustrating a data processing system according to one embodiment. [Figure 16AB] 16A and 16B are diagrams illustrating modules for encoding resonance, depth, height, and distance into audio signal components, according to respective embodiments. [Figure 16CD] 16C and 16D are diagrams illustrating modules for encoding resonance, depth, height, and distance into audio signal components, according to respective embodiments. [Figure 17] FIG. 1 is a diagram of a panning matrix according to one embodiment. [Figure 18A] 1A-1C illustrate modules for adding resonance characteristics to audio signal components according to respective embodiments. [Figure 18BC] 18B and 18C illustrate modules for adding resonant characteristics to audio signal components, according to respective embodiments. [Figure 18D] 1A-1C illustrate modules for adding resonance characteristics to audio signal components according to respective embodiments. [Figure 18E] 1A-1C illustrate modules for adding resonance characteristics to audio signal components according to respective embodiments. [Figure 18FG] 18F and 18G illustrate modules for adding resonant characteristics to audio signal components, according to respective embodiments. [Figure 19] 3A-3C illustrate modules for adding depth characteristics to audio signal components according to respective embodiments. [Figure 20] 3A-3C illustrate modules for adding distance characteristics to audio signal components according to respective embodiments. [Figure 21] FIG. 1 illustrates a shape generator for determining the dimensions, positions, and orientations of a virtual object and its virtual points, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0075] In the figures, like reference numbers indicate similar or identical elements. Additionally, elements drawn with dashed lines are optional elements.

[0076] The present disclosure relates to methods and systems for generating reverberant audio signals. A real reverberant audio signal can be understood to be formed by the reflection of sound by an object and subsequent vibrations in such an object. The generated reverberant audio signals described herein are associated with virtual objects in the sense that the reverberant audio signals have the same characteristics as an actually recorded audio signal that is subject to the reverberation of such an object. In principle, the virtual object can have any shape, size, and / or characteristics.

[0077] The methods described herein enable the generation of reverberant audio signals with the highest computational efficiency, enabling real-time processing. The ability to generate reverberant audio signals in real time allows for modifying properties of a virtual object, such as the shape, position, orientation, and / or material from which the virtual object is formed, and immediately generating a reverberant audio signal associated with the modified virtual object. In particular, conditions of the virtual acoustic space, such as the material composition of the virtual acoustic space, sound propagation conditions, spatial distance, height, and depth of sound sources reflecting off objects within the space and / or with other sound sources, can be modified, and these modifications can be processed rapidly.

[0078] The method and system for generating a reverberant audio signal uses relatively simple elements of artificial reverberation, such as delay lines and low-pass filters. The techniques disclosed herein are designed to minimize the number of delay lines required in the system, for example, in that smart signal distribution across multiple audio components is achieved, and are very computationally inexpensive.

[0079] The techniques disclosed herein make it possible to add characteristics to an input audio signal so that the input audio signal appears as if it was generated and / or recorded in a real acoustic space of a given shape, size, and materiality.

[0080] The method is optimally efficient and resolution-scalable for CPU requirements ranging from low to high. The method enables interactive models for generating reverberation, as the process can be performed in real time while running at the optimal resolution given the CPU requirements used. The method allows for adaptability to any variety of speaker configurations and quantities of speakers used in the configuration, while also being backward compatible with existing audio playback formats, such as mono, stereo, and / or HTRF-based headphone sound playback, inspiring new approaches to speaker system design.

[0081] The method for obtaining delay times is not based on a fixed set of time delay values ​​with statistically valid results, but can be generated in real time from the operation of a "shape generator" that allows for the modeling of shapes of any type and / or nature by determining a virtual point resolution, i.e., a set of points defined in a virtual shape. The resulting time delays are filtered using a "value filter" operation that estimates essential data from the shape to accurately represent the acoustic shape of the first and early reflections of the reverberation, while simultaneously generating a reverberation tail, i.e., a delayed reverberation, that meets the perceptual criteria of smoothness, temporal density, and modal density using the same set of time delays, without the need to incorporate a feedback delay network (FDN). Furthermore, the present invention introduces a new approach for automating the optimization of density and sample rate requirements for different user conditions set by the user.

[0082] As explained above, a method for generating a reverberant audio signal involves a representation of a virtual object 2, such as a square plate as shown in FIG. 1. The representation defines a number of virtual points, numbered {1, 2, 3, ..., N} in FIG. 1. In the present disclosure, "N" denotes the total number of virtual points. The virtual points have respective virtual positions relative to each other and relative to a center point 4 of the virtual object 2. The virtual points belong to a symmetry group of virtual points, and virtual points that are symmetrically positioned with respect to the center point belong to the same symmetry group of virtual points. This will be explained in more detail with reference to FIG. 9.

[0083] The virtual points defined by the representation also have a specific position relative to the observer 6. Thus, a virtual object may be located, for example, at a specific distance from the observer, at a depth below the observer, or at a height above the observer. Furthermore, the virtual points may also have a position relative to the sound source 8.

[0084] The observer 6 may perceive the virtual object 2 as having a shape, i.e., a distinct geometry, size, and materiality, and may perceive the sound source 8 and the reverberant space at distinct heights, depths, and distances relative to the observer. Such perception is analogous to exploring how sound is experienced in a real space of such shape, size, and material configuration, as well as how it moves through this space and sounds from any position and angle. Thus, the listener may virtually and / or physically move through the space and experience the resulting reverberation from any position and angle inside and / or outside the reverberant space, i.e., the acoustics of a space of a particular size, shape, and materiality, and experience the results of the reverberant excitation of a virtual object located at any virtual location.

[0085] FIG. 2 is a flowchart illustrating a method for generating a reverberant audio signal according to one embodiment. Here, an input audio signal x is provided to a multi-band filter configured to attenuate respective frequency bands in the input audio signal using respective attenuation coefficients. The attenuation coefficients may be understood to define the degree of attenuation for a particular frequency band. For example, a multi-band filter may be configured to attenuate a first frequency band using an attenuation coefficient of 0.7 and a second frequency band using an attenuation coefficient of 0.5. As a result, the intensity of frequencies in the first frequency band is attenuated by a factor of 0.7, and the intensity of frequencies in the second frequency band is attenuated by a factor of 0.5. Such a filter 10 may be referred to as an absorption filter because it may be used to model the absorption of sound by virtual objects. In one example, the absorption filter is an eight-octave band equalizer.

[0086] 2, the output of the filter 10 is then provided to a first reflection module 12. Such a module preferably comprises multiple parallel signal flows, one for each virtual point of the virtual object 2, to determine a virtual point audio signal component y_n for each virtual point, as shown. The modules may also be referred to as flow processes, flowcharts, etc.

[0087] In the embodiment of FIG. 2, the virtual point audio signal component y_n is a composite audio signal denoted by 14.

[0088]

number

[0089] are combined, eg summed, to obtain

[0090] Additionally, the composite audio signal 14 is provided to a second filter 16 which may be identical to the filter 10 .

[0091] The output of the second filter 16 is then provided to a module 18 for generating a reverberation tail. The module 18 comprises multiple parallel signal flows, one for each distance audio signal to be generated, as described herein. The module 18 determines one or more distance audio signals d_k+ / −, i.e., {d_1+, d_1−, d_2+, d_2−, ..., d_K+, d_K−} (not shown). As used in this disclosure, “K” denotes the total number of distinct distances in the one or more further sets of distances, as described herein.

[0092] The module 18 outputs several symmetric group audio signals s_m+ / -, i.e., {s_1+, s_1-, s_2+, s_2-, ..., s_M+, s_M-}, where "M" as used in this disclosure denotes the total number of symmetric groups.

[0093] In the embodiment of Figure 2, the symmetric group audio signal s_m+ / - is combined with the virtual point audio signal component y_n, which results in a reverberant audio signal according to one embodiment.

[0094] 2 further illustrates optional modules: a resonance module 20, a depth module 22, a height module 24, a distance module 26, and a panning system 28. These modules are not required for generating a reverberant audio signal, but are required for coherent projection of the reverberant audio signal to an observer, i.e., for the reverberation to be perceived by the observer at distinct depths, heights, and distances and angles. The resonance module 20 is configured to perform a spatial wave transformation on the audio signal components to add resonance characteristics to the audio signal components, the sum of which may be a mixdown 21 of the audio input signal for (second) signal processing, as described herein. The depth module 22 is configured to encode depth into the audio signal components. The height module 24 is configured to encode height into the audio signal components. The distance module 26 is configured to encode distance into the audio signal components.

[0095] Adding resonance characteristics, which may be performed by the resonance module 20, may include modifying the audio signal component to obtain a first modified audio signal component (see FIG. 16A ). This modification of the audio signal component optionally includes a signal inversion operation 74, a signal delay operation 75 that introduces a time delay, and optionally a signal feedback operation 73, as shown. In the illustrated embodiment, the fed-back signal is attenuated as indicated by an amplifier 76 having a gain less than one. The first modified audio signal component is then combined with the audio signal component via summation 78 to obtain a second modified audio signal component. Furthermore, the second modified audio signal is further modified by an attenuation operation 79 and, optionally, a high-pass filter operation 80 to obtain an audio signal component y_n′ associated with the virtual point of the virtual object.

[0096] The formula for determining the time delay to be introduced to determine the modified audio signal component is: Δt=Vx n / v where V is the dimensional volume of the shape and x n where x represents the coefficients for points on the virtual shape, each point having a relative spatial position given in Cartesian coordinates (x, y, z), and v is a constant related to the speed of sound in the medium. The determination of audio signal components is also described in Dutch Patent Applications Nos. 2024434 and 2025950, the contents of which are to be considered as included in this disclosure in their entirety.

[0097] The attenuation operation 79 after the summation operation 78 may include reducing the gain G of the audio signal by −6 bB.

[0098] In this disclosure, the values ​​within the triangles, i.e., values ​​in an attenuation or amplifier operation, may be understood to indicate constants by which a signal is multiplied. These constants are often denoted by "a" or "b." Thus, if such a value is greater than 1, signal amplification is performed. If such a value is less than 1, signal attenuation is performed.

[0099] The cutoff frequency f of the high-pass filter depends on the point n on the virtual shape. c teeth, r n / R≦0.5 for f c =v / V2(1-r n / R) r n / f for R>0.5 c =v / V2(r n / R) where v is a constant related to the speed of sound in the medium, V is the dimensional volume of the virtual shape, and r n denotes the spherical radius from the center of the virtual shape to point n, and R denotes the spherical radius from the center of the shape passing through the vertex where two or more sides of the virtual shape intersect. If there are two or more values ​​of R, the largest R is considered.

[0100] It should be understood that instead of the flowchart shown in FIG. 16A, any of the flowcharts shown in any of FIGS. 18A-18G may be used instead to add resonance characteristics having the same parameter values.

[0101] Adding depth characteristics to an audio signal component as may be performed by module 22 may include modifying the audio signal component y_n of interest using a time delay operation 86 that introduces a time delay, a signal attenuation 88, and a signal feedback operation 90 to obtain a modified version of the audio signal component, and combining (92) the modified version of the audio signal component with the audio signal component of interest (see FIG. 16B ). The signal attenuation 88 is performed depending on the virtual depth below the target of the virtual point associated with the audio signal component of interest.

[0102] In this embodiment, the signal attenuation is defined by the parameter "b": for a value b=0, the depth of the virtual point below the object is not encoded, and for a value b=1, the maximum depth of the virtual point associated with the audio signal component is encoded.

[0103] The value "a" is used to optionally attenuate or amplify the result of the combination of the modified audio signal with the input audio signal (94). a=(ab)x where x is a multiplication factor to correct the signal gain G depending on the amount of signal feedback b which affects the steepness of the high frequency dissipation curve. By varying the value b, preferably between 0 and 1, depth variations are added to the audio signal.

[0104] Preferably, the time delay Δt introduced by the time delay operation is as short as possible, e.g., less than 0.00007 seconds, preferably less than 0.00005 seconds, more preferably less than 0.00002 seconds, and most preferably about 0.00001 seconds. For a digital sample rate of 96 kHz, the time delay may be 0.00001 seconds.

[0105] It should be understood that instead of the flowchart shown in FIG. 16B, any of the flowcharts shown in FIG. 19 may be used instead, using the same values ​​for the parameters.

[0106] Adding a height characteristic to an audio signal component, as may be performed by height module 24, includes modifying the audio signal component of interest using a signal inversion operation 140, a signal delay operation 142 that introduces a time delay, and a signal attenuation 144 to obtain a modified version of the audio signal component, and combining (146) the modified version of the audio signal component with the audio signal component of interest (see FIG. 16C), where the signal attenuation 144 is performed depending on the virtual height of the virtual sound source.

[0107] In this embodiment, for a value b=0, no height characteristic is added to the audio signal component. For a value b=1, the maximum height of the virtual point is perceived. When the first attenuation operation is performed, the gain G of the value "a" of the optional attenuation 148 is: a=(1-b)x where x is a multiplication factor to correct the signal gain G according to the amount of attenuation b that affects the steepness of the low frequency dissipation curve. By varying the value b, preferably between 0 and 1, pitch variations are added to the audio signal.

[0108] Preferably, the time delay Δt introduced by the time delay operation 142 is as short as possible, e.g., less than 0.00007 seconds, preferably less than 0.00005 seconds, and more preferably less than 0.00002 seconds. Most preferably, it is about 0.00001 seconds. For a digital sample rate of 96 kHz, the time delay may be 0.00001 seconds.

[0109] Adding distance characteristics to audio signal components as may be performed by module 26 includes modifying the audio signal component of interest using a first time delay operation 160 that introduces a first time delay to obtain a first modified version of the audio signal component, a first signal attenuation operation 162, and a signal feedback operation 164, combining (166) the first modified version of the audio signal component with the audio signal component of interest to obtain a second modified version of the audio signal component, and performing a second signal attenuation 168 and, optionally, a second signal delay operation 170 that introduces a second time delay to the second modified version of the audio signal component (see FIG. 16D ), where the first signal attenuation 162 and the second signal attenuation 168 are performed depending on the virtual distance from the target.

[0110] Depending on the distance of the virtual point associated with the audio signal component in question, the value of the attenuation constant b for operation 162 and the value of the attenuation constant a for operation 168 are varied. The constants may be understood to indicate the constant by which the signal is multiplied. Thus, if such a value is greater than 1, signal amplification is performed. If such a value is less than 1, signal attenuation is performed. If b=0 and a=1, no distance is coded, and if b=1 and a=0, the maximum distance is coded. The gain G for the value a is: a=(1-b)x where the value of x is a multiplication factor applied to the amount of signal feedback that affects the steepness of the high frequency dissipation curve.

[0111] Preferably, the time delay Δt1 introduced by the time delay operation 160 is as short as possible, e.g., less than 0.00007 seconds, preferably less than 0.00005 seconds, and more preferably less than 0.00002 seconds. Most preferably, it is about 0.00001 seconds. For a digital sample rate of 96 kHz, the time delay may be 0.00001 seconds.

[0112] The optional time delay Δt2 introduced by the time delay operation 170 accounts for the Doppler effect associated with the movement of the virtual sound source. Δt2=r / v where r is the distance between the location of a virtual point associated with the audio signal component in question, shown in Cartesian coordinates (x, y, z), and the object, which may be represented as a viewpoint (x, y, z), and v is a constant representing the speed of sound in the medium.

[0113] It should be understood that instead of the flowchart shown in FIG. 16D, either of the flowcharts shown in FIG. 20 may be used, using the same parameter values.

[0114] The panning module 28 is configured to attenuate and sum the modified audio signal components y'_n to generate audio output signals z_p, i.e., {z_1, z_2, z_3, ..., z_P}, each associated with a distinct speaker p. As used herein, "P" refers to the total number of speakers.

[0115] The panning module is further described with reference to Figure 17, which is a flow chart illustrating a method for determining a speaker audio signal z_p for each speaker p of a plurality of speakers P. The illustrated method and system may also be referred to as a signal distribution matrix 28 or a panning matrix 28. In this embodiment, a speaker audio signal z_p is determined for each speaker p (not shown) of the plurality of speakers. Inputs to the signal distribution matrix are a plurality of audio signal components associated with respective virtual points of the virtual sound source, where the plurality of audio signal components y_n have been determined according to the methods described herein.

[0116] Each speaker p is associated with a speaker coefficient a_p. In the illustrated embodiment, determining a speaker audio signal z_p for speaker p includes attenuating each audio signal component y_n based on the speaker coefficient a_p to obtain a speaker-specific set of attenuated audio signal components. The speaker coefficient for a speaker may be determined based on the distance between the speaker in question and the virtual point. Attenuating each audio signal component y_n based on the speaker coefficient a_p may simply involve multiplication y_n*a_zp. In such a case, the speaker-specific set of attenuated audio signal components for speaker p may be described by {y_1*a_p; y_2*a_p; y_3*a_p;...; y_N*a_p}, where N indicates the total number of virtual points defined for the virtual sound source. The audio signal components in this set are then combined, e.g., summed, to arrive at a speaker audio signal z_p for speaker p. This method is performed for all speakers P.

[0117] The signal distribution matrix 28 may have a multiplier and a summation at each location where an input line, to which the multiplier's output signal is supplied, intersects with an output, as shown in FIG. 17 . The multiplier attenuates the signal received from the input line by a predetermined speaker coefficient specified by the controller, such as a value generated for each speaker amplitude by a panning system commonly known in the art, and outputs the resulting signal to the summation. The process in which the multiplier multiplies the signal by a predetermined coefficient is sometimes referred to as a "three-dimensional panning process." That is, the controller may provide appropriate values ​​for the associated coefficients corresponding to each output system so that the resulting audio signals provided to the target by the multiple speakers have a shape and a position in space, e.g., angle, distance, depth, and height relative to the target. As a result of the multiplier's processing, sound is properly simulated in terms of directional and dimensional propagation from the virtual sound source to the target. The summation supplies the audio output signals of the multiplier to respective output lines, each associated with a speaker in the speaker configuration. Each output line is assigned an attenuation coefficient. a=1 / N 2 where N is the number of audio signal components y in the signal distribution matrix. n is the number of The attenuation obtained for a is G(dB)=10log 10 (a) This is converted to gain G in decibels as

[0118] FIG. 3 is a flow chart illustrating the inputs for each module described in FIG. 2 according to one embodiment.

[0119] The method may include a shape generator 30 for determining the dimensions, position, and orientation of the virtual object, which may also be referred to as "shape data." The shape generator may output a set of virtual points that make up the virtual object, with each virtual point in the set of virtual points associated with a virtual position. The virtual points may be input to respective modules as shown. The shape generator is further described below with reference to FIG. 21.

[0120] 21 illustrates a method for determining a representation that defines a virtual object. The representation indicates the spatial dimensions of the virtual object, i.e., its shape and size, and its position relative to an object, and optionally, the density of the virtual object.

[0121] The virtual points may be evenly distributed on a surface or throughout a volume of a virtual object, with a higher density of virtual points on such a surface or throughout such a volume corresponding to a higher resolution.

[0122] It should be understood that a virtual object may be defined to be hollow. In such cases, the representation does not define virtual points "inside" the virtual object, but only on the exterior surfaces and edges of the virtual object. A virtual object may also be "solid." In such cases, the representation defines virtual points "inside" the virtual object, which may be evenly distributed throughout the interior volume of the virtual object, in addition to virtual points on the exterior surfaces and edges of the virtual object.

[0123] In one embodiment, the virtual object may have a geometric shape, i.e., a purely dimensional shape, or a semi-geometric, irregular, or organic shape. It should be understood that the virtual object may have any shape, and that any method may be used to determine the shape of the virtual object and the virtual points that make up the shape of the virtual object.

[0124] The density of the virtual points is sometimes referred to as the resolution of the virtual points and / or the "grid resolution."

[0125] 21 shows that obtaining the representation may include obtaining dimensions 210 of the virtual object and virtual point positions 212. Obtaining the shape dimensions 210 may include a shape generator generating a container 214 of scalable dimensions (x, y, z) and determining shape coordinates 216 and a shape volume within the bounds of the scaled dimensions to obtain the dimensions of the virtual object. In the illustrated example, the virtual object is shaped like a pyramid. Further, obtaining the virtual point positions 212 may include a grid generator determining a grid 218 to obtain the virtual point positions within the shape, where three major grids are introduced according to the dimensions of the selected shape, and determining a virtual point density 220 by defining the resolution of points along each of the introduced grids.

[0126] The infinite lattice L is L=a.(Z.v_1+Z.v_2+Z.v_3) where Z is the ring of integers, v_1, v_2, v_3 describe three vectors, and the constant a is = {point (x,y) such that x=an(v_1.x)+am(v_2.x), y=an(v_1.y)+am(v_2.y), n, m are integers} As such, it is related to the minimum increment.

[0127] Since sound is considered to propagate symmetrically in all directions, a pattern of overlapping or tangent circles generated by a grid is considered, where a sphere is centered at each imaginary point of the grid. The radius of the circles can be increased to affect the generated pattern of sound propagation in space.

[0128] In the illustrated example, a grid K=3 is shown, which means that three virtual points are defined along each axis of the grid.

[0129] Furthermore, the method may include a "sample rate interpolation" operation 32. This operation is preferably performed depending on the position data (x,y,z) of the virtual input source in order to modify the signal processing performed in the first reflection module 12. The operation is also performed depending on the obtained shape data, which serves to modify and optimize the signal processing of the reverberation module 18. Sample rate interpolation is further explained below with reference to FIG. 13.

[0130] The method for generating a reverberant audio signal may also include a "value filter" operation 34 and a "time density scaler" operation 36. These operations are performed depending on the obtained shape data and may serve to modify and optimize the signal processing of the reverberation module 18. Operations 34 and 36 are further described below with reference to FIG. 12.

[0131] The method may further include obtaining controller data relating to the position (x,y,z) and rotation (x,y,z) of the input source and a "point of view" indicative of the virtual and / or actual position of the listener, and applying the obtained data as input to a number of modules of digital signal processing.

[0132] The described audio signal processing in FIG. 2, as well as the shape data acquisition and modification, sample rate interpolation, value filtering and time density scaling in FIG. 3, may be performed in real time, for example, by executing software programs or code portions.

[0133] One aspect of the present disclosure relates to a data processing system configured to perform the methods for generating a reverberant audio signal described herein, such a data processing system may be connected to an audio output port and, optionally, to an audio input port for obtaining an audio input signal in real time.

[0134] It should be understood that one embodiment of the present invention includes performing one, some, and / or all (portions of) the modules described in FIGS. 2 and 3, and that the modules may be performed in a different order and / or may be performed repeatedly.

[0135] 4 shows a method for generating a reverberant audio signal according to one embodiment, where, in addition to determining a reverberant audio signal 40a by performing method 42a, a further reverberant audio signal 40b is determined for a further virtual object by performing method 42b. A further reverberant audio signal, for example signal 40c, can be generated by performing method 42c. Here, methods 42a, 42b, and 42c can be methods for determining a reverberant audio signal described herein. These methods can all be, for example, methods as shown in FIG. 2.

[0136] It should be understood that each method 42 shown in FIG. 4 is associated with a combination of a virtual sound source and a virtual object.

[0137] In the embodiment shown in Fig. 4, the signal generated by method 42a is taken as the input audio signal for determining the further reverberant audio signal 40b. Preferably, signal 21 shown in Fig. 2 is used as the input audio signal for generating the further reverberant audio signal. However, in principle, any signal generated while performing method 42 can be provided to method 42b as an audio input signal. Preferably, method 42b also receives as input the virtual position of the virtual object associated with method 42a, i.e., the virtual object for which method 42a generates reverberant audio signal 40a. This means that it is possible to determine the virtual position of a virtual point of the virtual object associated with method 42b relative to a "virtual sound source," the sound source for method 42b being the virtual object associated with method 42a.

[0138] A further method 42c for determining a reverberant audio signal 40c may then be performed using as input audio signals any of the signals generated while performing method 42b, preferably signal 21 shown in Figure 2. In such a case, method 42c also receives as input the position of the virtual object associated with method 42b, i.e., the position of the virtual object for which reverberant audio signal 40b is generated.

[0139] In principle, any number of signals may be input to any number of methods 42_x for determining a reverberant audio signal, either in succession or simultaneously. A method for generating a reverberant audio signal may use a combination of signals as audio input signals.

[0140] Method 42a may generate a signal that is input to a further method 42b for generating a further reverberant audio signal, while the signal generated while performing method 42b may simultaneously generate a signal that is used again as an input audio signal for method 42a. Thus, a feedback operation is possible between method 42a and method 42b, with both objects associated with these methods reflecting each other and generating reverberation depending on the audio input signal that was originally inserted into either method 42a or 42b. Of course, each method 42a and 42b is associated with a different reverberant virtual object and therefore has its own shape data as well as spatial position and rotation (x, y, z) as input.

[0141] Thus, the virtual objects referred to in this disclosure may form the virtual sound sources referred to in this disclosure for further virtual objects.

[0142] 4 thus allows for establishing a relationship between a sound source and a virtual object. A virtual object, as used herein, may be understood to virtually represent an object that reflects and reverberates sound from an excitation source, also called a "sound source" or "input source."

[0143] It should be appreciated that the reverberant audio signal 40a and the further reverberant audio signal 40b, and optionally further reverberant audio signals such as signal 40c, may be combined, optionally via a panning system, to determine a final reverberant signal that may be fed to a set of one or more loudspeakers.

[0144] FIG. 5 shows an embodiment in which an audio input signal x is used as an input audio signal for a first reverberant audio signal generation method according to an embodiment, and the resulting reverberant audio signal is then used as an input audio signal for a second reverberant audio signal generation method according to an embodiment.

[0145] Both methods produce a separate audio output signal z_p for each speaker p in the speaker configuration that is first summed and optionally attenuated and / or amplified by a multiplier in the range 0 to 1 before being fed to the speakers.

[0146] 6 is a detailed flowchart illustrating a method for generating a reverberant audio signal according to an embodiment, which comprises receiving an input audio signal (see left side of FIG. 6), which is optionally a sum of audio signal components determined during another method for generating a reverberant audio signal according to an embodiment as described in FIG.

[0147] The embodiment of Figure 6 includes providing an input audio signal to a multi-band filter 10 (see also Figure 2). Thus, this embodiment includes filtering the input audio signal, including applying the multi-band filter 10. Applying the multi-band filter 10 includes attenuating respective frequency bands in the input audio signal using respective attenuation coefficients, the respective attenuation coefficients being determined based on the material of the virtual object.

[0148] In a specific example, the multi-band filter consists of an eight-octave band equalizer in which the attenuation coefficient a (dB) is determined separately for each frequency band f. The values ​​of the attenuation coefficients are calculated using the standard formula for converting dB to power ratios: G(dB)=10log 10 (P t / P i ) where P t is the power level, P i (1) is the reference power level. G(dB) is the power ratio or gain in dB, a(dB)=G(dB), and the power ratio is α=1-(P t / P i ) is converted into an absorption coefficient as follows:

[0149] The value of α can be obtained from the standard ISO 354 for absorption coefficients, data which includes standardized methods of material testing (Bork, 2005b).

[0150] In one embodiment, the virtual object is a limestone wall. The octave bands in frequency f (Hz) are calculated using the random incidence absorption coefficient α of the limestone wall in ISO 354: f(125Hz) α=0.02, f(250Hz) α=0.02, f(500Hz) α=0.03, f(1000Hz) α=0.04, (f2000Hz) α=0.05, (f4000Hz) α=0.05, (f8000Hz) α=0.05, f(16000Hz) α=0.05 is given by

[0151] By applying an attenuation a (dB) for each octave band at frequency f (Hz), the input audio signal is modified so that the difference between the one-sided intensity of the reflected sound (Ir~Pt) and the one-sided intensity of the incident sound (Ii~Pi) is the absorption of sound energy by the limestone wall. The resulting reverberation therefore constitutes a distinct materiality characteristic.

[0152] The method further includes, for each virtual point of the virtual object (in the first reflection module 12, see also FIG. 2), determining a virtual point audio signal component y(t)_n based on a filtered version of the input audio signal. Here, determining each virtual point audio signal component y(t)_n includes performing a virtual point-specific operation on the (filtered) audio input signal. In the illustrated embodiment, the virtual point-specific operation includes performing an attenuation operation 52 to attenuate the signal, a low-pass filter operation 54 to filter frequencies higher than a threshold frequency, and a time delay operation 56 to introduce a time delay. As used herein, filtering frequencies above or below a threshold frequency, also referred to as a cutoff frequency, can be understood to attenuate gradually increasing frequencies up to and above such threshold frequency, or up to and below such threshold frequency, respectively. Thus, filtering does not mean that frequencies higher than the cutoff frequency or lower than the cutoff frequency, respectively, are completely removed and / or not removed.

[0153] The first reflection module 12 generates a first reflection of the sound. n is associated with a distinct virtual point of the virtual object. Furthermore, each virtual point audio signal component y nis determined based on the virtual position of its associated virtual point, in particular based on its virtual distance from a virtual sound source, i.e., the generated reverberant audio signal can be understood to reflect an audio signal originating from this virtual sound source, the position of which is defined, for example, in a virtual representation of a virtual object.

[0154] It should be understood that the attenuation operation 52, low-pass filtering operation 54, and time delay operation 56 for each virtual point may be performed in any order, and that one or more steps may be omitted, repeated, modified, and / or added.

[0155] The attenuation a (dB) in the calculation 52 of the audio signal component y_n depending on the distance r between the associated virtual point and the sound source is calculated by the transformation G → Pt / Pi and Pt=x(1 / r 2 ) where the sound intensity is given by I~Pt 2 where x is a multiplication factor between 0 and 1 that controls the amount of attenuation applied.

[0156] It should be understood that an optional multiplication factor x, "scaler," may be added at any other step in the described audio signal processing to provide a parameter scaling function, which may be implemented by the user sending controller data to modify aspects of the reverberant sound output, i.e., to increase or decrease the output value and thus to magnify or reduce the effect of the operation.

[0157] The low-pass filter 54 filters the audio signal component y, which depends on the distance r between the virtual point and the virtual sound source. n where the attenuation a (dB / m) of a frequency f (Hz) is a function of the absorption of sound propagating through a medium, including the high frequency dissipation of sound before reflection.

[0158] Cutoff frequency of low-pass filter 54

[0159]

number

[0160] is commonly defined as the frequency at which the ratio of output to input sound pressure amplitudes has a magnitude of 0.707, where:

[0161]

number

[0162] is the sound pressure level (SPL),

[0163]

number

[0164] and a(dB / m)=-3 / r is.

[0165] In one embodiment, the propagation medium for sound in a space is air at a given temperature and humidity. By rearranging the equation for atmospheric absorption of sound, we can compare frequency as a function of attenuation:

[0166]

number

[0167] can be determined as, where p α teeth,

[0168]

number

[0169] is the ambient atmospheric pressure at

[0170]

number

[0171] is the reference ambient atmospheric pressure, T is the ambient atmospheric temperature, and T0 = 293.15K is the reference ambient atmospheric temperature.

[0172]

number

[0173] is the oxygen relaxation frequency,

[0174]

number

[0175] is the nitrogen relaxation frequency,

[0176]

number

[0177] is the molar concentration of water vapor,

[0178]

number

[0179] is the saturated water vapor pressure, where

[0180]

number

[0181] is the relative humidity as a percentage, and T 01 = 273.16 K is the triple point isotherm temperature (Zuckerwar, Meredith, 1984).

[0182] frequency

[0183]

number

[0184] An approximate solution to the equation defining

[0185]

number

[0186] is implicit in the sine wave equation, and the damping a is related to an exponential with a dependence on frequency, which can be derived from first principles: y=A=A0e -a cos(wt) where A is the amplitude in dB, A0 is the initial amplitude in dB, and a is the attenuation coefficient in dB / m.

[0187] a(dB / m) → α a(dB / m)=10log 10 (Pt / Pi) and α=1-(Pt / Pi) Rearranging and analyzing the absorption coefficient as a function of frequency, a polynomial of degree = 3 is found as a minimum. Analyzing both the absorption coefficient and frequency separately gives the absorption coefficient:

[0188]

number

[0189] and the frequency is found as

[0190]

number

[0191] where

[0192]

number

[0193] are the data points. Combining the two equations,

[0194]

number

[0195] is found, where y is a coefficient specific to external variables such as temperature and humidity. To achieve maximum correlation with the absorption coefficient at higher frequencies, the coefficients are averaged as y~5×10 -8 and is found to vary for different conditions of temperature and humidity.

[0196] Absorption coefficient as a function of frequency

[0197]

number

[0198] Rearranging the equation to get

[0199]

number

[0200] and the correction factor y for different temperatures and humidity is, for example,

[0201] [Table 1]

[0202] Each virtual point audio signal component y n The time delay operation 56 for Δt(ms)=(r / v)10 3 The time delay Δt depends on the distance r, as given by n(ms), where v is the speed of sound propagating through a medium, which is 343 m / s in air at a temperature of 20°C and an average humidity of 50%, and r is the distance between the virtual point and the virtual sound source, specifically between the virtual point and the center of the virtual sound source.

[0203] Therefore, for each virtual point of the virtual object, a virtual point specific operation including an attenuation operation 52, a low pass filter operation 54, and a time delay operation 56 is applied to generate a virtual point audio signal component y n is generated. n is similar to the first reflection of a sound originating from a sound source and reflected by a virtual object depending on the position of the virtual object relative to the sound source. n is also determined according to the conditions that affect the propagation of sound through a medium, such as air at a particular temperature and humidity.

[0204] The virtual point audio signal component y(t)_n, also called y(t)_n, resulting from the first reflection operation 12 n are passed as audio signal components to be summed with the symmetric group audio signal resulting from the reverberation operation 18 (i) as shown in FIG. 6C, and (ii) summed to obtain a composite audio signal (see combiner 53).

[0205] The composite audio signal is then multiplied by a factor a=1 / N 2 (see attenuation operation 55), where N is the number of virtual point audio signal components, and the attenuation a(dB)=G(dB) is converted to G→Pt / Pi, where the power ratio is given in dB with respect to the gain G.

[0206] It should be understood that whenever two or more audio signals are summed within the described audio signal process, the above attenuation operation may be applied depending on the number of audio signals summed.

[0207] The composite audio signal is then filtered by performing a second multiband filtering operation 16, which may be an absorption filter as described above. The value selected for the first absorption filter 10 from ISO 354 is preferably the same value for the second absorption filter 16.

[0208] The filtered composite audio signal is then attenuated by performing an attenuation operation 57. This may be understood to be the same operation as operation 55, except that the attenuation is a=1 / K 2 The difference is that it depends on the number of distinct distances in the further set of one or more distances K, such as

[0209] The embodiment further comprises determining one or more distance audio signals d_k, in this example two distance audio signals for each distinct distance in the further set of one or more distances, i.e. a first distance audio signal d_k+ and a second distance audio signal d_k−, as described herein.

[0210] Here, determining the first distance audio signal for the distinct distances includes modifying the composite audio signal by performing a time delay operation 64 that introduces a time delay and a signal feedback operation 58. Determining the first distance audio signal also includes performing an attenuation operation 60 and a low pass filter operation 62.

[0211] In the illustrated embodiment, determining the second distance audio signal for the distinct distances includes modifying the composite audio signal by performing a second time delay operation 72 that introduces a second time delay, a signal inversion operation 68, a signal attenuation operation 68, a low pass filter operation 70, and a second signal feedback operation 66. In principle, operations 64 and 72 are identical, and operations 62 and 70 are identical. It should also be noted that the attenuation performed by each of operations 60 and 68 is also identical, with operation 68 inverting the signal and operation 60 not inverting the signal.

[0212] Performing the signal feedback operation may include, for example, recursively adding the distance audio signal to the input itself before the decay operation, as shown. It should be understood that box 18a may be part of the reverberation module 18 shown in FIG. 2.

[0213] It should be understood that in determining the distance audio signal d_k+−, the attenuation operation, the signal inversion operation (if performed), the low-pass filtering operation, and the time delay operation may be performed in any order, and one or more steps may be omitted, repeated, modified, and / or added. The signal feedback operation is preferably performed last, and the summation of the distance audio signal with the input is preferably performed first.

[0214] Note that in the illustrated embodiment, a 2K distance audio signal is generated.

[0215] The first distance audio signal for the distinct distance may be referred to as a non-inverted distance audio signal, and the second distance audio signal for the distinct distance may be referred to as an inverted distance audio signal.

[0216] The time delay calculation 64 / 72 determines the time delay Δt depending on the distance from which the audio signal is determined. n (ms). This time delay is Δt(ms)=(r / v)10 3 where r is the distance over which the distance audio signal d_k is determined and v is the speed of sound propagating through the medium, for example 343 m / s in air at a temperature of 20° C. and an average humidity of 50%.

[0217] The difference is that the distance at which the distance audio signal is determined should be taken as the distance r, but as explained above with respect to operation 54, the low pass filter operation 62 / 70 constructs an attenuation function distance audio signal depending on the distance at which the distance audio signal d_k is determined and the conditions of sound propagation through the medium.

[0218] The attenuation operation 60 / 68 reduces the time delay Δt introduced by the time delay operation 64 / 72. n Depending on a(dB)=-Δtx where x is the total decay time of the reverberant audio signal Dt(s) and the decay constant e ax is a variable of x=(1 / Dt) / e ax and e ax is the attenuation constant of a wave propagating through a medium per unit distance from the sound source. (Federal Standard 1037C, 1996) It is the real part of the propagation constant and is measured in Neper per meter (Np / m). A Neper is approximately 8.686 dB. The attenuation constant is therefore defined by the amplitude ratio: e ax =A0 / A x =1Np=approx.8.686

[0219] The value Dt (in seconds) determines the actual time over which the total energy of the reverberant audio signal decays. If the lower loudness threshold of hearing is considered standard at -72 dB, then for practical purposes the above formula can be written as x=~1 / 2(1 / Dt) / e ax can be adjusted to

[0220] As a result, the actual amplitude A of each distance audio signal after applying the attenuation a (dB) will vary based on the length of the delay time Δt, with shorter lengths typically resulting in higher initial amplitudes, and may also vary significantly based on the frequencies fn present in the audio input signal, which is the fundamental frequency F and its harmonics which are multiples of integer F (f1, f2, f3, etc.). F(Hz)=1 / (2Δt) This is because, although related to the time delay of the feedback operation as above, the time amplitude M(t) of the feedback of each distance audio signal is equal for each distance audio signal at the loudness threshold of -72 dB, and therefore the magnitude M(t) of each distance audio signal is equal to Dt(s). This condition satisfies the reverberation density, which needs to remain constant throughout the entire time the reverberation decays.

[0221] In one embodiment, the virtual object has the shape of a circular chapel with walls made of limestone. It has been determined that this virtual space has an audible reverberation decay time of approximately 2.2 seconds, and therefore Dt(2.2) is applied to modify all distance audio signals.

[0222] Additionally, the damping calculation 60 / 68 may be further modified depending on the compensation for high frequency dissipation resulting from the absorption of sound propagating through the medium, i.e., during the decay of the reverberation, higher frequencies dissipate relatively faster than lower frequencies, and therefore the shorter the delay time Δt, the more the time amplitude M(t) may be reduced. a(dB)=x-(Δt0 / Δtn)x where x is a variable function of the total decay time Dt and Δt0 is the reference time delay, which is the longest time delay in the system, i.e., the longest time delay used in block 18a shown in Figure 6A. As a result, the attenuation of a delay line with time delay Δtn = Δt0 is equal to 0.

[0223] Combining the two equations, a(dB)=x-(Δt n x)-(Δt0 / Δt n) x This becomes:

[0224] Thus, by applying attenuation, low pass filtering, and time delay operations, audio signal components are produced that, once summed, resemble the coherent reverberation of sound sources within spaces and / or objects of distinct shapes, sizes, and materials, according to conditions affecting the propagation of sound through a medium, such as air at a particular temperature and humidity.

[0225] 6B is a flowchart illustrating how the symmetry group audio signal s_m+- is determined based on the distance audio signal d_k+- according to one embodiment. This embodiment includes determining a first symmetry group audio signal and a second symmetry group audio signal for each symmetry group m. The determining the first and second symmetry group audio signals further includes selecting distance audio signals from all pairs of the first and second distance audio signals, each pair determined for a respective distance from a set of one or more symmetry group distances associated with the symmetry group in question, and combining the selected distance audio signals to determine the first symmetry group audio signal and combining the unselected distance audio signals from all pairs of the first and second distance audio signals to determine the second symmetry group audio signal.

[0226] To illustrate, for the determination of s_1-, the distance audio signal d_1- and the distance audio signal d_2+ (and other signals) are combined. Thus, for the determination of s_1+, the distance audio signal d_1+ and the distance audio signal d_2- (and other signals) are combined. Note that this means that the distance associated with the distance audio signal d_1 and the distance associated with the distance audio signal d_2 are within a set of one or more symmetry group distances for the group m=1, where s_1- and s_1+ are symmetry group audio signals. If these distances are not within the set of symmetry group distances for the symmetry group m=1, they are not added to s_1- or s_1+.

[0227] In one embodiment, the virtual object is a 40x40x40m hollow cube with a grid k (=3) that provides 54 points defined on the shape, i.e., a "virtual point resolution" of 3x3 points evenly distributed on the cube's six surfaces. In this embodiment, the one or more further sets of distances include 18 distinct distances. Each distinct distance is associated with a unique delay time Δt and is generated relative to the propagation of sound at a speed of sound of 343 m / s, i.e., through air at a temperature of 20°C and an average humidity of 50%. Thus, each pair of distance audio signals is associated with a distinct distance within the one or more further sets of distances. As a reminder, the one or more further sets of distances referred to in this disclosure include all distances within all sets of one or more symmetry group distances. Each set of one or more symmetry group distances is associated with a symmetry group. In this embodiment, each virtual point of the virtual object belongs to one of three symmetry groups. Thus, the audio distribution matrix resulting from performing a summation of the inverted and non-inverted versions of the distance audio signals for the cube may be defined according to the following table: Here, the columns indicate symmetry group signals 1.1 (s_1−), 1.2 (s_1+) for symmetry group 1, 2.1 (s_2−), 2.2 (s_2+) for symmetry group 2, and 3.1 (s_M−), 3.2 (s_M+) for symmetry group 3, where the symmetry group quantity in this embodiment is M=3. Each row of this column is associated with a distinct distance within a further set of one or more distances, where the distances are indicated by an associated time delay.

[0228] [Table 2]

[0229] Thus, this table should be read such that, to determine symmetric group audio signal 1.1, an inverted version of the distance audio signal for distance "27.487 ms", a non-inverted version of the distance audio signal for distance "38.873 ms", an inverted version of the distance audio signal for distance "47.609", etc. are added together to form symmetric group signal 1.1.

[0230] It should be noted that the method for determining a symmetry group audio signal based on a distance audio signal, as shown in FIG. 6B, is preferably implemented in the reverberation module 18 shown in FIG.

[0231] FIG. 6C is a detailed flowchart illustrating how the reverberant audio signal is determined based on the symmetric group audio signal s_m and the virtual point audio signal component y_n.

[0232] Optionally, determining the reverberant audio signal comprises combining the symmetry group audio signal with the virtual point audio signal component. Optionally, such combining comprises determining a modified audio signal component, wherein determining the modified audio signal component comprises adding a symmetry group audio signal of the symmetry group in question to each virtual point audio signal component determined for the virtual point belonging to the symmetry group. Illustratively, in FIG. 6C , a modified audio signal component y′ is thus obtained for each virtual point of the virtual object. n Optionally, the corrected component y' is obtained. n Determining {overscore (x)} may include, for example, attenuating the virtual point audio signal component and / or the symmetric group audio signal before adding the symmetric group audio signal to the virtual point audio signal component as shown in FIG. 6C.

[0233] Optional attenuation of the virtual point audio signal component is controlled by variable parameter a, which includes a gain (dB) scaled from 0 to 1 (∞ to 0 dB line out). Optional attenuation of the symmetric group audio signal is controlled by variable parameter b, which includes a gain (dB) scaled from 0 to 1. This provides the user with optional control to independently adjust the audio output level of each of the first reflections of the sound source and the reverberation. It should be understood that additional multipliers to attenuate or amplify the gain of the audio signal may be added at any point in the signal processes described in Figures 6A-6C.

[0234] It should be understood that each modified audio signal component y'_n obtained after summation, eg combination, of the virtual point audio signal component and the symmetry group audio signal is associated with a virtual point of the virtual object.

[0235] Optionally, each audio signal obtained after adding the appropriate symmetry group audio signal to the virtual point audio signal component is further modified. Thus, in such an embodiment, determining the modified audio signal component may also include further modifications by a resonance module, a depth module, a height module, and / or a distance module, such as those present in module 80 (see FIG. 2), as described above with reference to FIG.

[0236] The finally obtained modified audio signal components y′, each associated with a virtual point of the virtual object, n are inputs to a panning system, such as a panning matrix 28, for distributing the modified audio signal components y'n to form individual audio output signals z_p, each associated with a speaker p in the speaker configuration. Figure 17 shows a detailed embodiment of a panning system.

[0237] An advantage of applying the computations as part of the reverberation computation 18, resonance computation 20, depth computation 22, height computation 24, distance computation 26, and panning matrix computation 28 is that coherent sound projections of sound sources reflecting off virtual objects are generated, with both the sound sources and virtual objects being independently controllable and scalable with respect to the actual output medium, i.e., speakers. This means that the experience of a dynamic virtual space that can be configured by defining multiple virtual objects as described in this disclosure and that a listener can navigate through and auditorily explore is, in principle, independent of the quantity and configuration of speakers. Thus, sound sources and virtual objects can be scaled, rotated, tilted, and (parts of) the space can be expanded in close-up or moved far away, and can be positioned at any distance, height, and depth with respect to the observer and with respect to another virtual object and / or virtual sound source / virtual object, without the need to reconfigure speakers.

[0238] FIG. 7 illustrates an embodiment in which the speaker configuration is a mono sound system, i.e., a speaker system with one individual output channel. In this case, in the reverberation calculation 18 as described in FIG. 6A, only an inverted or non-inverted version of each distance audio signal is generated based on the composite audio signal. In such an embodiment, determining the symmetric group audio signal may be omitted in its entirety. Instead, as described with reference to FIGS. 2, 5, and 6A, all determined distance audio signals d_k resulting from the reverberation calculation 18a are summed and optionally attenuated or amplified, and all virtual point audio signal components y_n resulting from the first reflection calculation 12 are summed, attenuated using an equation depending on the number of summed audio signals, optionally further attenuated or amplified, and summed to the audio output of the speaker together with the audio signal resulting from summing the distance audio signals and together with an audio input signal x, which may also serve as an input signal for further reverberation audio signal determination methods.

[0239] FIG. 8 illustrates an embodiment in which the speaker configuration is a stereo sound system, i.e., a speaker system having two separate output channels with a left (L) 30a and a right (R) 30b speaker setup for the left and right ears of a centrally located (virtual) listener. In one embodiment, such a system may be a pair of headphones. In this case, for the reverberation calculation as described in FIG. 6A, both inverted and non-inverted versions of each distance audio signal are determined. In this embodiment, the determination of symmetric group audio signals may be omitted in its entirety. Instead, the output of a first inverted version of a delay line is summed with a non-inverted version of a second delay line, an inverted version of a third delay line, etc. to form a separate left output signal (L), and the output of the first non-inverted version of a delay line is summed with an inverted version of a second delay line, a non-inverted version of a third delay line, etc. to form a separate right output signal (R). With respect to the R output, L ≠ R output, and both the L and R output signals may be optionally further attenuated or amplified before being fed to the L and R speakers.

[0240] In this example, the audio input signal x, which is the input audio signal for the reverberant audio signal determination method described herein, is defined as a stereo signal having an L and R channel, where either L=R or L≠R can be true. Both the L and R output signals of the initial audio input signal can be optionally further attenuated or amplified before being fed to the L and R speakers.

[0241] As a result, for every virtual point audio signal component resulting from the first reflection operation 12, an L and R version of the delay line is generated for each delay time Δt, where either L=R or L≠R can be true. Then, all L versions of the virtual point audio signal components resulting from the first reflection operation 12 are summed, and all R versions are summed, and both L and R are attenuated using an equation depending on the number of audio signals summed, and optionally further attenuated or amplified before being fed to the L and R speakers.

[0242] It should be understood that Figures 7 and 8 represent possible embodiments illustrating the application of the described invention with respect to backward compatibility with existing audio standards. Any changes in output routing with respect to the described audio signal processing adjusted to prior art standards should be considered to be included herein.

[0243] The virtual points of the virtual object belong to a symmetry group of virtual points which may be determined as described in Figure 9. The virtual points are preferably distributed equally throughout and / or on the virtual object.

[0244] A number N of virtual points may be defined, and the virtual points may be understood to define a virtual object. In one embodiment, the virtual object is a square plate, as shown in FIG. 9, which is a two-dimensional shape with N=25 evenly distributed virtual points defined on the shape. In this example, the center point of the square plate coincides with virtual point #13. Being a square plate, a 90 degree rotation about the center point will again result in the square plate having the same configuration, i.e., the same position and orientation.

[0245] The diagram in the upper right corner of Figure 9 indicates which symmetry group each imaginary point belongs to; i.e., the numbers between the parentheses indicate the symmetry group of the imaginary point. In this example, there are six symmetry groups. As can be seen in Figure 9, a single point, imaginary point #13, which in Figure 9 belongs to symmetry group g6, may form its own symmetry group in one embodiment. For non-geometric or more irregularly shaped shapes, the lack of rotational symmetry may, in one embodiment, allow several or many single points to form their own symmetry group. Thus, geometric or regular polygonal shapes tend to have fewer symmetry groups containing many points, while irregular shapes tend to have more symmetry groups containing fewer points, i.e., a minimum of one point.

[0246] Therefore, every virtual point n defined on the shape belongs to one symmetry group g, which means that virtual points falling within the same symmetry group will have the same reverberation conditions in the virtual object, and virtual points falling within this symmetry group will share the same set of distances with any other point defined on the shape.

[0247] Each symmetry group is associated with a set of one or more symmetry group distances. Figure 9 shows the associated symmetry group distances for each of the symmetry groups 1, 2, 3, 4, 5, and 6.

[0248] As can be seen from Figure 9, the symmetric group g n is a distance r that is greater or less than the other groups. g(n)<->n and many groups may have the same distance r n<->n In the example of the square plate of Figure 9, there are seven distinct distances within the further set of one or more distances. These seven distances within the further set of one or more distances are shown in Figure 10.

[0249] In one embodiment, the shape is a square plate with a grid k (=5) and sides of length l (m), where L = √2l = 1 m, and sound propagates through the plate at a velocity v = 343 m / s. As introduced in Figure 9, there are a total of seven distinct distances in the set of one or more additional distances. Each distinct distance is separated by a time delay Δt (ms) = (rn<->n / v) 10 3 is associated with.

[0250] With respect to the number of virtual points defined for a virtual object, in accordance with the example set forth in FIG. 10, the described method comprises a highly efficient way to reduce the computational overhead required to generate a reverberant audio signal having the desired coherence, density, and smoothness without losing the essential information inherent in a virtual object of a particular shape and / or particular size and materiality.

[0251] For example, a straight-line signal path for each distance between all existing pairs of virtual points of a virtual object, such as a rectangular plate with 25 virtual points as introduced in FIG. 9, theoretically results in a total of 25 × 24 = 600 distances associated with 600 distance audio signals, or as many as 1200 distance audio signals if two distance audio signals are determined for each distance, to generate the first reflection of reverberation after the first reflection from the virtual sound source to the virtual object. Furthermore, the sum of the distance audio signals does not satisfy the criterion of smoothness of the reverberation tail, i.e., the reverberation tail generated by a feedback operation as described in FIG. 6A or, more generally, a feedback delay network (FDN) as is commonly used to generate reverberation in the prior art. The criterion of smoothness of the reverberation modal density is generally described as resembling a pink noise signal, in which higher frequencies dissipate faster than lower frequencies in the time dimension. Since many identical distances and integer multiple distances are found in the set of distances generated in this way, this results in a certain frequency F dominating, upsetting the balance of the smoothness of the tail across the entire frequency range. Therefore, such techniques for generating reverberation tails are ignored in prior art methods because they do not meet certain conditions.

[0252] Instead, following the proposed method described with reference to FIG. 9, in order to ensure the desired smoothness of the reverberation while maintaining those distance values ​​to generate a reverberant audio signal containing essential information specific to the shape of the virtual object, the number of signal paths required is significantly reduced by a factor of 85 (600:7), to only seven distinct distances, for which seven distance audio signals would be determined in comparison, or 14 distance audio signals if two distance audio signals are determined for each distinct distance.

[0253] In the embodiment of FIG. 10, a simple distribution for each of the 25 points of the seven distinct distances obtained together with the distances associated with the delay times is calculated as follows for the distance audio signal: 25(×Δt1)+25(×Δt2)+25(×Δt3)+24(×Δt4)+24(×Δt5)+16(×Δt6)+12(×Δt7)=151 This results in a signal path of

[0254] Alternatively, by including virtual points within the symmetry group and associating distance audio signals with the symmetry group rather than directly with each virtual point, the number of required signal paths can be further reduced. 6(×Δt1)+6(×Δt2)+6(×Δt3)+5(×Δt4)+5(×Δt5)+3(×Δt6)+2(×Δt7)=33 signal paths+ 4(×g1)+8(×g2)+4(×g3)+4(×g4)+4(×g5)+4(×g6)=61 signal paths

[0255] Therefore, the amount of signal paths required for reverberation is further reduced to X2.5 (151:61) by introducing an intermediate step of summing the distance audio signals over the symmetry group as described in Figure 10 without any loss of quantity / information in the resulting signal paths after the operation.

[0256] The method described herein therefore comprises an efficient solution for optimizing the required power consumption by using a minimum of processing data and signal paths, and a novel method that meets known criteria for density and smoothness of high quality reverberant signals, while at the same time introducing new qualities to the reverberant signal in terms of expressing its shape and materiality that are not achieved by methods for generating artificial reverberation known in the art.

[0257] It can be seen from the above that as the virtual point resolution, i.e., the number of defined points per shape, increases, the amount of distinct distances in the further set or sets of distances in the reverberation calculation 18 also increases, and more distance audio signals should be generated. This in turn increases the density of the reverberation, i.e., the "time density" comprising the amount of echoes per second, and the "modal density" associated with frequencies F (f1, f2, f3, etc.) across the frequency range, where each unique F is considered to be the result of delay times Δt that are not integer multiples of each other.

[0258] The present invention further includes a novel method for optimizing modal density of a virtual shaped reverberation with a balanced distribution of extremes (+ / -) for a specific distance rn<->n.

[0259] The inverted, time-delayed feedback signal amplifies the odd harmonics of F (f1, f3, f5, etc.) based on the time delay Δt, while the non-inverted, time-delayed signal with feedback amplifies the even harmonics of the same F (f2, f4, f6, etc.) based on the same time delay. Thus, by distributing the signal as symmetric even and odd harmonics of the exact same harmonic series, which are the resonant components resulting from the shape of the reverberant space, the modal density of the resulting reverberation increases by a factor of two for the same number of distances as found for the same virtual point resolution of the shape.

[0260] In one embodiment, the virtual object is a square plate with four virtual points and therefore one symmetry group. For this one symmetry group, there are two symmetry group distances. For each symmetry group distance, two distance audio signals are determined: an inverted version and a non-inverted version.

[0261] Since several perpendicular or parallel distances of the same length may connect a single point, a chessboard-like distribution of polar codes is proposed to achieve an optimal symmetric spread of the poles across the complexity of points on the shape, as also explained with reference to Figure 6B.

[0262] In an embodiment according to FIG. 11A, this gives the symmetric group audio signals shown in the table below. Each row in this table is associated with a distinct distance in one or more further sets of distances. In the columns, the distances are indicated by the associated time delays. Furthermore, each column is associated with a symmetric group signal.

[0263] [Table 3]

[0264] Thus, to determine the symmetry group audio signal for symmetry group 1.1, an inverted version of the distance audio signal for distance Δt1 and a non-inverted version of the distance audio signal for distance Δt2 are added together. Further, to determine the symmetry group audio signal for symmetry group 1.2, a non-inverted version of the distance audio signal for distance Δt1 and an inverted version of the distance audio signal for distance Δt2 are added together.

[0265] In one embodiment, the shape is a plate of equal sides but unequal length and width, with four imaginary points and two symmetry groups. Three distinct distances are found in one or more further sets of distances, and for each distinct distance, two distance audio signals are determined. The distance r according to FIG. 11B n-<->n+ The equal distribution across both poles of

[0266] [Table 4]

[0267] Give.

[0268] In one embodiment, the shape is a plate with unequal sides, length, and width, with four imaginary points and four symmetry groups. In this example, each imaginary point defined on the shape is its own symmetry group. Therefore, a symmetry group has only one version of itself instead of two. In a further set of one or more distances, six unique distances are found. The distance r from Figure 11C n-<->n+ The equal distribution across both poles of

[0269] [Table 5]

[0270] Give.

[0271] It should be noted that in this embodiment, the distribution deviates from a simple chessboard-like pattern for Δt2→g3 and Δt3→g4, since there are only symmetrically opposed points between groups rather than points falling within exactly the same group.

[0272] Figure 12 is a flow process for a "value filter" operation 34 to obtain the desired time delay depending on the distance between virtual points as described in Figure 9, and a "time density scaler" operation 36 that automates the optimization of the desired time density of the reverberation based on a variable threshold. The steps up to the step "Determine Symmetry Group" may be performed in the value filter operation 34, and the steps following this step in each iteration may be performed in the time density scaler operation 36.

[0273] In particular, determining the location of the virtual point may include receiving grid and shape data, the output of which is an NxN dimensional array, each element of which holds a distance value from the virtual point to every other virtual point.

[0274] The ordering logic of the generated virtual points constitutes a right-handed coordinate system. const{virtual points}=pd.get virtual points Distances()

[0275] The virtual point numbering and coordinates (x,y,z) associated with the virtual point position in virtual space can be generated from a custom shape script file. A simple virtual object script file contains how vec3 objects are specified. # ... def getPositions(self, density, hollow, speedAdjustedTime): positions = [] #pyramid base 50×50 positions.append(nap.vec3(-25., 0., -25.)) positions.append(nap.vec3(0., 0., -25.)) positions.append(nap.vec3( 25., 0., -25.)) # ... positions.append(nap.vec3(12.5, 17.67767, 12.5)) positions.append(nap.vec3(0., 35.355339, 0.)) return positions # ...

[0276] The computer program can then extract the values ​​in the function calls to generate the virtual point coordinates. const config={ type: "scriptFile", name: "pyramid", fileName: "scripts / pyramidShape" } const pd2 = new virtual points Distances()

[0277] In the next step, the distance between pairs of virtual points is calculated. Since each virtual point has an (x,y,z) coordinate, the standard distance equation is used. r=√((x2-x1) 2 +(y2-y1) 2 +(z2-z1) 2 )

[0278] The time delay is then converted from distance based on receiving the value of the speed of sound (v). For a virtual object with N virtual points, the result is an N x N dimensional matrix, where the entries e ij is the virtual point p i and p j For a 1x1 meter plate with 3x3 imaginary points evenly distributed on its surface, this is

[0279] [Table 6]

[0280] Give.

[0281] This initial matrix may be obtained by defining, for each virtual point of the plurality of virtual points, a set of one or more virtual distances that includes a respective virtual distance between the virtual point in question and each other virtual point of the plurality of virtual points. In this example, the set of one or more virtual distances associated with virtual point #1 is shown in row "1", the set of one or more virtual distances associated with virtual point #2 is shown in row "2", etc.

[0282] The initial matrix is ​​then passed through a multiplicity filter to remove some of the delay times. The filter can be applied to each virtual point, where one virtual point is represented by each horizontal row in the matrix. Because some of the generated time delay values ​​may be close to each other, equality is determined by taking the absolute value of the delay time difference according to the sample rate and checking whether the difference is less than one sample time. Therefore, such nearly identical distances are not treated as separate distances. function equals(a, b, sampleRate) { return Math.abs(a -b) < (1 / sampleRate) }

[0283] [Table 7]

[0284] In the first instance of the filter applied to the virtual points, we are only interested in the first occurrence of the time delay, and all overlaps are filtered out. In this example, the distance between virtual points p1 and p4 is the same as the distance between virtual points 1 and 2, so it is filtered out. Note that the virtual distance between points 1 and 2 is maintained.

[0285] In the second instance of the filter, delay times, i.e., distances, that are integer multiples of other delay times are also filtered. In this example, the distance between virtual points p1 and p3 is filtered because it is twice the value of the distance between virtual points p1 and p2.

[0286] After this filtering step, for each virtual point, a set of one or more virtual point-specific distances is obtained.

[0287] Such a virtual point-specific set may also be arrived at by first, for each set of one or more distances associated with a virtual point, i.e., for each row in the initial matrix, removing distances that are integer multiples of any other distance within the set, i.e., within the row in question, to obtain a further set of one or more distances associated with the virtual point. A virtual point-specific set may then be determined as the distinct distances within each further set. Illustratively, the virtual point-specific set for point 1 includes 0.972, 1.374, and 2.173.

[0288] After obtaining the delay times for each virtual point that satisfies the filter, the set of values ​​in the matrix are sorted into ascending order of the distinct delay times, i.e., distinct distances, from shortest to longest delay lines. These distinct delay times are also referred to herein as distinct distances of one or more further sets of distances. After these distinct distances are found, other properties, such as the occurrence of the delay times associated with the distinct distances per symmetry group, can be analyzed.

[0289] Virtual points that have the same set of virtual-point-specific distances are determined to belong to a symmetry group. For example, point 3 has 0.972, 1.374, and 2.173 as the set of virtual-point-specific distances that are the same as the set of virtual-point-specific distances for point 1. Therefore, points 1 and 3 belong to the same symmetry group.

[0290] In this example, the time delay of 0.972 milliseconds (i.e., the virtual distance of 0.333 m) is associated with two symmetry groups: symmetry group 1, which is composed of virtual points 1-, 2+, 3-, 4+, 6-, 7+, 8-, 9+, and symmetry group 2, which is composed of virtual point 5-. The time delay of 1.374 ms (i.e., virtual distance 0.471 m) is associated with two symmetry groups: object group 1, consisting of virtual points 1+, 2-, 3+, 4-, 6+, 7-, 8+, 9-, and object group 2, consisting of virtual point 5+. A time delay of 2.173 ms (i.e., a virtual distance of 0.745 m) is associated with one symmetry group, namely object group 1, which is composed of virtual points 1-, 2+, 3-, 4+, 6-, 7+, 8-, 9+.

[0291] A new matrix is ​​created in which all symmetry groups are arranged in two versions on the horizontal axis: 1.1, 1.2, 2.1, 2.2, 3.1, 3.2, etc. If a group consists of only one virtual point, only one version of the group exists, and therefore the group appears only once in the matrix. On the vertical axis, all distinct distances in one or more further sets of distances formed by all symmetry group distances, where distinct distances are represented as time delays, are sorted from shortest (top) to longest (bottom). Then, both extremes (+ / -) are added to the matrix in a simple chessboard-like fashion. Finally, the virtual points in each symmetry group are distributed alternately between group n.1 and group n.2 according to the process described in Figures 11A-11C.

[0292] The second step in the command chain involves a "time density scaler" operation, which refers to building block 36 in FIG. 3. The time delay values ​​for the symmetry groups of the virtual points may be performed as follows: Each symmetry group of the virtual point is associated with one or more symmetry group distances. Furthermore, for each symmetry group distance, one or more distance audio signals are determined using a time delay operation (see 64 / 72 in FIG. 6A). The time delay value, also called the density index, for a particular symmetry group is: di=Σ{1 / Δt1,1 / Δt2,...,1 / Δt Q} where Δt1 is the time delay (in seconds) introduced by the time delay operation for determining one or more distance audio signals for a first one of the symmetry group distances, Δt2 is the time delay introduced by the time delay operation for determining one or more distance audio signals for a second one of the symmetry group distances, etc. Q denotes the number of symmetry group distances for the symmetry group for which the density index is determined.

[0293] The given time density for each symmetry group is compared to a variable threshold: if the calculated value for d is lower than the threshold, the filter is not satisfied, instructing the process to increase the virtual point resolution and triggering a re-execution of the command chain until the density index for each symmetry group is equal to or greater than the threshold, thus satisfying the filter.

[0294] The amount of echoes per second required to achieve sufficient reverberation density is generally 1000 s -1 10,000 seconds depending on the type and nature of the audio input signal to the audio signal process. -1 The actual di is further affected by the sum of all audio output signals at the loudspeaker as shown in Figure 6C and the amount of virtual point audio signal components generated in the first reflection operation 12, which increases the actual di x N (= the number of delay lines in 12). Therefore, the time density threshold needs to be a variable parameter that can be adjusted by the user according to different situations.

[0295] In this way, the system for generating reverberation within a virtual object is automatically adjusted to the most optimal density conditions for a given virtual object in given conditions. In prior art applications of artificial reverberation systems, parameters such as delay time are carefully selected to meet a time density criterion. The present invention provides a novel method for determining optimal time density without requiring a set of fixed values ​​such as delay time selected in the system in advance; instead, such values ​​may depend on geometric attributes, i.e., dimensional shape, size, and materiality, of the reverberation within the virtual object.

[0296] Specifically, Figure 13 shows the high-frequency and ultra-high-frequency cutoff values ​​f that can occur depending on the reverberation settings, including the speed of sound through the medium, temperature, humidity, and other factors. c, which shows a flow process for a "sample rate interpolation" operation 32 to obtain the desired low pass filtering included within the first reflection operation 12 and the reverberation operation 18, depending on the size of the virtual object and the virtual point resolution that determine the scale of the distance between the virtual points. The value f obtained as part of the delay line implemented in either the first reflection operation 12 or the reverberation operation 18 is c may be (much) above the threshold of the human audible frequency range (~20 kHz), and more specifically, may be greater than the Nyquist frequency (= 0.5 x sample rate).

[0297] The Nyquist frequency essentially determines the upper threshold for the assignable cutoff frequency in a digital low-pass filter, meaning that filtering above Nyquist has no audible effect. Nevertheless, f above the Nyquist frequency, as shown in Figure 6A, c The effect of the distance dependent attenuation function, which may include values ​​of , may also involve significant audible effects from the attenuation of frequencies below the Nyquist frequency. As an approach to optimizing the attenuation functions in the first reflection and reverberation operations, a command flow is proposed that takes into account the relationship between the sample rate of an audio output device connected to (or part of) a computer processing unit executing a program or code portion and the generated frequency cut-off values ​​in the first reflection operation 12 and the reverberation operation 18, by locally increasing the sample rate to complete the desired filtering operation and then decreasing the sample rate to the sample rate of the audio output device by sample interpolation.

[0298] As a first step, all necessary distances are calculated between the input source and each virtual point in the case of a first reflection calculation, or between all virtual points in the case of a reverberation calculation, as described in Figure 12. Then, f is calculated as described in detail in Figure 6A with reference to operation 54. c can be determined.

[0299] The second step involves a first filter to determine if further action is required regarding sample interpolation. c If the value obtained for is greater than the Nyquist frequency, it does not fill the filter and commands a process to locally increase the sample rate by interpolation until it fills the filter, and commands a low-pass filtering process to be performed at the locally optimized sample rate.

[0300] After the low-pass filtering operation is completed, the second filter checks whether the local sample rate matches the sample rate of the audio output device. If it does not satisfy the filter, it will reduce the local sample rate by interpolation until the sample matches the sample rate of the audio output device.

[0301] As a result, the effect of the frequency cutoff on frequencies within the human hearing range is accurately encoded in the signal after sample interpolation, even though the frequency cutoff in question is (well) above the human hearing range. Applicant has found that this has a significant impact on the accuracy and smoothness of the high frequency decay that is constructed in reverberant audio signals as described.

[0302] Figure 14 illustrates a user interface according to one embodiment of the present invention.One embodiment of a method includes generating a user interface as described in the present invention.

[0303] In one embodiment, a user interface for the described system comprises modules for controlling the virtual object, e.g., its position relative to the observer and / or "eye point", the shape of the virtual object, the materials comprising the virtual object, the conditions of the medium selected for sound propagation, certain other attributes such as the attributes of the reverberation itself, resonances resulting from standing waves within a virtual object of a particular shape, as well as modules for controlling the audio output of the audio signal process, including the master output level and send levels of the audio output signal or "audio mixdown" from the audio signal process to provide as input audio signals to other audio signal processes determined for other virtual objects.

[0304] The illustrated user interface comprises an input section that allows a user to control audio output signals or audio mixdowns from other audio signal processes determined for other virtual objects using input channels as input audio signals for the audio signal process determined for said virtual object. The input channels may consist of multiple audio channels that receive audio signals from audio signal processes determined for other virtual objects that are combined together as input audio signals for the audio signal process determined for said virtual object, either by optionally performing the described method or by an external audio source. The user interface allows a user to control the amplification of each input channel, for example, by using a gain knob.

[0305] The user interface may further comprise an output section that enables a user to route a summed audio output signal or an audio mixdown of the audio signal processes determined for the virtual object as an input audio signal to determine audio signal processes for other virtual objects.

[0306] The output module may further comprise a master level fader that may determine the level of optional attenuation (value "a") of the audio output signal of the audio signal process that is fed to the separate speakers, as described in FIG.

[0307] The user interface may further comprise a virtual object definition section that allows the user to input parameters related to the virtual object, such as its shape, for example by selecting a shape via a drop-down menu, and / or whether the virtual object is hollow or solid via an on / off button, and / or adjust the scale, i.e., size, of the virtual object via a knob, and / or input its dimensions, e.g., its Cartesian dimensions via numeric boxes for dimensions x, y, and z, and / or input its rotation, and / or input a resolution for determining the amount, i.e., density, of virtual points defined on the shape of the virtual object via numeric boxes. This allows the user to control the amount of computation required in the audio signal processing determined for the virtual object.

[0308] The input means for inputting the parameters related to the rotation may be presented as endless rotary knobs for the dimensions x, y and z.

[0309] The user interface may further comprise a position section that allows the user to input parameters related to the position of the virtual object. The position of a shape in three-dimensional space may be expressed in Cartesian coordinates + / - x, y, z, with the virtual center of the space designated as 0, 0, 0, and may be presented as a visual three-dimensional field in which the virtual object can be placed and moved. This three-dimensional control field may be scaled in size by adjusting the radius of the field.

[0310] Thus, the separate audio output signals for each speaker resulting from the reverberation audio signal process determined for the virtual object can be automatically controlled by i) modeling the shape of the virtual object, ii) the rotation of the shape in three-dimensional space, and iii) the position of the shape in three-dimensional space.

[0311] The user interface may further include an attributes section that allows the user to control various parameters, such as knobs for adjusting the bandwidth and amount of resonance, which determines the attenuation (value "b") of the optional feedback signal in resonance calculation 20 as described in FIG. 16A; a knob for scaling the perceived distance, which determines the multiplication factor x in the equation that determines the attenuation calculation of distance calculation 26 as described in FIG. 16D; a knob for scaling the perceived height, which determines the multiplication factor x in the equation that determines the attenuation calculation of either depth calculation 22 as described in FIG. 16B or height calculation 24 as described in FIG. 16C; and a knob for scaling the amount of Doppler effect, which is a scaler for modifying the equation that determines the second time delay of distance calculation 26 as described in FIG. 16D.

[0312] The user interface may further include a section for selecting the material of the virtual object via a drop-down menu with several pre-programmed options. The material selection then determines the selected set of ISO 354 values ​​in the absorption filter operation as described in FIG. 6A. The "Absorption knob" and "Reflectance knob" provide a proportional scaler of the absorption coefficient from the selected ISO 354 value to increase or decrease the absorption characteristics of the selected material so that the resulting reflections and reverberations constitute less absorption and more dense reflections of sound.

[0313] The user interface may further include a section for controlling the conditions of the selected medium via a drop-down menu with several pre-programmed options. The selection of the medium, which in one embodiment may be air, may configure several custom options specific to the medium, considered parameters of the sound propagation behavior in that medium, such as a numeric box for setting the temperature in °C and a knob for increasing / decreasing humidity. The speed of sound is a value resulting from the selection of the medium and related parameters, but can be manually adjusted in a controllable numeric box to deviate from the calculated standard. The parameter values ​​set in the conditions section determine the frequency-dependent attenuation of the low-pass filter operation, the time delay in the time delay operations of the first reflection operation 12 and the reverberation operation 18 as described in FIG. 6A, the calculation of the time delay in the value filtering operation as described in FIG. 12, and the frequency cutoff associated with the Nyquist frequency as described in FIG. 13.

[0314] The user interface may further comprise sections for controlling attributes of the reverberation, such as a knob for controlling the output gain of the first reflection, which determines the level of optional attenuation (value "a") of the audio signal components resulting from the first reflection operation 12 as described in FIG. 6C; a knob for controlling the output gain of the reverberation tail, which determines the level of optional attenuation (value "b") of the audio signal components resulting from the reverberation operation 18 as described in FIG. 6C; a knob for controlling the reverberation decay time, which determines the coefficient x in the equation that determines the decay operation of the reverberation operation 18 as described in FIG. 6A; a knob for controlling the reverberation decay, which may modify and / or scale the coefficient y in the equation that determines the frequency cutoff of the first reflection 12 and / or the low-pass filtering operation as part of the reverberation operation 18, and further modify, i.e., increase or decrease, the effect of the correction formula for high frequency dissipation used in the decay operation of the reverberation operation 18 as described in FIG. 6A; and a knob for controlling the reverberation density, which sets a time density threshold used to automate the adjustment of the reverberation system to an optimal time density as described in FIG. 12.

[0315] User input received via the user interface can be used to determine appropriate values ​​for the parameters according to the methods described herein, thus translating all functional operations of the reverberation system into audible manipulation of sound sources reverberating within a virtual space with front-end user characteristics, i.e., dimensional shape, size, and materiality.

[0316] It should be understood that application of the present invention is in no way limited to the layout of this particular interface example, is subject to numerous approaches in system design, can include numerous levels of control for shaping and positioning sound sources within virtual space, and is not limited to any particular platform, medium, or visual design and / or layout.

[0317] FIG. 15 shows a block diagram illustrating a data processing system according to one embodiment.

[0318] 15, data processing system 100 may include at least one processor 102 coupled to memory elements 104 via a system bus 106. As such, the data processing system may store program code in memory elements 104. Furthermore, processor 102 may execute program code accessed from memory elements 104 via system bus 106. In one aspect, the data processing system may be implemented as a computer suitable for storing and / or executing program code. However, it should be understood that data processing system 100 may be implemented in the form of any system including a processor and memory capable of performing the functions described herein.

[0319] The memory elements 104 may include one or more physical memory devices, such as, for example, a local memory 108 and one or more mass storage devices 110. Local memory may refer to random access memory or other non-persistent memory devices typically used during the actual execution of program code. The mass storage devices may be implemented as hard drives or other persistent data storage devices. The processing system 100 may also include one or more cache memories (not shown) that provide temporary storage of at least some program code to reduce the number of times the program code must be retrieved from the mass storage device 110 during execution.

[0320] Input / output (I / O) devices, depicted as input devices 112 and output devices 114, may optionally be coupled to the data processing system. Examples of input devices may include, but are not limited to, a keyboard, a pointing device such as a mouse, a touch-sensitive display, etc. Examples of output devices may include, but are not limited to, a monitor or display, speakers, etc. The input and / or output devices may be coupled to the data processing system directly or through intervening I / O controllers.

[0321] In one embodiment, the input and output devices may be implemented as a hybrid input / output device (indicated in FIG. 15 by the dashed line surrounding input device 112 and output device 114). One example of such a hybrid device is a touch-sensitive display, sometimes referred to as a "touchscreen display" or simply a "touchscreen." In such an embodiment, input to the device may be provided by the movement of a physical object, such as a stylus or a user's finger, on or near the touchscreen display.

[0322] Network adapters 116 may also be coupled to the data processing system to enable it to be coupled to other systems, computer systems, remote network devices, and / or remote storage devices through intervening private or public networks. A network adapter may comprise a data receiver for receiving data transmitted by the systems, devices, and / or networks to data processing system 100, and a data transmitter for transmitting data from the data processing system to the systems, devices, and / or networks. Modems, cable modems, and Ethernet cards are examples of various types of network adapters that may be used with data processing system 100.

[0323] As shown in Figure 15, memory element 104 may store application 118. In various embodiments, application 118 may be stored in local memory 108, in one or more mass storage devices 110, or separately from local memory and mass storage devices. It should be understood that data processing system 100 may further execute an operating system (not shown in Figure 15), which may facilitate the execution of application 118. Application 118, implemented in the form of executable program code, may be executed by data processing system 100, for example, by processor 102. In response to executing the application, data processing system 100 may be configured to perform one or more operations or method steps described herein.

[0324] In one embodiment of the present invention, the data processing system 100 may represent a first reflection module 12 and / or an absorption filter 16 and / or a reverberation module 18 and / or a resonance module 20 and / or a depth module 22 and / or a height module 24 and / or a distance module 26 and / or a panning system 28 as described herein.

[0325] Additionally, data processing system 100 may represent a shape generator 30 and / or a sample rate interpolator 32 and / or a value filter 34 and / or a time density scaler 36 as described herein.

[0326] Various embodiments of the present invention may be implemented as a program product for use with a computer system, the program of the program product defining the functions of the embodiments (including the methods described herein). In one embodiment, the program may be contained on various non-transitory computer-readable storage media, where the phrase "non-transitory computer-readable storage medium" as used herein includes all computer-readable media with the sole exception of transitory propagating signals. In another embodiment, the program may be contained on various transitory computer-readable storage media. Exemplary computer-readable storage media include, but are not limited to, (i) non-writable media on which information is permanently stored (e.g., a read-only memory device within a computer, such as a CD-ROM disk readable by a CD-ROM drive, a ROM chip, or any type of solid-state non-volatile semiconductor memory), and (ii) writable storage media on which changeable information is stored (e.g., flash memory, a floppy disk or hard disk drive in a diskette drive, or any type of solid-state random-access semiconductor memory). The computer program may be executed on the processor 102 described herein.

[0327] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. It will be further understood that as used herein, the terms "comprises" and / or "comprising" specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0328] Corresponding structure, materials, acts, and equivalents of all means or step-plus-function elements in the following claims are intended to include any structure, material, or acts for performing the function as specifically claimed in combination with other claimed elements. The description of the embodiments of the present invention is presented for illustrative purposes and is not intended to be exhaustive or limited to the implementation in the form disclosed. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the invention. The embodiments were chosen and described to best explain the principles and some practical applications of the invention and to enable those skilled in the art to understand the invention in various embodiments with various modifications suited to the particular uses contemplated.

[0329] The inventors would like to thank Ia Mgvdliashvili and Dr. Amira Val Baker for their contributions to this disclosure. [Explanation of symbols]

[0330] 2 Virtual Objects 4 center point 6. Observer 8. Sound Source 10 Filter, multi-band filter, first absorption filter 12 First reflection module, first reflection operation 14 Composite Audio Signals 16 Second filter, second multiband filtering operation 18 Modules, Reverberation Modules, Reverberation Calculation 18a Box, Block 20 Resonance module, resonance calculation 21 Mixdown 22 Depth module, depth operation 24 Height module, height calculation 26 Distance module, distance calculation 28 Panning system, panning module, signal distribution matrix, panning matrix, panning matrix calculation 30 Shape Generator 30a Left side (L) 30b Right side (R) 32 "Sample Rate Interpolation" operation, Sample Rate Interpolator 34 "Value Filter" operation, operation, value filter operation, value filter 36 "Time Density Scaler" Operation, Operation, Time Density Scaler Operation, Building Block, Time Density Scaler 40a Reverberant audio signal 40b reverberant audio signal 40c signal 42 Method 42a method 42b Method 42c method 52 Attenuation calculation, calculation 53 Combiner 54 Low-pass filter operation, low-pass filter, operation 55 Attenuation calculation, calculation 56 Time Delay Calculation 57 Attenuation Calculation 58 Signal Feedback Calculation 60 Attenuation calculation, calculation 62 Low-pass filter operation, operation 64 time delay calculation, calculation 66 Second signal feedback calculation 68 Signal inversion calculation, signal attenuation calculation, calculation 70 Low-pass filter operation, operation 72 second time delay operation, operation, time delay operation 73 Signal Feedback Calculation 74 Signal inversion operation 75 Signal Delay Calculation 76 Amplifier 78 Summation and addition operations 79 Attenuation Calculation 80 High-pass filter operation, module 86 Time Delay Calculation 88 Signal Attenuation 90 Signal Feedback Calculation 92 Combining 94 Attenuated or amplified 100 Data Processing System 102 processors 104 memory elements 106 System Bus 108 Local Memory 110 Mass Storage Devices 112 Input Devices 114 Output Devices 116 Network Adapter 118 Applications 140 Signal inversion operation 142 Signal delay calculation, time delay calculation 144 Signal Attenuation 146 Combining 148 Attenuation 160 first time delay operation, time delay operation 162 first signal attenuation calculation, first signal attenuation, calculation 164 Signal Feedback Calculation 166 Combining 168 Second signal attenuation, calculation 170 Second signal delay calculation, second signal attenuation 210 Virtual object dimensions and shape dimensions 212 Virtual point position 214 Container 216 Shape Coordinates 218 Lattice 220 Virtual Point Density

Claims

1. 1. A method for generating a reverberant audio signal associated with a virtual object, comprising: storing a representation of the virtual object, the representation defining a plurality of virtual points that make up the virtual object, the virtual points having respective virtual positions relative to one another, the virtual points belonging to a symmetry group of virtual points, the symmetry group of the virtual points comprising: defining, for each virtual point of said plurality of virtual points, a set of one or more virtual distances comprising a respective virtual distance between the virtual point in question and each other virtual point of said plurality of virtual points; for each set of one or more distances associated with the virtual point, removing distances that are integer multiples of any other distance in the set to obtain a further set of one or more distances associated with the virtual point; for each further set of one or more distances associated with the virtual point, determining distinct distances within the further set in question to form a virtual point-specific set of one or more distances associated with said virtual point; determining virtual points having the same respective virtual point-specific set of one or more distances to form a symmetry group of the virtual points, the symmetry group of the virtual points being associated with the same set of one or more symmetry group distances as the virtual point-specific set of the virtual points; and said one or more sets of symmetry group distances, each set associated with a symmetry group, together form one or more further sets of distances; The method comprises: receiving and / or storing and / or generating an input audio signal; for each virtual point, determining a virtual point audio signal component based on the input audio signal or a filtered version thereof; - combining the virtual point audio signal components thus obtained to obtain a composite audio signal; for each distinct distance in the further set of one or more distances, determining one or more distance audio signals based on the composite audio signal; determining the reverberant audio signal based on the one or more distance audio signals and the virtual point audio signal component; further comprising: method.

2. for each symmetry group, determining one or more symmetry group audio signals based on the determined distance audio signals; determining the reverberant audio signal based on the symmetry group audio signal and the virtual point audio signal component; The method of claim 1 further comprising:

3. determining a virtual point audio signal component for each virtual point based on the input audio signal or a filtered version thereof, for each virtual point, performing a virtual point specific operation on the input audio signal or a modified version thereof, wherein performing said virtual point specific operation comprises performing a time delay operation to introduce a time delay, said introduced time delay being approximately equal to the virtual distance between the virtual point in question and a virtual sound source divided by the speed of sound; 3. The method according to claim 1 or 2.

4. determining one or more distance audio signals for each distinct distance in the further set of one or more distances; for each distinct distance in the further set of one or more distances, determining a first distance audio signal and a second distance audio signal; determining the first distance audio signal for the distinct distances includes modifying the composite audio signal by performing a first time delay operation that introduces a first time delay, a signal attenuation operation, a low pass filter operation, and a signal feedback operation; determining the second distance audio signal for the distinct distances includes modifying the composite audio signal by performing a second time delay operation that introduces a second time delay, a signal inversion operation, a signal attenuation operation, a low pass filter operation, and a second signal feedback operation; 4. The method according to any one of claims 1 to 3.

5. The method of claim 4 , wherein the first time delay introduced by the first time delay operation is equal to the distinct distance divided by the speed of sound.

6. determining, for each symmetry group, one or more symmetry group audio signals based on the distance audio signals, determining, for each symmetry group, a first symmetry group audio signal and a second symmetry group audio signal; determining the first and second symmetry group audio signals comprises selecting distance audio signals from all pairs of first and second distance audio signals, each pair determined for a respective distance in the set of one or more symmetry group distances associated with the symmetry group in question; combining the selected distance audio signals to determine the first symmetry group audio signal; and combining unselected distance audio signals from all pairs of first and second distance audio signals to determine the second symmetry group audio signal.

6. The method of any one of claims 2, 4 and 5.

7. determining the reverberant audio signal based on the symmetry group audio signal and the virtual point audio signal component, combining the symmetry group audio signal and the virtual point audio signal components to determine the reverberant audio signal; combining the symmetry group audio signal and the virtual point audio signal components to determine the reverberant audio signal, determining modified audio signal components, wherein determining the modified audio signal components comprises adding the first or second symmetry group audio signal of the symmetry group in question to each virtual point audio signal component determined for the virtual point belonging to the symmetry group; The method of claim 6.

8. 8. The method of claim 1, further comprising: performing the method of any one of claims 1 to 7 to generate a further reverberant audio signal for a further virtual object; the determined reverberant audio signal associated with the virtual object is used as an input audio signal.

8. The method according to any one of claims 1 to 7.

9. combining the reverberant audio signal associated with the virtual object and the further reverberant audio signal associated with the further virtual object; providing the combination to one or more speakers; 9. The method of claim 8, further comprising:

10. The method of claim 1 , further comprising providing the reverberant audio signal to one or more speakers.

11. providing the reverberant audio signal to the one or more loudspeakers includes providing the reverberant audio signal to a panning system configured to distribute the reverberant audio signal to a plurality of loudspeakers. The method of claim 10.

12. The method further includes filtering the input audio signal before determining a virtual point audio signal component for each virtual point, the filtering of the input audio signal comprising: applying a multi-band filter comprising attenuating respective frequency bands in the input audio signal using respective attenuation coefficients, the respective attenuation coefficients being determined based on a material of the virtual object; 12. The method according to any one of claims 1 to 11.

13. determining one or more distance audio signals for each distinct distance in the further set of one or more distances; determining a distance audio signal for each distinct distance in the further set of one or more distances, the distance audio signal comprising modifying the composite audio signal by performing a time delay operation to introduce a time delay, a signal attenuation operation, a low pass filter operation, and a signal feedback operation, the method comprising: determining a density index for at least one symmetry group of the hypothetical point, determining, for each distance from the set of one or more symmetry group distances associated with the at least one symmetry group, how many feedback operations have been performed per unit of time to determine a distance audio signal for the distance in question, thus obtaining, for each distance from the set of one or more symmetry group distances associated with the at least one symmetry group, a respective number of feedback operations performed; summing the respective numbers of the performed feedback operations to obtain the density index for the symmetry group of the hypothetical point; wherein the method comprises the steps of: receiving a threshold value for the density index; determining that the determined density index is less than the threshold; modifying the stored representation based on the determination by increasing the number of virtual points that make up the virtual object; further comprising:

13. The method according to any one of claims 1 to 12.

14. The low-pass filter operation is determining that the signal to which the low pass filter operation is applied is associated with a Nyquist frequency that is lower than a cutoff frequency associated with the low pass filter operation; based on this determination, upsampling the signal to be filtered so that it is associated with a Nyquist frequency that is greater than or equal to the cutoff frequency; low-pass filtering the upsampled signal; Optionally, determining that the filtered signal is associated with a sample rate that is higher than an output sample rate, the output sample rate being a sample rate that can be output by an output system, and down-sampling the filtered signal based on this determination. Including, 14. The method of claim 4 or 13.

15. a computer-readable storage medium having computer-readable program code embodied therein; A computer comprising: a processor coupled to the computer-readable storage medium, the processor configured to perform the method of any one of claims 1 to 14 in response to executing the computer-readable program code.

16. 15. A non-transitory computer-readable storage medium storing at least one software code portion, the software code portion being configured to perform the method of any one of claims 1 to 14 when executed or processed by a computer.

Citation Information

Patent Citations

  • Sound reproduction device and sound reproduction method

    JP2013073027A

  • Apparatus and method for generating an output signal based on an audio source signal, an acoustic reproduction system, and a loudspeaker signal

    JP2017537574A

  • Generating an audio signal associated with a virtual sound source

    NL2024434A

  • NL2025950

  • Early Reflection Method for Enhanced Externalization

    US20080273708A1