System for transforming an acoustic space by generating audio from environmental analysis
Patent Information
- Application Number
- PCT/US2026/020788
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-25
- Filing Date
- 2026-03-25
- Publication Date
- 2026-10-01
Smart Images

Figure US2026020788_01102026_PF_FP_ABST
Abstract
Description
FELT AU-001SYSTEM FOR TRANSFORMING AN ACOUSTIC SPACE BY GENERATING AUDIO FROM ENVIRONMENTAL ANALYSISI. CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] The present application claims the benefit of U.S. Provisional Patent Application No. 63 / 777,474, filed on March 25, 2025, entitled “System for Synthetic Alteration of Psychoacoustic Space,” the entire disclosure of which is hereby incorporated herein by reference.II. FIELD OF THE INVENTION
[0002] The present invention relates to systems and methods for transforming the perceived acoustics of a space by capturing environmental audio, extracting parameters from the captured audio, generating spectrally novel audio content based on those parameters, and emitting the generated audio into the same acoustic space or into an individualized acoustic space perceived by a listener. More specifically, the invention relates to the use of modal resonance processing, harmonic generation, and configurable parameter mapping to produce audio output containing spectral content not present in the captured environmental audio.III. BACKGROUND OF THE INVENTION
[0003] Active acoustic interventions have taken multiple forms. In professional settings, concert halls and lecture halls install distributed speaker arrays and signal processing equipment to reinforce an original sound source with fidelity and accuracy. These active sound reinforcement systems — manufactured by companies such as Meyer Sound, Yamaha, L- Acoustics, and CARMEN — correct for deficiencies in acoustic dispersion and diffusion to create a more homogeneous, controllable, and tuned acoustic space. These systems are designed to reproduce and spatialize the original sound, not to transform it. Their processing is linear or convolutionbased, meaning the output spectrum is proportional to the input spectrum.
[0004] In a different domain, generative audio systems such as those developed by Endel, Brain. fm, Mubert, and ASTI generate audio content algorithmically for purposes of focus, relaxation, and wellness. These systems create music and other audio from internal algorithms, user preferences, or biometric inputs. They do not capture environmental audio with microphones for the purpose of deriving generative parameters from it, and they do not re-emit generated spectral content into the acoustic space where the environmental audio was captured.FELT AU-001
[0005] Sound masking systems, such as those made by Cambridge Sound Management, Lencore, and Soft dB, emit static or quasi-static noise into a space to reduce the intelligibility of unwanted speech. These systems do not generate novel spectral content derived from the acoustic environment in real-time.
[0006] What is lacking in the prior art is a system that captures environmental audio, extracts parameters from that audio, maps those parameters to a generative processing engine, and produces spectrally novel audio content that is emitted into the same acoustic space or into an individualized acoustic space synchronously perceived by a listener. The present invention addresses this gap.IV. SUMMARY OF THE INVENTION
[0007] The present invention describes a system that transforms the perceived acoustics of a space. The system captures environmental audio, analyzes it to extract metric and event signals, maps those extracted signals to parameters of a generative audio engine via a configurable parameter mapping module, and generates audio output containing spectral content not present in the captured environmental audio. The generated audio is emitted into the acoustic space via audio output transducers. The captured environmental audio, after feedback suppression, serves as the source of extracted parameters for the generative engine and, in certain implementations such as modal resonance processing and frequency- selective convolution, as the excitation signal for the generative processing. In all cases, the output of the generative engine is perceptually new content rather than clear reproduction of the captured audio.
[0008] The system is architecturally distinct from both active acoustic reinforcement systems and generative audio systems. Active acoustic reinforcement systems process audio linearly or proportionally to preserve the spectral character of the source. The present system includes an intermediate parameter extraction layer, of metrics and events derived from captured environmental audio, and uses those parameters to configure generative processing that produces spectral content absent from the input. Alternatively, generative audio systems create content from internal algorithms without direct environmental audio input. The present system requires environmental audio capture as the source of its real-time generative parameters.
[0009] In various embodiments, the system can employ modal resonance processing configured by extracted spectral parameters to produce tonal content from sub-thresholdFELTAU-00Ibroadband energy, harmonic generation from detected environmental fundamentals, feedback-aware parameter mapping that integrates acoustic feedback data with environmental parameters, and time-varying parameter control from composed media sequences or external data feeds. The acoustic space may be a shared space such as a room in which loudspeakers emit audio, or an individualized acoustic space perceived by a wearer of headphones.
[0010] One aspect of the disclosure is directed to a system for transforming an acoustic space, comprising: one or more processors; and memory having stored thereon instructions, wherein the instructions cause the one or more processors to: receive environmental audio captured from an acoustic space; extract one or more metric signals from the environmental audio, the one or more metric signals representing ongoing acoustic characteristics of the acoustic space; map the one or more metric signals to processing parameters according to a configurable mapping; generate audio output comprising spectral content not present in the environmental audio; and output the generated audio output for emission into the acoustic space.
[0011] In some examples, the one or more signals may include one or more of: amplitude envelope, spectral centroid, spectral flux, spectral rolloff, zero-crossing rate, harmonic-to-noise ratio, fundamental frequency estimate, and sub-band energy levels.
[0012] In some examples, the instructions may further cause the one or more processors to extract one or more event signals from the captured environmental audio, the event signals including one or more of: amplitude threshold crossings, frequency detection events, transient onset detection, silence detection, and sound classification results.
[0013] In some examples, the instructions may cause the one or more processors to generate the audio output using one or more of: additive synthesis, subtractive synthesis, sample-based synthesis, physical modeling synthesis, modal resonance processing, and convolution with frequency- selective impulse responses.
[0014] In some examples, the instructions may cause the one or more processors to: map the one or more metric signals to the processing parameters based further on input from at least one of a composed media sequence and an external data feed including one or more of: weather data, time-of-day data, calendar event data, and sensor data from devices external to the system, and wherein the generated audio output varies over time in response to changes in the external dataFELT AU-001feed, and generate the audio output using a combination of the one or more audio input-derived metrics and at least one composed media sequence or external data feed.
[0015] In some examples, the acoustic space may be an individualized acoustic space perceived by a wearer of headphones.
[0016] In some examples, the instructions may further cause the one or more processors to analyze audio signals for resonant feedback between one or more microphones from which the environmental audio is received and one or more audio output transducers at which the generated audio output is emitted; and produce feedback-suppressed audio. The one or more metric signals may be utilized by processors to reduce resonant frequences identified in the audio input.
[0017] In some examples, the instructions may further cause the one or more processors to adjust the configurable mapping repeatedly such that the spectral content of the generated audio output varies in response to changes in the one or more metric signals, causing the perceived acoustic character of the acoustic space to change responsively to its own acoustic content.
[0018] Another aspect of the disclosure is directed to a system for generating audio in an acoustic space, comprising: one or more processors; and memory having stored thereon instructions, wherein the instructions cause the one or more processors to: receive environmental audio from an acoustic space; extract spectral parameters from the captured environmental audio, the spectral parameters comprising at least frequency content and energy distribution information; receive the spectral parameters and derive therefrom modal resonance configuration parameters; receive an audio signal; and produce an audio output in which spectral energy at one or more frequencies of the audio output exceeds the spectral energy of the audio signal at those frequencies by at least a threshold amount measurable via spectral analysis of the respective audio signals, wherein the audio output is produced using a bank of resonant filters whose center frequencies, Q factors, and gains are set by the modal resonance configuration parameters.
[0019] In some examples, the instructions may further cause the one or more processors to produce feedback analysis data characterizing feedback energy detected between one or more microphones from which the environmental audio is received and one or more audio output transducers at which the audio output is emitted, extract feedback-derived metric signals from theFELTAU-00Ifeedback analysis data, and integrate the feedback-derived metric signals with the spectral parameters to derive the modal resonance configuration parameters.
[0020] In some examples, the instructions may further cause the one or more processors to shift center frequencies of resonant filters included in the bank of resonant filters away from frequencies at which feedback energy is detected, and adjust Q factors of resonant filters included in the bank of resonant filters at frequencies at which feedback energy is not detected.
[0021] In some examples, the bank of resonant filters may correspond to a non-physical resonance model whose parameters are not constrained to values corresponding to physically realizable acoustic structures.
[0022] In some examples, the instructions may further cause the one or more processors to update the modal resonance configuration parameters repeatedly in response to changes in the spectral parameters extracted from the captured environmental audio, such that the tonal content of the audio output tracks changes in the environmental audio.
[0023] Yet a further aspect of the disclosure is directed to a system for augmenting an acoustic space with generated harmonic content, comprising: one or more processors: and memory having stored thereon instructions, wherein the instructions cause the one or more processors to: capture environmental audio from an acoustic space; identify one or more fundamental frequencies present in the captured environmental audio: and generate a harmonic audio output at multiples or divisions of each identified fundamental frequency, with harmonic amplitudes determined by a configurable harmonic profile, in which the generated harmonic audio output comprises spectral energy at frequencies that exceed the spectral energy present in the captured environmental audio at those frequencies.
[0024] In some examples, the configurable harmonic profile may specify, for each generated harmonic, an amplitude ratio relative to the corresponding identified fundamental frequency, a maximum number of harmonics to generate, and a decay curve across the harmonic series.
[0025] In some examples, the instructions may further cause the one or more processors to generate audio at non-integer-related frequencies derived from the identified fundamental frequencies.FELTAU-00I
[0026] In some examples, the instructions may further cause the one or more processors to derive a plurality of parameters from the captured environmental audio; process the plurality of parameters: and mix the generated harmonic audio output and the processed plurality of parameters prior to emission by one or more audio output transducers.
[0027] In some examples, the instructions may further cause the one or more processors to identify the one or more fundamental frequencies using one or more of: fast Fourier transform, autocorrelation, and harmonic product spectrum analysis.
[0028] In some examples, any of the example systems described herein may further include: one or more microphones configured to capture environmental audio from an acoustic space; and one or more audio output transducers configured to emit the generated audio output into the acoustic space. In some examples, the system may further include a housing, in which the one or more processors, the memory and at least one of the one or more microphones or the one or more audio output transducers is contained within the housing.
[0029] Yet another aspect of the disclosure is directed to a method performed by one or more processors, in which the one or more processors execute any of the various combinations of instructions stored in the memory as described herein.
[0030] One more aspect of the disclosure is directed to a non-transitory computer-readable medium having programmed thereon any of the various combinations of instructions stored in the memory as described herein.V. APPLICATIONS AND USE CASES
[0031] Here we will outline a few examples which illustrate how this invention can be useful for users in various circumstances, in which one or more embodiments could be employed to alter, augment, or enhance the experience of the user’s acoustic space. While these use cases are not exhaustive, they illustrate how the same system architecture can be applied for many different purposes depending on users’ needs and intents.
[0032] As the world becomes increasingly louder with technology and urbanization, the need for people to create quieter conditions for relaxation, meditation, and sleep is a pressing issue. While this invention does not physically cancel noises from the surrounding environment likeFELTAU-00Inoise-cancelling headphones do, it can use various means to reduce or mask the psychoacoustic impact of these noises. By adding new harmonically related and dynamically responsive sounds to acoustic spaces, the disturbing effects of external sounds can be reduced and softened significantly.
[0033] Multiple systems and techniques can also be combined. For example, even with noisecancelling headphones it is often possible to hear a conversation next to you. It is thus common for users to play music in their headphones to mask such conversations. Alternately, noisecancelling headphones can create a relatively quiet listening space which this invention can then augment with harmonically synthesized audio in perceived real-time. In other words, this invention can further improve the effectiveness of existing noise-cancelling technologies by dynamically masking sounds that still manage to leak through those systems. As an alternative to music, this invention could be used to create a reactive sound curtain that blurs the intelligibility of external speech so that a user can focus on their task at hand.
[0034] In addition, this invention could equally be used to enhance conversation or other meaningful frequencies in a space. In this case, instead of blurring unwanted sounds spectrally or temporally, the system could accentuate certain frequencies from the environment around it. For example, whether using machine learning or pre-configured settings, this system could make speech more intelligible in a loud acoustic environment without the feedback normally associated with proximate audio amplification systems.
[0035] Another use case for this invention involves intentionally enhancing the aesthetic quality of an acoustic space for the purpose of enjoyment, calm, or other mood-alteration. For example, a parent may want to place this system in a child’s bedroom to create an “acoustic blanket” of comforting tones which are reactive to the existing sounds in the room, building, or outside vicinity. Instead of using a static noise machine to bluntly mask those sounds, this invention can enhance environmental sounds to make them more calming for a child or adult who is trying to sleep.
[0036] Another use case for this invention is to help reduce stress in newborn children and infants. Parents of newborns are often challenged by an infant’s trouble with sleep. Noise machines are currently a standard strategy, but they are not responsive to the real-time variations inFELTAU-00Ienvironmental audio. This invention can employ machine intelligence to react to the infant’s level of stress by altering the sound in the room while measuring the intensity of the infant’s cries. Rather than a generic approach, the invention can learn what sound modifications and enhancements (including recordings of parent’s voices) are most effective in calming an individual infant. This method can also be used to quiet a barking dog, a meowing cat, or a squawking pet bird.
[0037] This same technology can be used for other purposes such as creating a productive focus environment, a lively entertainment room that reactively simulates the din of a restaurant, an energizing environment to help exercise or wake up in the morning, and many more examples. Similarly, by enhancing the perceived calmness or comfort of a space, various uses may be found in healing and medicine to improve the acoustic environments of patients (human or animal) and their caregivers.
[0038] This invention can also be integrated with various technology systems such as smart home devices, automotive sound systems, and teleconferencing software. For instance, by integrating with existing smart sensors in a home a user could be sonically alerted to activity in their backyard without a traditional disruptive “alarm” going off. In another example, by integrating with automotive sensors a user could hear augmented audio of the vicinity around their vehicle in new ways, such as a fast-approaching car being heard as a subtle sweeping tone.VI. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] FIG. 1 is a simplified block diagram of the system architecture showing the functional categories, within which certain functional modules would be placed depending on the specific audio processing approach and embodiment.
[0040] FIG. 2 is a more detailed block diagram of the system architecture showing the functional categories, examples of functional modules within these categories, their interconnections, and signal flow paths from environmental audio capture through analysis, parameter mapping, audio generation, and output.FELTAU-00I
[0041] FIG. 3 is a signal flow diagram of energy concentration synthesis techniques, showing the path from environmental audio capture through spectral analysis, parameter mapping, modal resonance or frequency-selective convolution processing, and emission into the acoustic space.
[0042] FIG. 4 is a signal flow diagram of the environmental harmonic generation subsystem, showing frequency analysis, harmonic generation at integer multiples or divisions of detected fundamentals, and emission into the acoustic space.
[0043] FIG. 5 is a signal flow diagram of the feedback-integrated parameter mapping subsystem, showing how feedback analysis data is extracted as feedback-derived metrics and events that are integrated with environmental audio metrics in the parameter mapping module.
[0044] FIG. 6 is a simplified block diagram of the system architecture showing an example hardware arrangement for performing the methods and techniques described in the various embodiments herein.VII. DEFINITIONS
[0045] As used herein, the following terms have the meanings set forth below:
[0046] “Generating spectral content” means producing audio output in which spectral energy at one or more frequencies exceeds the captured audio’s energy at those frequencies by a perceptually significant margin, such that the output contains tonal, harmonic, or resonant content not perceptually detectable in the captured audio. The underlying generative mechanism common to several techniques in the present system is the concentration of distributed sub-threshold energy at specific frequencies into perceptible tonal content. Environmental audio — particularly broadband noise — contains energy at all audible frequencies, but at per-Hz levels below the threshold of perception. Processing that concentrates this distributed energy into narrow frequency bands can raise it above perceptibility, producing tonal content the listener perceives as new. By way of example, environmental broadband noise at approximately -60 dB per Hz can be passed through processing with a sharp resonant peak (Q=100) centered at 220 Hz. The processing concentrates energy in a narrow bandwidth around 220 Hz, producing output approximately 40 dB above the per-Hz input level — an audible tone not perceptually present in the input. Generating spectral content includes, without limitation: (a) modal resonance processing using configurableFELT AU-001resonant filter banks to concentrate energy at parameterized center frequencies; (b) convolution with frequency- selective impulse responses that concentrate energy at specific frequencies rather than scaling the spectrum proportionally: (c) harmonic generation at multiples or divisions of detected fundamentals; (d) additive, subtractive, or sample-based synthesis creating waveforms at specified frequencies; and (e) any other processing producing output frequencies whose energy exceeds the input energy at those frequencies by a perceptually significant margin. The distinction between generative processing and conventional room- simulation processing (such as convolution with a broadband room impulse response) is whether the processing produces perceptually new tonal content or proportionally scales the existing spectrum to reproduce spatial character.
[0047] “Metric signal” means any time-varying signal abstracted from an audio or data input that represents an ongoing measurable property of that input. Examples include, without limitation: amplitude envelope, spectral centroid, spectral flux, spectral rolloff, zero-crossing rate, harmonic -to-noise ratio, fundamental frequency estimate, sub-band energy levels, and any other varying parameter derivable from the input signal by standard signal processing techniques.
[0048] “Event signal” means any binary or momentary signal abstracted from an audio or data input that indicates the occurrence or non-occurrence of a detectable condition. Examples include, without limitation: amplitude threshold crossings, frequency detection exceeding defined thresholds, transient onset detection, silence or pause detection, sound classification results (e.g., speech, music, mechanical noise), and any other binary determination derivable from the input signal.
[0049] “Acoustic space” means the physical or perceptual space within which a listener experiences audio emitted by the system. The acoustic space may be a shared space such as a room, building, or outdoor area in which one or more loudspeakers emit audio perceptible to multiple listeners, or an individualized acoustic space such as that perceived by a wearer of headphones or earbuds, in which the audio output transducer creates a personal listening environment for a single listener. In the headphone embodiment, the system captures environmental audio from the shared space via microphones and emits generated audio into the wearer’s individualized acoustic space, transforming the wearer’s perception of the shared space without altering the shared space itself.FELT AU-001VIII. DETAILED DESCRIPTION
[0050] Referring now to FIGS. 1 and 2, the system consists of connected functional modules for sensing, analyzing, generating, and emitting audio toward the utility of transforming the perceived acoustic character of a space for listening by an audience of one or more people. The sequential order of the modules is important for the signal chain to function appropriately, though the specific components within each module can vary based on unique embodiments.
[0051] The general function of the system is to receive environmental audio, analyze it for relevant metrics and events, generate audio according to abstracted parameters, and emit the generated audio to listeners in their acoustic space. The system architecture includes an intermediate parameter extraction layer between audio capture and audio generation, and a generative processing stage that produces spectral content perceivably absent from the captured input. The combination of these elements is not present in active acoustic reinforcement or generative audio systems identified in the prior art.Audio Inputs 100
[0052] The first module is Audio Inputs 100, whereby environmental sounds are received by the system through acoustic transduction. A Single Microphone 110 captures proximate sound, and different polarities such as omnidirectional, cardioid, and bi-directional can assist in audio source detection as well as feedback management. An Array of Microphones 120 is a configuration of multiple microphones whose positions are calculated in relation to each other to achieve location detection via phase and other relationships, enabling source separation and spatial analysis of the acoustic environment. Streamed Audio Input 130 enables the system to receive real-time audio data from remote sources via wired or wireless connections such as Wi-Fi, Bluetooth, or cellular protocols, supporting embodiments where the environmental audio is captured by a remote microphone. Audio Input Mixer 140 enables the system to blend one or more audio inputs into a combined Audio Input Signal 150, which is sent to other modules as one or more channels in analog or digital forms.FELT AU-001Audio Feedback Suppression 200
[0053] Audio Feedback Suppression 200 receives input from Audio Input Signal 150 and reduces or eliminates resonant feedback resulting from the reception and re-emission of audio in the same acoustic space. In embodiments involving live microphone input and loudspeaker output in the same space, feedback suppression is required for stable system operation.
[0054] Audio Feedback Suppression 200 produces two outputs. Feedback-Suppressed Audio 210 is the audio signal with identified feedback frequencies attenuated or removed. Feedback Analysis Data 220 is a data output characterizing the feedback state of the system, including the frequencies at which feedback energy is detected, the magnitude of that energy, and the rate of change of feedback conditions.
[0055] Feedback-Suppressed Audio 210 can be routed to two destinations: Audio and Data Analysis 400, where it is analyzed for extraction of Metric Signals 410 and Event Signals 420; and Audio Synthesis 600, where it serves as the excitation signal for generative techniques that process audio through parameterized filters (such as Modal Resonance 610 and Frequency- Selective Convolution 620). Feedback Analysis Data 220 is routed solely to Audio and Data Analysis 400. The Audio and Data Analysis 400 module extracts Metric Signals 410 and Event Signals 420 from both the suppressed audio and the feedback data. The extracted signals flow to Parameter Mapping 500, which configures the generative processing engine (Audio Synthesis 600). In implementations where captured audio excites the generative filters (modal resonance, convolution), the output constitutes generated spectral content as defined herein — the listener perceives tonal or resonant content not present in the input, rather than a linear reproduction of the captured audio.Data Inputs 300
[0056] Data Inputs 300 enables the system to receive non-audio data from local or remote sources, whether in the form of real-time streams or stored formats.
[0057] Sensor Data 310 includes data from sensors within the device or connected to it through wired or wireless protocols, including without limitation optical, motion, environmental,FELT AU-001and biometric sensors. These data are sent to Audio and Data Analysis 400 for extraction as Metric Signals 410 and Event Signals 420.
[0058] Composed Media 332 represents saved multimedia formats intended to be translated into system parameters over time, including MIDI files, video files, written notation, graphical timelines, generative data engines, and other scripted sequences or programs. A composed media file can drive the parameter mapping module over time, causing the generative processing to evolve according to a predetermined sequence without requiring continuous environmental audio change as the sole trigger. For example, a composed media file for an art installation could script a 30-minute parameter evolution that gradually shifts modal resonance tuning from warm, low-frequency emphasis to bright, high-frequency emphasis, while environmental audio parameters continue to modulate the moment-to-moment character of the generated output.
[0059] External Data 333 includes data not composed specifically for audio translation, but which the system can transduce into parameter modulations. Examples include weather data (temperature, barometric pressure, wind speed), time-of-day, calendar events, and other environmental or contextual data feeds.Audio and Data Analysis 400
[0060] Audio and Data Analysis 400 receives Feedback-Suppressed Audio 210 and Feedback Analysis Data 220 from Audio Feedback Suppression 200, and non-audio data from Data Inputs 300. This module is the extraction point for all Metric Signals 410 and Event Signals 420 used by the system.
[0061] Metric Signals 410 extracted from the audio include, without limitation: amplitude envelope, spectral centroid, spectral flux, spectral rolloff, zero-crossing rate, harmonic-to-noise ratio, fundamental frequency estimate, sub-band energy levels, and any other continuously varying parameter derivable from the audio signal.
[0062] Event Signals 420 extracted from the audio include, without limitation: amplitude threshold crossings, frequency detection exceeding defined thresholds, transient onset detection, silence or pause detection, and sound classification results.FELT AU-001
[0063] The Audio and Data Analysis 400 module can also extract Metric Signals 410 and Event Signals 420 from Feedback Analysis Data 220. These feedback-derived metrics include the frequencies at which feedback energy is detected, the magnitude of that energy, and the rate of change of feedback conditions. Feedback-derived events include threshold crossings indicating that feedback energy at a given frequency has exceeded a defined level. These feedback-derived metrics and events are made available to the Parameter Mapping 500 alongside all other extracted metrics and events.Parameter Mapping 500
[0064] Once audio and data have been abstracted into Metric Signals 410 and Event Signals 420 by Audio and Data Analysis 400, they become modulations for parameters within Audio Synthesis 600. Parameter Mapping 500 enables the system to dynamically map these signals — including feedback-derived metrics — to functional elements of the generative engine.
[0065] For example, an event signal representing an amplitude threshold crossing could trigger a stored sample at a proportional volume level. In a more sophisticated mapping, the spectral centroid extracted as a metric signal could be mapped to the center frequency of a modal resonator bank, causing the modal resonators to track the spectral center of gravity of the environmental audio in real time. This flexible mapping enables a wide range of generated sounds to be created as a result of environmental conditions, rather than reproducing incoming audio in its original form.Audio Synthesis 600
[0066] Audio Synthesis 600 generates perceptually novel audio (Synthesized Audio 690) using Metric Signals 410 and Event Signals 420 from Audio and Data Analysis 400, mapped to synthesis parameters via Parameter Mapping 500. In certain implementations — including Modal Resonance 610 and Frequency- Selective Convolution 620 — the module also receives Feedback-Suppressed Audio 210 as the excitation signal for its generative filters. In these implementations, the feedback-suppressed audio excites parameterized resonant filters or convolution kernels, and the output is generated spectral content as defined herein: tonal or resonant energy not perceptually present in the input, hr other implementations — including additive synthesis, subtractive synthesis,FELT AU-001physical modeling, and sample-based synthesis — the module receives parameters only and generates audio without an external excitation signal. These and similar synthesis methods are referenced in Other Synthesis Methods 630.
[0067] The purpose of Audio Synthesis 600 is to produce sounds (Synthesized Audio 690) which are reactive to environmental audio characteristics but are spectrally novel — containing tonal, harmonic, or resonant content at levels not perceptually present in the captured audio. The specific generation technique is not limiting; any technique capable of receiving parametric data and generating audio in real time can be used.
[0068] Energy-Concentration Techniques: Referring now to FIG. 3, a category of generative techniques within this module operates by concentrating distributed subthreshold energy at specific frequencies to produce perceptible tonal content, as described in the definition of generating spectral content above. Two implementations are disclosed:
[0069] Modal Resonance 610 uses a bank of resonant filters whose center frequencies, Q factors (bandwidth), and gains are configurable parameters. When configured by spectral parameters extracted from the environmental audio via Audio and Data Analysis 400 and Parameter Mapping 500, the modal resonators concentrate sub-threshold broadband energy at their center frequencies, producing audible tonal content from energy that was imperceptible in the captured audio. The modal resonator parameters are derived from environmental audio analysis, causing the generated spectral content to track the acoustic environment. This is distinguished from conventional modal reverberation used in active acoustics, which uses fixed parameters for room simulation.
[0070] Frequency- Selective Convolution 620 with frequency- selective impulse responses achieves a similar generative result through a different implementation. An impulse response with sharp resonant peaks at specific frequencies and deep attenuation elsewhere — for example, an impulse response that resonates at frequencies in the key of C — concentrates broadband noise energy at those frequencies, producing audible pitched content from sub-threshold input. The impulse response may be selected or configured based on parameters extracted from the environmental audio by Audio and Data Analysis 400 and Parameter Mapping 500, enabling the generated spectral content to track theFELT AU-001acoustic environment. A bank of modal resonant filters and a frequency-selective impulse response are mathematically related, and both produce similar perceptual results: tonal content not present in the input. The distinction from conventional room- simulation convolution is the frequency selectivity of the impulse response: a room IR scales the spectrum broadly and proportionally, reproducing spatial character; a frequency- selective IR concentrates energy at specific frequencies, generating perceptually new content.
[0071] Environmental Harmonic Generation: Referring now to FIG. 4, an additional synthesis technique within this module detects fundamental frequencies present in the captured environmental audio using pitch detection methods via Frequency Analysis 440 (e.g., fast Fourier transform, autocorrelation, harmonic product spectrum, and other techniques) and generates audio at multiples (e.g. 2x, 3x, 4x) or divisions of each detected fundamental via Harmonic Generator 660. The generated harmonic energy exceeds the environmental baseline energy at those frequencies. The Harmonic Generator 660 module accepts a configurable harmonic profile specifying, for each generated harmonic, an amplitude ratio relative to the detected fundamental, a maximum number of harmonics, and a decay curve across the harmonic series. By way of example, an HVAC system might produce a 120 Hz fundamental. Harmonic Generator 660 detects this fundamental and generates content at 240 Hz, 360 Hz, 480 Hz, and higher harmonics with a decaying amplitude profile. The listener might perceive a richer, more comforting quality to the ambient mechanical sound. Alternatively, with a non-integer frequency profile, the system can generate audio at frequencies not harmonically related to the detected fundamental, distributing energy across the spectrum to mask the fundamental rather than enrich it.Audio Outputs 700
[0072] Audio Outputs 700 receives Synthesized Audio 690 (the output of Audio Synthesis 600) and distributes this generated spectral content to one or more destinations. The generated audio may optionally be processed by standard audio effects (such as reverb, spatialization, or filtering) before reaching the output transducers.
[0073] Loudspeakers 710, including single speakers and speaker arrays, emit the generated audio into a shared acoustic space. Headphones 720 (including earbuds) emit the generated audioFELT AU-001into an individualized acoustic space perceived by the wearer. Streamed Audio Output 730 enables the system to transmit generated audio to remote devices via Bluetooth, Wi-Fi, or other network protocols, supporting multi-room embodiments and wireless headphone configurations.Acoustic Space 800
[0074] Acoustic Space 800 represents the acoustic space within which the listener experiences the generated audio. The acoustic space may be a shared space such as a room, building, or outdoor area in which loudspeakers emit audio perceptible to multiple listeners, or an individualized acoustic space perceived by a wearer of headphones or earbuds. In the shared-space embodiment, the system captures environmental audio from the space via microphones, generates audio based on extracted parameters, and re-emits the generated audio into the same space. In the headphone embodiment, the system captures environmental audio from the shared space via microphones and emits generated audio into the wearer’s individualized acoustic space, transforming the wearer’s perception of the shared space without altering the shared space itself.IX. FUNCTIONAL SUBSYSTEMS
[0075] The system architecture described above enables multiple functional subsystems, each extending the core invention of transforming an acoustic space by capturing, analyzing, and generating audio based on extracted parameters.Subsystem 1: Spectrally Generative Transformation
[0076] Referring to FIGS. 1 and 2, this subsystem captures environmental audio (Audio Inputs 100), suppresses feedback (Audio Feedback Suppression 200). extracts Metric Signals 410 and Event Signals 420 (via Audio and Data Analysis 400), maps those signals (via Parameter Mapping 500) to parameters of a generative engine (Audio Synthesizer 600), generates spectrally novel audio (Synthesized Audio 690), and emits the output into Acoustic Space 800 via audio output transducers (Audio Outputs 700).
[0077] By way of example, the system can be placed in a room with ambient conversation and mechanical noise. Audio and Data Analysis 400 extracts the spectral centroid, the amplitude envelope, and the harmonic-to-noise ratio. These metrics are mapped via Parameter Mapping 500FELT AU-001to parameters of a synthesis engine: the spectral centroid controls the filter cutoff of a subtractive synthesizer, the amplitude envelope modulates the output level, and the harmonic-to-noise ratio adjusts the ratio between tonal and noise components in the generated output. The resulting synthesized audio is spectrally related to the environmental sound’s statistical properties but contains none of Audio Input Signal 150’s waveform content. The data analysis, parameter mapping and audio synthesis may be performed within a short enough span of time so that the resulting synthesized audio is perceived as being heard in the acoustic space together with the ambient sound used to produce the synthesized audio. As the ambient sound evolves, the synthesized audio evolves with it. Thus, the synthesized audio is effectively produced in “real time.” The listener perceives a space whose novel acoustic character tracks and responds to the ambient sound, producing a continuously evolving sonic environment.Subsystem 2: Modal Resonance Generation
[0078] Referring to FIG. 3, this subsystem captures environmental audio via Audio Inputs 100, uses Spectral Analysis 430 to extract parameters including frequency content and energy distribution, and maps those parameters to the configuration of a Modal Resonance 610 module. The Modal Resonance 610 module comprises a bank of resonant filters whose center frequencies, Q factors, and gains are set by the extracted spectral parameters.
[0079] When environmental audio — including broadband noise — passes through the modal resonators, high-Q resonators concentrate energy at their center frequencies, producing audible tonal content from sub-threshold input energy. The output spectrum differs from the input spectrum in ways determined by the environmental conditions and mapped parameters.
[0080] By way of example, the system can be placed in a room with HVAC noise and occasional conversation. Audio and Data Analysis 400 detects that the spectral energy is concentrated below 500 Hz (from the HVAC) with intermittent broadband energy above 1 kHz (from speech). Parameter Mapping 500 maps the sub-band energy distribution to the center frequencies of a 12-resonator bank: six resonators are tuned to frequencies in the 200-500 Hz range with Q values of 80-120, and six are tuned to 1-4 kHz with Q values of 40-60. The low-frequency resonators produce warm, organ-like tonal content from the HVAC noise, while the high-frequency resonators produce shimmering tonal content from conversational energy. WhenFELTAU-00Ithe conversation stops, the high-frequency resonators receive only noise-floor energy and their output drops below audibility, leaving only the low-frequency tonal content. The acoustic transformation tracks the room’s activity in real time.
[0081] This subsystem is distinguished from conventional modal reverb in that the modal resonator parameters are continuously derived from environmental audio analysis, not fixed for room simulation. The system does not simply reproduce the acoustic characteristics of a modeled room; instead, it generates perceptually new tonal content whose character is determined by the acoustic content of the actual room at that time.Subsystem 3: Environmental Harmonic Generation
[0082] Referring to FIG. 4, this subsystem captures environmental audio, analyzes the captured audio via Frequency Analysis 440 to identify fundamental frequencies present in the environment, generates audio via Harmonic Generator 660 at multiples or divisions of those fundamentals with configurable harmonic profiles, and emits the generated harmonics into the acoustic space.
[0083] By way of example, a building’s HVAC system produces a persistent 120 Hz hum. The frequency analysis module within Audio and Data Analysis 400 detects this 120 Hz fundamental. The Harmonic Generator 660 module within Audio Synthesis 600 generates content at 240 Hz, 360 Hz, 480 Hz, 600 Hz, and 720 Hz, with amplitude ratios of 0.5, 0.3, 0.2, 0.12, and 0.08 relative to the fundamental, representing a natural harmonic decay profile. The listener now perceives the mechanical hum as a rich harmonic drone.
[0084] In an alternative embodiment, the Harmonic Generator 660 module produces audio at non-integer-related frequencies derived from the detected fundamental, distributing energy across the spectrum to mask the fundamental rather than enrich it.Subsystem 4: Feedback-Integrated Parameter Mapping
[0085] Referring to FIG. 5, in this subsystem, Feedback Analysis Data 220 from Audio Feedback Suppression 200 is routed to Audio and Data Analysis 400, where it is extracted into feedback-derived Metric Signals 410 and Event Signals 420. These feedback-derived signals areFELT AU-001then made available to Parameter Mapping 500 alongside the environmental audio metrics and events. Parameter Mapping 500 uses the combined set of environmental and feedback-derived metrics when configuring the generative processing engine.
[0086] By way of example, Audio and Data Analysis 400 detects increasing feedback energy at 800 Hz from Feedback Analysis Data 220 and outputs a feedback-frequency metric at 800 Hz and a feedback-magnitude metric showing a rising trend. Parameter Mapping 500 receives these feedback-derived metrics alongside the environmental spectral parameters and shifts the center frequencies of the modal resonator bank away from 800 Hz while simultaneously increasing the Q factors of resonators at nearby frequencies where feedback energy is not detected (e.g., 690 Hz and 950 Hz). The modal engine’s output shifts its tonal character to avoid the feedback frequency, and the listener perceives a smooth evolution of the generated sound rather than an abrupt notch or dropout. In this embodiment, the feedback condition can optionally be managed within Audio Synthesis 600 rather than or in addition to Audio Feedback Suppression 2OO.Referring to FIG. 6, the various techniques and routines described herein may be implemented using the example System 900. The System 900 may include one or more Processing Devices 910 configured to execute a set of instructions or executable programs. The processors may be dedicated components such as general-purpose CPUs, or application specific integrated circuit ("ASIC"), or may be other hardware-based processors. Although not necessary, specialized hardware components may be included to perform specific computing processes faster or more efficiently. For example, operations of the present disclosure may be carried out in parallel on a computer architecture having multiple cores with parallel processing capabilities.
[0087] The System 900 may further include one or more storage devices or Memory 920 for storing Instructions 930 and programs executed by the one or more Processor(s) 910. The Instructions 930 may include any of the various software modules described herein, including but not limited to feedback analysis, audio spectral analysis, audio frequency analysis, audio parameter mapping, audio synthesis, and so on. Additionally, the Memory 920 may be configured to store Data 940, such as sensor data from the Data Inputs 300, analysis results on audio analysis functions such as the feedback analysis of Audio Feedback Suppression 200 or the determined metric and event signals of the Audio and Data Analysis 400 modules.FELT AU-001
[0088] In some examples, the Instructions 930 may be stored on a computer-readable medium such as a storage disc or memory drive, and may be accessed and executed by the one or more Processor(s) 910. In this manner, the Instructions 930 may be thought of as a method or routine containing a series of steps that, when executed, achieve the various objectives of the present disclosure.
[0089] The System 900 may further include one or more Interfaces 950 for input and output of data. For example, Audio Inputs 100 may be received by the Processor(s) 910 via an Input Interface 952. For further example. Audio Outputs 700 may be output by the Processor(s) 910 via an Output Interface 954. In some cases, analog-to-digital conversion circuitry may be included in the Interfaces 950 to receive analog audio inputs and transmit analog audio outputs. It should be appreciated that the Interfaces 950 may further be capable of receiving digital audio inputs, such as a Streamed Audio Input 115, as well as outputting digital audio output, such as Streamed Audio Output 730. Additionally, aside from audio inputs and audio outputs, other parameters and instructions may be provided to and from the system via Interfaces 950.
[0090] In the example of FIG. 6, the System 900 is shown as being contained in a Housing 901. The housing may include each of the Processor(s) 910. Memory 920. and Interfaces 950. In some examples, either one or both of Audio Inputs 100 (such as microphones in the Acoustic Space 800) and Audio Outputs 700 (such as speakers or transducers in the Acoustic Space 800) may also be housed within the same Housing 901. Containing all processing circuitry along with some or all of the interface circuitry in a single housing may facilitate easy setup of the System 900. Additionally, containing the interface circuitry and processing circuitry in the common housing is a way of providing direct, and sometimes dedicated, communication connections between the inputs, processors and outputs. These connections can help to reduce latency costs, which is important for real-time audio applications in which complex sounds must be analyzed and modified in a short window of time to avoid perception of a lag or delay in the final audio result. Reducing latency in the audio processing may in turn allow for higher processing bandwidth, which in turn can result in more complex or robust sounds.
[0091] Alternatively, in some examples, the Housing 901 may contain only the Processor(s) 910, Memory 920, and Interfaces 950, and latency between input and output components (e.g., Audio Inputs 100, Audio Outputs 700) and the Processor(s) 910 can be addressed by providingFELTAU-00Istrong connections therebetween. For example, instead of the System 900 being a cloud-based device accessing remote processing and storage components, the System 900 may be a physical device that is connected, either physically via a wire or remotely via a wireless connection such as Bluetooth or Wi-Fi, to the Audio Inputs 100 and Audio Outputs 700 using a low latency connection. The choice of connection may depend on the degree of tolerable latency in a given system, which in turn may depend on the bandwidth of audio signals, the complexity of processing operations, or a combination thereof.
[0092] Although the invention herein has been described with reference to particular embodiments, it is to be understood that these embodiments are merely illustrative of the principles and applications of the present invention. It is therefore to be understood that numerous modifications may be made to the illustrative embodiments and that other arrangements may be devised without departing from the spirit and scope of the present invention as defined by the appended claims.
Claims
FELT AU-001IN THE CLAIMS1. A system for transforming an acoustic space, comprising:one or more processors; andmemory having stored thereon instructions, wherein the instructions cause the one or more processors to: receive environmental audio captured from an acoustic space; extract one or more metric signals from the environmental audio, the one or more metric signals representing ongoing acoustic characteristics of the acoustic space; map the one or more metric signals to processing parameters according to a configurable mapping;generate audio output comprising spectral content not present in the environmental audio; andoutput the generated audio output for emission into the acoustic space.
2. The system of claim 1, wherein the one or more signals comprise one or more of: amplitude envelope, spectral centroid, spectral flux, spectral rolloff, zero-crossing rate, harmonic -to-noise ratio, fundamental frequency estimate, and sub-band energy levels.
3. The system of claim 1, wherein the instructions further cause the one or more processors to extract one or more event signals from the captured environmental audio, the event signals comprising one or more of: amplitude threshold crossings, frequency detection events, transient onset detection, silence detection, and sound classification results.
4. The system of claim 1, wherein the instructions cause the one or more processors to generate the audio output using one or more of: additive synthesis, subtractive synthesis, samplebased synthesis, physical modeling synthesis, modal resonance processing, and convolution with frequency-selective impulse responses.
5. The system of claim 1, wherein the instructions cause the one or more processors to:map the one or more metric signals to the processing parameters based further on input from at least one of a composed media sequence and an external data feed comprising one or more of: weather data, time-of-day data, calendar event data, and sensor data from devices external toFELT AU-001the system, and wherein the generated audio output varies over time in response to changes in the external data feed, andgenerate the audio output using a combination of the one or more audio input-derived metrics and at least one composed media sequence or external data feed.
6. The system of claim 1, wherein the acoustic space comprises an individualized acoustic space perceived by a wearer of headphones.
7. The system of claim 1, wherein the instructions further cause the one or more processors to analyze audio signals for resonant feedback between one or more microphones from which the environmental audio is received and one or more audio output transducers at which the generated audio output is emitted; andproduce feedback-suppressed audio, wherein the one or more metric signals are utilized by processors to reduce resonant frequences identified in the audio input.
8. The system of claim 1, wherein the instructions further cause the one or more processors to adjust the configurable mapping repeatedly such that the spectral content of the generated audio output varies in response to changes in the one or more metric signals, causing the perceived acoustic character of the acoustic space to change responsively to its own acoustic content.
9. The system of claim 1, further comprising:one or more microphones configured to capture environmental audio from an acoustic space: andone or more audio output transducers configured to emit the generated audio output into the acoustic space.
10. The system of claim 9, further comprising a housing, wherein the one or more processors, the memory and at least one of the one or more microphones or the one or more audio output transducers is contained within the housing.
11. A system for generating audio in an acoustic space, comprising:one or more processors; andFELT AU-001memory having stored thereon instructions, wherein the instructions cause the one or more processors to:receive environmental audio from an acoustic space;extract spectral parameters from the captured environmental audio, the spectral parameters comprising at least frequency content and energy distribution information; receive the spectral parameters and derive therefrom modal resonance configuration parameters;receive an audio signal; andproduce an audio output in which spectral energy at one or more frequencies of the audio output exceeds the spectral energy of the audio signal at those frequencies by at least a threshold amount measurable via spectral analysis of the respective audio signals, wherein the audio output is produced using a bank of resonant filters whose center frequencies, Q factors, and gains are set by the modal resonance configuration parameters.
12. The system of claim 11, wherein the instructions further cause the one or more processors toproduce feedback analysis data characterizing feedback energy detected between one or more microphones from which the environmental audio is received and one or more audio output transducers at which the audio output is emitted,extract feedback-derived metric signals from the feedback analysis data, andintegrate the feedback-derived metric signals with the spectral parameters to derive the modal resonance configuration parameters.
13. The system of claim 12, wherein the instructions further cause the one or more processors to shift center frequencies of resonant filters included in the bank of resonant filters away from frequencies at which feedback energy is detected, andadjust Q factors of resonant filters included in the bank of resonant filters at frequencies at which feedback energy is not detected.FELT AU-00114. The system of claim 11, wherein the bank of resonant filters corresponds to a nonphysical resonance model whose parameters are not constrained to values corresponding to physically realizable acoustic structures.
15. The system of claim 11, wherein the instructions further cause the one or more processors to update the modal resonance configuration parameters repeatedly in response to changes in the spectral parameters extracted from the captured environmental audio, such that the tonal content of the audio output tracks changes in the environmental audio.
16. A system for augmenting an acoustic space with generated harmonic content, comprising:one or more processors; andmemory having stored thereon instructions, wherein the instructions cause the one or more processors to:capture environmental audio from an acoustic space;identify one or more fundamental frequencies present in the captured environmental audio; andgenerate a harmonic audio output at multiples or divisions of each identified fundamental frequency, with harmonic amplitudes determined by a configurable harmonic profile,wherein the generated harmonic audio output comprises spectral energy at frequencies that exceed the spectral energy present in the captured environmental audio at those frequencies.
17. The system of claim 16, wherein the configurable harmonic profile specifies, for each generated harmonic, an amplitude ratio relative to the corresponding identified fundamental frequency, a maximum number of harmonics to generate, and a decay curve across the harmonic series.
18. The system of claim 16, wherein the instructions further cause the one or more processors to generate audio at non-integer-related frequencies derived from the identified fundamental frequenciesFELT AU-00119. The system of claim 16, wherein the instructions further cause the one or more processors to:derive a plurality of parameters from the captured environmental audio;process the plurality of parameters; andmix the generated harmonic audio output and the processed plurality of parameters prior to emission by one or more audio output transducers.
20. The system of claim 16, wherein the instructions cause the one or more processors to identify the one or more fundamental frequencies using one or more of: fast Fourier transform, autocorrelation, and harmonic product spectrum analysis.