Method and system for generating sound field capable of perceiving rhythm characteristics
By directly stimulating the physical resonator, a sound field with natural acoustic characteristics and controllable physiological rhythms is generated, which solves the problem of lack of physical authenticity in electronic synthetic acoustic technology and achieves human perception matching and comfort improvement in high-end acoustic environments.
Patent Information
- Application Number
- CN202510948769.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-09-12
AI Technical Summary
Existing electronic synthetic acoustic technology lacks physical authenticity, resulting in a mismatch between the acoustic environment and the human perception system, making it difficult to maintain natural adaptability and auditory comfort in high-end acoustic environments.
By directly and controllably exciting the physical resonator, a sound field with natural acoustic characteristics and controllable physiological rhythms is generated. The excitation mechanism is used to drive the physical resonator network to form a perceptible rhythmic characteristic sound field covering the frequency range of 1-100 Hz.
It creates a highly realistic, natural and comfortable acoustic experience that conforms to the laws of human ear perception, meets the acoustic comfort zone requirements of the ISO 3382-1 standard, and enhances the immersion and adaptability of the acoustic environment.
Smart Images

Figure CN120636362A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the interdisciplinary technical field of new-generation information technology and high-end equipment manufacturing, and in particular to a method and system for real-time generation and control of physical acoustics based on artificial intelligence. Background Art
[0002] In the field of acoustic experience and acoustic assisted adjustment, creating an acoustic environment that sounds natural and conforms to human perception is crucial. Existing technology mainly relies on electronic speakers to synthesize sound waves, but this technology has inherent limitations due to the lack of physical acoustic authenticity.
[0003] Electronic speakers are limited by the electro-acoustic conversion mechanism. The sound waves they generate are essentially different from the sounds produced by the vibration of real physical entities in terms of the richness of the timbre spectrum, the natural evolution of harmonics, and the characteristics of sound energy attenuation. For example, when a physical resonator (such as a crystal vessel or a wooden component) is excited, its sound energy attenuation presents a nonlinear process determined by the internal resistance and geometric topology of the material, accompanied by the dynamic evolution of the harmonic components. Electronically synthesized sound waves usually use simplified mathematical models (such as exponential decay models), resulting in a step-like truncation characteristic of sound energy attenuation.
[0004] Studies have shown (see Zwicker et al., 1999; Patterson et al., 2010) that the nonlinear attenuation sound waves produced by physical resonators are more consistent with the time-domain integration characteristics of the human auditory system; whereas, the mechanical truncation characteristics of electronically synthesized sound may increase the adaptive load of the auditory system (Moore, 2012); at the level of rhythm perception, physical acoustic rhythms of 1-100 Hz can produce an effective synchronization effect with the body's intrinsic perceptual rhythm (Large et al., 2009), optimizing the synchronization between acoustic rhythms and the perceptual system; such unnatural acoustic characteristics may interfere with the adaptive response of the auditory system, resulting in a mismatch between the sound field environment and perceptual expectations.
[0005] This is especially true in the construction of high-end acoustic environments (such as: collaborative innovation spaces for creative corporate teams (for example, creating an acoustic environment that promotes "collective flow" among team members); high-end experiential inspiration spaces (for example, acoustic devices in museums and art galleries used to stimulate deep thinking and emotional resonance); original acoustic design of luxury spaces (such as the reconstruction of the soundscape of hotel lobbies / villa courtyards); optimization of cognitive states in cutting-edge educational spaces (for example, existing technologies lack acoustic environment solutions that can effectively guide learners to enter and maintain efficient learning states such as "flow"); optimization of the intrinsic rhythm of nighttime sound field environments (in compliance with ANSI / ASAS12.2-2022 architectural acoustics specifications)). Electronically synthesized rhythms lack the authenticity of the physical vibration source, making it difficult to maintain the time-domain phase-locked state between the acoustic rhythm and the perception system. Similarly, in education and collaborative innovation scenarios, existing technologies also lack an acoustic environment solution that can actively guide individuals or groups into specific cognitive states without incurring additional auditory burden.
[0006] Therefore, there is an urgent need for a technology based on the intrinsic characteristics of physical vibrations that can generate a perceptible rhythmic sound field, so as to improve the matching degree between acoustic parameters and perception systems and meet the natural adaptability requirements of high-end acoustic environments. Legal and Security Boundary Statement
[0007] To clarify the application scope and safety attributes of this invention, we hereby declare:
[0008] 1. Non-medical device classification: This invention is strictly limited to the scope of optimizing the perceptual experience of healthy people. It does not have the function of diagnosing, treating or intervening in any mental disorder, and does not meet any functional requirements of "medical device" in the "Medical Device Supervision and Administration Regulations"; any physiological effects produced by the system are the natural response of the human body to physical acoustics and do not constitute medical-grade biofeedback.
[0009] 2. Technical Achievement Limits: All core acoustic properties (such as timbre, attenuation, and rhythmic effects) of this invention's unique, rhythmically perceptible acoustic layer are derived from the physical vibrations of the material itself, rather than from the synthesis or modulation of digital signals via electronic speakers. This invention does not preclude the ability to combine this physically generated acoustic layer with other audio (such as original music or vocals) played through conventional sound reinforcement equipment to create a composite sound field. Summary of the Invention
[0010] The present invention provides a sound field generation method that is completely different from existing electronic synthetic acoustics in terms of technical principles. It creates a real sound field with natural acoustic characteristics and controllable physiological rhythms through direct and controllable excitation of physical resonators, fundamentally solving the technical problems inherent in existing electronic synthetic acoustics technology, such as the lack of physical reality and unnatural auditory experience. To this end, it provides a new adaptive sound field generation method and system based on physical resonators.
[0011] To achieve the above-mentioned objectives, the technical solution adopted by the present invention includes: determining a set of acoustic guidance parameters including at least one of rhythm, fundamental frequency or harmonic components; driving one or more excitation mechanisms according to the set of parameters to stimulate the physical resonator network to produce an acoustic environment whose core acoustic characteristics (such as timbre spectrum and sound energy attenuation characteristics) are dominated by the physical properties of the resonator itself; and through the acoustic beat effect or isochronous sound effect endogenous to the acoustic environment, forming a physical sound field with perceptible rhythmic characteristics covering the frequency range of 1-100 Hz in the environment.
[0012] Compared to existing technologies, this invention offers the following advantages: By returning to the fundamental principles of physical vibration, it creates a highly realistic, natural, and comfortable acoustic experience. Its sound energy attenuation, driven by physical resonators, better aligns with the human ear's perception. For example, its attenuation curve meets the acoustic comfort zone requirements defined by the ISO 3382-1 standard, creating a more immersive acoustic environment for users. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 It is a functional block diagram showing a preferred overall architecture of the system of the present invention.
[0014] Figure 2 It shows Figure 1 Data flow diagram of an optimal workflow within the parameter determination module.
[0015] Figure 3 It is a schematic diagram showing the core structure of the physical platform described in [Embodiment 1] of the present invention.
[0016] Figure 4 It is a principle diagram for explaining the dynamic tuning principle and controllable beat generation method in [Embodiment 1] of the present invention. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical solutions and advantages of the present invention clearer, the following Figures 1 to 4 The technical solution of the present invention is described in detail.
[0018] First see Figure 1, which shows a preferred overall architecture of the system of the present invention. The system may include a central control engine (100) as a core processing unit. The engine is used to receive data from one or more input modules and generate final control instructions. These input modules may include an audio content source (110) or a user target setting module (130). The central control engine (100) may include a parameter determination module (101) and an instruction generation module (102) for generating an acoustic generation script. The script is sent to the physical excitation module (140).
[0019] Please continue reading Figure 1 The physical excitation module (140) may include an excitation mechanism array (141) and a physical resonator network (142). The excitation mechanism (141) excites one or more physical resonators in the physical resonator network (142) according to the acoustic generation script, thereby generating a "physical acoustic environment" with specific acoustic characteristics in the physical space.
[0020] See also Figure 2 , which shows in detail Figure 1 A preferred workflow within the parameter determination module (101). In one embodiment, the module starts from an audio file (200) and can extract multiple dimensions of features such as beat / BPM (221), harmony (222) and melody contour (223) in parallel through the underlying acoustic feature extraction (210) step. These underlying features are then fed into the form / emotional structure analysis (230) step to form the final structured music feature data (240) containing information such as timestamps. This set of data will serve as the basis for generating acoustic guidance parameters. Those skilled in the art should understand that the above Figure 2 The workflow shown is only an exemplary description of achieving the purpose of the present invention. Any other algorithm or method that can extract rhythm, frequency or harmonic features that can be used to generate acoustic guidance parameters from audio content should fall within the protection concept of the present invention.
[0021] In the present invention, the term "physical resonator" (142) can be understood as a physical object that can be excited by external energy and produce characteristic vibrations and sounds according to its own material, shape and structure. To illustrate the technical solution described in claim 3, the physical resonator (142) includes but is not limited to: a solid resonator (preferably a precisely tuned acoustic resonant stone bar made of natural stone, which has long resonant sound (T60>2.5s) and low damping acoustic characteristics through the optimization of geometric structure and material internal resistance), a tensioned string / diaphragm, and a gas / liquid cavity within a limited volume. In a preferred embodiment (as described in Example 1), the physical resonator (142) can be configured as a replaceable module, which allows the user to change the timbre or tonality of the system as needed.
[0022] In the present invention, the term "stimulation mechanism" (141) can be understood as a device that can convert digital instructions in the acoustic generation script into precise physical actions to stimulate the resonant body. To illustrate the technical solution described in claim 2, the specific implementation of the stimulation mechanism can be diverse, for example, including: a liquid dripping device controlled by an electromagnetic valve (a fluid stimulation mechanism); a micro-hammer driven by a solenoid or a motor with a striking head made of different materials (such as a felt head, a rubber head) (a mechanical contact stimulation mechanism); or an electromagnetic stimulation mechanism that drives the vibration of the resonant body in a non-contact manner by changing the attraction and repulsion between the electromagnetic field and the permanent magnet, etc.
[0023] To illustrate the technical solution of claim 4, in a preferred embodiment, the central control engine (100) can also output a digital signal of an external audio content (such as music or voice) synchronously with the generated "acoustic generation script". The digital audio signal can be played through an external playback device such as a conventional speaker, while the acoustic environment generated by the physical resonance body serves as the basic physical sound field in the space, and the two are presented together to form a composite sound field that combines virtuality and reality.
[0024] To illustrate the technical solution described in claim 5, in another preferred embodiment, to generate an acoustic beat effect, the "acoustic generation script" may include at least two sets of parallel excitation instructions. For example, instruction A may drive the first excitation mechanism to continuously excite resonant body A at a fundamental frequency of 400 Hz, while instruction B may drive the second excitation mechanism to continuously excite resonant body B at a fundamental frequency of 410 Hz. This will endogenously generate a stable acoustic beat effect in the environment, with an intensity that periodically varies at 10 Hz.
[0025] To illustrate the technical solution described in claim 6, in another preferred embodiment, to create an isochronous sound effect, the "acoustic generation script" defines a series of discrete stimulus events. For example, the script may instruct a single or multiple stimulus mechanisms to perform a short tapping at a fixed interval (e.g., every 100 milliseconds), thereby generating a 10Hz, rhythmic, and physically rhythmic pulse in the environment.
[0026] To illustrate the technical solution described in claim 7, in a preferred embodiment, the acoustic guidance parameters are configured to dynamically and gradually change the rhythmic frequency. For example, the system can dynamically adjust the frequency difference of the acoustic beat effect or the excitation interval of the isochronous sound effect according to a preset frequency change curve (e.g., a linear decrease from 15 Hz to 10 Hz over 1 minute). This "gradual" change method can gently guide the user's neural state to smoothly transition from one cognitive state (such as the focused Beta band) to another cognitive state (such as the relaxed Alpha band).
[0027] To illustrate the technical solution of claim 9, in another preferred embodiment, the system may further include a dynamic tuning module. The module is used to dynamically adjust the resonant frequency of the physical resonator (142) by changing the effective physical parameters of the physical resonator (142) in real time. For example, for a hollow or liquid-filled resonator (such as a crystal cup), the module can inject or extract liquid medium into or out of the resonator through a precision pump to change its effective mass, thereby achieving a smooth and dynamic change in its resonant frequency.
[0028] Example 1: Physical Sound Field Generation Platform and Its Core Capability Verification
[0029] This example describes a method for generating an acoustic environment with specific perceptible rhythmic characteristics using pre-defined scripting rules. It not only demonstrates the basic feasibility of this invention but also details its core hardware components and multi-dimensional capabilities as a dynamically tunable physical acoustic platform.
[0030] 1. Hardware composition
[0031] See also Figure 3In this embodiment, the physical resonator network (142) can be specifically implemented as an array of 61 solid resonators covering five full octaves, wherein each solid resonator is preferably an array of precisely tuned acoustic resonant stone bars (310). The stone bars are carefully selected and tuned to ensure that they have excellent acoustic properties of long resonant sound (for example, a T60 reverberation time greater than 2.5 seconds in the 200-800 Hz frequency band) and low damping, and can produce harmonically rich and pure physical sounds, with a range corresponding to C3 (130.81 Hz) to C8 (4186.01 Hz). The excitation mechanism array (141) can be specifically 61 independent electromagnetic valve drip heads (320) corresponding to the stone bars, and a constant pressure water supply system (330) ensures that the energy and initial conditions of each water drop excitation are highly consistent, thereby ensuring the stability of the acoustic output. It is understandable that the actual sound range characteristics may have individual differences due to the physical properties of the material, and can be calibrated through conventional acoustic measurement methods during specific implementation.
[0032] 2. Core Capability Verification 1. Generation of a "Multi-Frequency Isochronous Tone" Effect (Further Description of Claim 6) Implementation: By editing a script containing rhythmic execution rules, isochronic sound effects can be generated in a highly musical manner. The core technique is to apply the target rhythm (e.g., 10 Hz) as a dynamic "performance method" to a changing sequence of notes. For example, for a musical phrase consisting of C4 (1 second) - G4 (2 seconds) - E4 (1 second), the system will rewrite it into a series of high-frequency pulse excitation instructions: 10 excitations of the C4 resonator in the first second, 20 excitations of the G4 resonator in the next 2 seconds, and finally 10 excitations of the E4 resonator in the next 1 second.
[0033] Technical Effect: The resulting auditory effect is a melody flowing between C4, G4, and E4, with each note pulsating in intensity at a frequency of 10 Hz. This is the unique multi-frequency isochronic tone created by this invention. Its frequency structure (varying pitch) carries the musical melody, while its temporal structure (constant pulse rhythm) provides a clear rhythmic basis, thus avoiding the monotony and auditory fatigue of traditional single-frequency electronic isochronic tone.
[0034] 2. Dynamic Physical Tuning Principle and Controllable Beat Generation (Further explanation of the solutions described in claims 5 and 9) A core innovation of the present invention is its ability to dynamically tune a physical resonator, thereby generating a physical acoustic beat effect with precisely controllable frequency.
[0035] See also Figure 4, which shows the core principle and structure for achieving dynamic tuning and beat generation. The core principle of the dynamic tuning of the present invention is to actively and reversibly change the effective physical parameters of the variable resonance body (420) through a controllable liquid (421), thereby changing its resonant frequency; wherein the variable resonance body (420) is composed of a glass cup firmly fixed on an acoustic resonance stone bar, and the liquid level control subsystem (430) accurately controls the liquid (421) injected into the glass cup.
[0036] Feasibility verification: The principle has been verified by a prototype: after actual measurement, for a composite resonator composed of a glass cup and a stone bar with a length of 85 cm and a fundamental frequency of D4 (293.66 Hz), after 80 ml of water was injected into the glass cup, the resonant frequency of the system decreased by about 4.3 Hz.
[0037] Beat generation application example: Based on this capability, to generate an 8 Hz beat, the system can perform the following operations: select two physically independent resonators, one of which is used as a reference resonator (410) (for example, a stone bar with a fundamental frequency of ~293 Hz), and the other as a variable resonator (420). The "variable resonator" is operated through the liquid level control subsystem (430) to reduce the frequency of the variable resonator by about 8 Hz to ~285 Hz. At the same time, the "reference resonator" and the tuned "variable resonator" are excited respectively. Ultimately, the sound waves generated by the two sound sources in space physically interfere with each other, and theoretically, the interference of the two resonators should produce a beat of ~11 Hz.
[0038] 3. Unexpected acoustic aesthetic effect: During the testing of the above-mentioned expansion scheme, an unexpected and highly valuable technical effect was discovered: when the excitation mechanism (drip head) drips water into the glass of the variable resonance body (420) and collides with the liquid (421) therein, in addition to exciting the resonance body to vibrate and produce pure musical sound, it also produces a "bubble sound" with rich high-frequency details when the water drop collides with the water surface and the bubbles burst. These natural and random acoustic details are harmoniously interwoven with the stable musical sound, creating a highly immersive listening experience that is rich in musicality and feels like being under a quiet roof listening to the natural rain scene.
[0039] III. Summary of Example 1: This example describes a powerful and flexible physical acoustics platform in detail, demonstrating its core capabilities across multiple dimensions, including rhythm generation, frequency tuning, beat generation, and acoustic aesthetics. It provides a solid hardware foundation and broad creative possibilities for the more advanced, AI-based analysis-based Example 2.
[0040] Example 2: Intelligent physical soundscape rendering system based on audio analysis.
[0041] To further illustrate the technical solution described in claim 4 (hybrid mode), this embodiment describes an advanced form of the present invention, which aims to coordinate the playback of an input audio signal (especially a signal containing human voice or instrument solo) with a physical sound field that is tailored for it, generated with a lag, and contains multi-part harmony and rhythmic information, thereby creating a new acoustic experience that combines the real and the virtual and is highly integrated.
[0042] 1. Core Concept The core of this embodiment lies in defining the physical sound field system of the present invention as an "intelligent music post-production and physical rendering engine." This engine performs multi-dimensional feature analysis on the input audio signal and, based on the analysis results, renders a physical acoustic layer that evolves dynamically and microscopically in sync with the original audio, ultimately achieving a coordinated presentation of the original audio signal and the generated physical acoustic events.
[0043] 2. System Composition and Collaboration To achieve this functionality, the entire application can be considered as two independent yet collaborative systems: 1. The physical sound field generation system described in this invention: Its core task is to receive commands and generate the physical sound field. It internally includes an audio analysis and command generation engine and a physical execution system. 2. An external high-fidelity sound reinforcement system: Its function is to play the original audio and has an adjustable delay function for final sound synchronization.
[0044] 3. Workflow: From Analysis to Multidimensional Rendering Step 1: The multi-dimensional feature analysis system first receives a complete input audio (for example, a recording of a human voice singing or a violin solo). Its internal audio analysis and instruction generation engine deeply processes the audio and comprehensively extracts features such as its pitch, beat, speed, and the changes in the loudness envelope of the overtone structure of the key musical notes during their duration. This step provides a data basis for determining the parameters of the target rhythm sequence defined in claim 6 (isochronous sound effect) and claim 7 (gradual effect). Step 2: "Acoustic texture reconstruction" script generation This is the core algorithm of this embodiment. The engine integrates all the analyzed feature information to generate a unified physical execution script that contains "acoustic texture reconstruction" information.
[0045] Take a melody segment "do-mi-fa" (corresponding to the pitch of C4-E4-F4) sung by a human voice as an example: a. Multi-part harmony mapping: When the analysis engine recognizes that the melody note is C4, its harmony generation module will decide to generate a C major chord consisting of three notes (C4, E4, G4) as a physical harmony based on the rules of music theory, and allocate three independent physical resonance body channels to these three pitches. The material of the resonance body can preferably be the crystal or stone defined in claim 3, and its excitation method can adopt the electromagnetic excitation mechanism defined in claim 2. b. Rhythmic playing method application: The engine sets a basic 8Hz rhythm and decides to apply this rhythm as a "playing method" to the entire C major chord to be generated. c. Velocity envelope synchronization: The engine maps the overtone intensity envelope of the original human voice C4 note analyzed into a series of pulse-excited velocity control envelopes. d. Final script generation: The engine encodes all of the above decisions into a complex, parallel sequence of instructions, instructing the physical system to synchronize pulse excitation at a frequency of 8Hz after tuning to C4, E4, and G4, and the intensity of all pulses must follow the preset velocity control envelope.
[0046] Step 3: Parallel Execution and External Synchronization: Physical Path: Upon receiving the script, the physical execution system immediately begins work, systematically performing dynamic tuning and "acoustic texture reconstruction"-style excitation. Electronic Path and Synchronization: Simultaneously, the original input audio signal is sent to an external high-fidelity sound reinforcement system. The system's delay function is set to a value that matches the execution delay of the physical path to ensure precise alignment of the two audio playback channels. This delay value can be set manually by an operator based on calibration data, or automatically by the system outputting a simple synchronization parameter. The specific implementation method does not constitute the core of this invention.
[0047] 4. Final Technical Effect and Creative Summary Final technical effect: After precise synchronization and coordination, the system outputs a composite sound field composed of the original foreground audio and the background physical acoustic layer. The background physical acoustic layer has the following measurable and objective technical characteristics: 1. Polyphonic harmonic structure: At any moment, it contains multiple fundamental frequency components that have a clear mathematical frequency relationship and sound simultaneously, providing rich harmonic support for the foreground audio in the spectrum.
[0048] 2. Isochronous rhythmic envelope: The amplitude of the sound pressure level of each fundamental frequency component that constitutes the acoustic layer is modulated by a periodic, high-frequency intensity envelope (for example, 8Hz), forming a clear and stable rhythmic pulsation in the time domain.
[0049] 3. Dynamic intensity correlation: The overall gain of the intensity envelope shown above shows a high temporal correlation with the dynamic changes in the energy of the foreground audio signal in a specific frequency band (harmonic region).
[0050] Creative summary of this embodiment: The innovation of this embodiment lies in its first proposal and implementation of a complete set of software and hardware collaborative methods that can achieve "from multi-dimensional audio feature analysis to multi-level physical acoustic event rendering." Its core creative contributions include:
[0051] a) Multi-part mapping capability: For the first time, a single melody pitch can be intelligently mapped to control logic for parallel tuning and excitation of multiple physical resonators with multiple pitches (chords).
[0052] b) Unified coding of rhythm and timbre: We innovatively proposed an “acoustic texture reconstruction” algorithm that combines rhythmic information and timbre dynamic information, encoding them uniformly as intensity envelope control of physical excitation actions. This allows the same physical sound event to simultaneously carry information on the three dimensions of harmony, rhythm, and timbre dynamics.
[0053] c) Feasible Collaborative Architecture: A feasible system architecture is proposed and verified that can coordinate a complex physical execution system with intrinsic physical delays and a digital audio playback system to achieve precise time alignment of their outputs.
[0054] d) Versatility of device implementation: The technical solution is implemented using a standard computing architecture, and its technical features include: (i) Compatible with various processor platforms; (ii) support all types of incentive mechanisms defined in claim 2; (iii) Adaptability to various signal transmission media. This method defines a programmable intelligent physical acoustic system capable of generating polyphonic chords and meticulously rendering the micro-dynamics of sound. Its technical complexity and implementation are unmatched by existing technologies.
[0055] Further configuration and expansion of the system The disclosed sound field generation system, based on physical resonators, provides a modular architecture and dynamic tuning capabilities, providing the core foundation for building a high-fidelity physical acoustics execution platform. Future development of this platform will focus on deep collaboration with a new intelligent content analysis engine, fundamentally aiming to provide the system with a set of highly contextually adaptive and application-specific acoustic guidance parameters.
[0056] Further configuration and expansion of the system The sound field generation system disclosed in the present invention, which is based on a physical resonator, has a modular architecture and dynamic tuning capabilities, providing a core foundation for building an open, high-fidelity physical acoustics execution platform. As a further functional extension of the system of the present invention, its parameter determination module (101) can be configured to work in conjunction with an external intelligent content analysis engine, the output form of which is multimodal and can generate both excitation parameters for driving the physical system of the present invention and digital audio signals for conventional playback devices.
[0057] Optimal technical solution for intelligent content analysis engine The "intelligent content analysis engine" may include one or more of the following modules working together, each of which is described below:
[0058] (i) Music structure analysis module This module, implemented using Music Information Retrieval (MIR) technology, deeply analyzes the musical elements of the input audio signal. The module is configured to receive an audio waveform as input and fuse the analyzed musical structure (e.g., pitch, harmony) with the target guiding rhythm. The module's output can be one or more of the following:
[0059] Physical Rendering Output: Generates a set of physical excitation parameters, which includes a set of timbre parameters defining pitch and timbre, as well as rhythm generation instructions. The rhythm generation instructions (anchoring claims 5, 6, and 7) can be acoustic beat instructions (including a static target frequency or a dynamic gradient curve) or isochronous tone sequence instructions (whose instantaneous frequency can be constant or dynamically gradient). The frequency range of both instructions is limited to 1-100Hz.
[0060] Digital rendering output: In an optional embodiment, the module can independently synthesize the parsed music structure and target rhythm into a new digital audio signal containing a guiding rhythm for driving headphones or speakers.
[0061] Note: The inventor reserves the right to claim protection for the proprietary technical solution for generating specific harmonic components in a separate application.
[0062] (ii) Narrative Rhythm Conversion Module: This module is implemented using natural language processing (NLP) technology. Its core task is to analyze speech content and map temporal emotional atmosphere (e.g., "calm" or "tense") to rhythm. This module is configured to receive audio signals containing human speech as input and output one or more of the following:
[0063] Physical Rendering Output: Generates a set of physically excited parameters containing base timbral instructions for matching the narrative atmosphere (e.g., a woody tone for a "calm" atmosphere), and acoustic beat / isochronous tone sequence instructions (anchoring claims 5, 6, and 7) with static or dynamic fading modes.
[0064] Digital Rendering Output: In an optional implementation, this module performs real-time dynamic processing on the input speech audio signal. Using signal processing techniques such as amplitude modulation, it applies the target rhythm corresponding to the identified emotional atmosphere directly to the original audio stream, generating a new digital audio signal that retains the original speech content while also incorporating subtle rhythmic pulsations.
[0065] Note: The inventor reserves the right to claim protection in a separate application for the proprietary algorithm model used to achieve accurate mapping of semantic emotions to rhythmic parameters.
[0066] (iii) Acoustic fingerprint generation and multimodal output module: As an advanced implementation form, this module can be implemented through a deep generative model. Its core task is to learn the "acoustic fingerprint" of the reference audio through deep learning and generatively reshape it to creatively integrate the guiding rhythm with the original acoustic features at the genetic level.
[0067] A typical application scenario is processing vocal singing. The workflow of this module includes the following steps: Deep learning: The system will first deeply learn the singer's unique vibrato characteristics, including its complete acoustic fingerprint such as rate, amplitude, timbre and texture.
[0068] 1. Intelligent analysis: The system will then conduct an in-depth analysis of these learned acoustic fingerprints to deconstruct the patterns and structures behind them that are used to express musical emotions.
[0069] Generative Output: Finally, based on this deep analysis, the module can creatively generate highly synchronized output in one or more of the following modes, depending on the system configuration:
[0070] Hybrid rendering output mode: In this mode, the module generates two signals in parallel: (a) A physical excitation instruction stream (to the physical resonator network): This generates a set of parameters containing timbre instructions and rhythm generation instructions. The rhythm generation instructions, anchored in claims 5, 6, and 7, can be acoustic beat instructions or isochronous tone sequence instructions that can achieve static or dynamic gradient effects.
[0071] (b) Digital audio signal flow (to external speakers / headphones): Generates a new audio waveform that retains the original characteristics in terms of timbre, but has been intelligently reshaped rhythmically (such as vibrato) to contain a guiding rhythm that complements or is synchronized with the aforementioned physical rhythm.
[0072] 2. Independent Digital Rendering Output Mode: In this mode, the module operates independently, generating only a digital audio signal. This signal has been processed to incorporate the guiding rhythm defined by the aforementioned rhythm generation instructions (anchored in claims 5, 6, and 7), with either isochronous or out-of-step effects (supporting both static and dynamic gradations).
[0073] Note: The inventor reserves the right to claim protection for the proprietary technical solution for extracting acoustic fingerprints in a separate application.
Claims
1. A method for generating a perceptible rhythmic characteristic sound field, characterized in that: include: (a) determining a set of acoustic guidance parameters, wherein the parameters include at least one of rhythm, fundamental frequency or harmonic components; (b) driving an excitation mechanism to excite one or more physical resonators according to the acoustic guidance parameters to generate an acoustic environment, wherein the timbre spectrum characteristics and sound energy attenuation characteristics of the acoustic environment are dominated by the material properties, geometric structure and damping characteristics of the one or more physical resonators; (c) forming a physical sound field environment with perceptible rhythmic characteristics through an acoustic beat effect or isochronous sound effect endogenous to the acoustic environment, wherein the frequency range of the acoustic beat effect or isochronous sound effect covers 1-100 Hz.
2. The method according to claim 1, wherein: The actuation mechanism includes at least one selected from the group consisting of: an electromagnetic actuation mechanism; a piezoelectric actuation mechanism; a fluid actuation mechanism; and a mechanical contact actuation mechanism.
3. The method according to claim 1, wherein: The one or more physical resonators include at least one selected from the group consisting of: a solid resonator made of crystal, glass, metal, stone, wood, ceramic or a composite thereof; a tensioned string; a tensioned membrane; a gas within a defined volume; and a liquid within a defined volume.
4. The method according to claim 1, wherein: The method also includes presenting a digital audio signal together with the acoustic environment generated by the physical resonant body to form a composite sound field, wherein the digital audio signal is selected from the group consisting of: a) the original audio content; and b) a rhythmically modulated audio signal generated by real-time dynamic processing of the original audio content to embed a second perceptible rhythmic feature defined by the acoustic guidance parameters therein.
5. The method according to claim 1, wherein: The acoustic guidance parameters are configured to drive an excitation mechanism to generate at least two vibrations of different frequencies that can interfere with each other on one or more physical resonators, thereby forming the acoustic beat effect through physical sound wave interference, and the frequency difference between the vibrations of different frequencies is between 1-100 Hz.
6. The method according to claim 1, wherein: The acoustic guidance parameters define a series of discrete excitation events, where the time intervals between adjacent events are a sequence of {t1, t2,..., t n The following conditions are met to form the isochronous sound effect: (1) Each time interval t k is in the interval [10, 1000] milliseconds; (2) the instantaneous stimulation frequency f corresponding to each time interval k = 1 / t k ; and (3) the instantaneous stimulation frequency f k The sequence constitutes a continuously changing target rhythm frequency sequence.
7. The method according to claim 1, wherein: The acoustic guidance parameters are further configured to enable the rhythmic frequency of the acoustic beat effect or isochronous sound effect to undergo continuous or step-wise dynamic gradual change according to a preset frequency change curve.
8. A system for generating a perceptible rhythmic characteristic sound field, characterized in that: The method comprises a processor and a memory, wherein the processor is configured to execute the method according to any one of claims 1 to 7.
9. The system according to claim 8, characterized in that Also included is a dynamic tuning module configured to dynamically adjust the resonant frequency of the physical resonator by changing one or more effective physical parameters of the physical resonator in real time; The effective physical parameters include: effective mass, effective length, effective stiffness or vibration boundary conditions.
10. A sound field rhythm control device, characterized in that: The device comprises hardware components and is configured to implement the system according to claim 8 or 9.
Citation Information
Cited By
Self-adaptive physical acoustic environment generation method and system
CN120932624A