Novel sound synthesizers

WO2026077365A1PCT designated stage Publication Date: 2026-04-16LIU PO-TING
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/126389
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-10-08
Filing Date
2025-10-08
Publication Date
2026-04-16

AI Technical Summary

Technical Problem

Existing techniques struggle to effectively reconstruct sound from auditory-coded pulse time-frequency representations, especially due to the low resolution of intermediate features in the time domain, leading to inaccurate reconstruction. Existing models suffer a 65% reduction in PESQ scores.

Method used

We employ an end-to-end ANN-based model to reconstruct sound directly from neuronal spiking activity. By utilizing high temporal resolution auditory neural maps and combining them with an improved sampler network and vocoder model, we achieve high-fidelity sound waveform reconstruction.

Benefits of technology

It achieves direct reconstruction of high-fidelity acoustic waveforms with a PESQ score exceeding 4.2 and zero-shot performance, surpassing the performance of existing vocoder models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025126389_16042026_PF_FP_ABST
    Figure CN2025126389_16042026_PF_FP_ABST
Patent Text Reader

Abstract

Sound reconstruction from corresponding auditory neurograms was a long-standing problem in computational neurosciences and auditory modeling fields, and was solved by my end-to-end artificial neural network models. This is a novel system for sound synthesis utilizes the modern neural vocoders and spiking-based time-frequency representations derived from auditory neurograms. This method introduces the auditory neurograms as the novel input, replacing the mel-spectrogram in conventional vocoding approaches. This method provides high-fidelity sound quality of synthesized sound, compared to the existing two-stage models on the sound reconstruction task and the modern neural vocoders. This method also provide zero-shot performance on unseen datasets. This method can be used for the applications of a novel Sound Synthesizer using auditory spiking data and Speech NeuroProsthesis using auditory spiking activities.
Need to check novelty before this filing date? Find Prior Art

Description

Novel sound synthesizers技术领域

[0001] This invention relates to sound synthesizers using deep learning algorithms, and more particularly to methods and systems for processing novel audio input types through innovative strategies of implementing re-sampler networks and sound synthesizer models to solve the reconstruction task from auditory-coded spiking time-frequency representation, and also to improve sound synthesizing model performance in audio signal processing applications.背景技术

[0002] The sound reconstruction task from auditory neurograms was a long-standing unsolved research question. The main challenge of sound reconstruction task is to inverse the effects of the stochastic responses of auditory nerve fibers (ANFs) and the adaptations by the highly non-linear and bidirectional interactions in auditory pathways.

[0003] The structure of existing models (Zai, et al., 2015; Liu, et al., 2024) to address this questions were two-stage model. Their two-stage structure was illustrated in Figure 5. They used intermediate features in between the stages; however, these intermediate features was limited by their lower resolution in time domain, leading to inaccurate reconstruction. Previous approaches using intermediate features resulted in a PESQ score reduction of 65% compared to direct synthesis.

[0004] Therefore, the novel of the proposed system is taking the advantage of auditory neurograms: the high time-resolution. I developed an ANN-based end-to-end model that reconstructs sound with high fidelity directly from neuronal spiking activities. My decoding model (Liu, P.T.B., 2024) has no equal competition, even compared to the latest work (Liu, et al., 2024). My models’ sound quality can match the performance of modern speech synthesis models using neural vocoders. Thus, my decoding model is the State-of-the-Art (SOTA) on this task. Furthermore, this proposed model has advanced neural engineering and auditory neuroscience by providing a deeper understanding of how the auditory periphery contributes to human speech perception.发明内容

[0005] The present application provides a system to synthesize the waveforms from time-frequency representation. Please refer to Figure 4, which depicts inputs and outputs of the provided system, i.e., an ANN (artificial neural network) model.

[0006] Input of the provided system (Figure 4): time-frequency representation from one or any combination of following: neural models, mathematical methods, or hardware devices, recoding from neurons, “auditory spectrogram” , neuro-gram, spike-gram, and nerve-gram.

[0007] Output of the provided system (Figure 4): waveforms.

[0008] Please also refer to Figure 1, which illustrates a system overview diagram of an embodiment of the present application. The provided system as well as the mentioned neural model(s) may be implemented in one computer or a plurality of interconnected computers. Person having ordinary skill in the art of computer architecture and computer organization can understand that the computer may include but not limit to processor(s), memories and input and output devices such as network interface, audio processor(s), microphone, speaker, etc. Date modulated or formatted in time-frequency representation may be received by one or more of the input devices of the computer(s). And the waveforms may be outputted by one or more of the output devices of the computer(s).

[0009] In some examples, the computer(s) may be commercially available or on-the-shelf computers. In some cases, the computer(s) may be constructed as ASIC (application specific integrated circuits). The present application does not limit how the system is implemented.

[0010] My artificial neural network models synthesize sounds from the neuronal spiking activities in the peripheral auditory pathways, which are simulated based on humans with normal hearing. To address the reconstruction problem from auditory spiking data, the models must handle the stochastic responses of auditory nerve fibers (ANFs) and compensate for the adaptations by the highly non-linear and bidirectional interactions in the pathways.

[0011] The provided system may be implemented by one of the following neural vocoder models, including various neural vocoder models. However, person having ordinary skill in the art can understand that the provided system does not limit to the following neural vocoder models or their variants and equivalents.

[0012] - WaveFit

[0013] - Fftnet

[0014] - Periodnet

[0015] - iSTFTnet

[0016] - DDSP

[0017] - Paranet

[0018] - Autoregressive Models

[0019] WaveNet

[0020] WaveRNN

[0021] SampleRNN

[0022] FeatherWave

[0023] - Source-filter models

[0024] LPCnet

[0025] - Diffusion models

[0026] Diffwave

[0027] WaveGrad

[0028] Prior Grad

[0029] - Glow / Flow models

[0030] WaveFlow

[0031] WaveGlow

[0032] FloWavenet

[0033] Squeezewave

[0034] Wavenode

[0035] ClariNet

[0036] - VAE

[0037] WaveVAE

[0038] - Transformer-based

[0039] GoodBye WaveNet --A Language Model for Raw Audio with Context of 1 / 2 Million Samples (2022)

[0040] A compact transformer-based GAN vocoder (2022)

[0041] - GAN models

[0042] (Parallel) WaveGAN

[0043] HiFiGAN

[0044] (Multi-band) MelGAN

[0045] VocGAN

[0046] BigVGAN

[0047] Key Advantages:

[0048] The proposed models solved the long-standing question, which is sound reconstruction from auditory neurograms.

[0049] The proposed models provide high-fidelity synthesized sound with the average PESQ scores above 4.2.

[0050] The proposed models provide zero-shot performance on unseen datasets.附图说明

[0051] Figure 1 is a diagram demonstrating the structure of proposed models’ .

[0052] Figure 2 is a chart showing objective metrics evaluated on the output waveforms from different variants of proposed models.

[0053] Figure 3 is a chart showing generalization test on unseen dataset.

[0054] Figure 4 is a diagram illustrating the novel structure of the proposed models.

[0055] Figure 5 is a diagram illustrating the structure of existing 2-stage models on the sound reconstruction task.具体实施方式

[0056] The Spiking-based time-frequency representation were computed by a peripheral auditory model or an auditory periphery model (APM) or any variants and equivalents with the normal hearing or various hearing conditions. The APM simulated thirty thousand ANFs, which is the number of ANFs in the normal human ear, at high-, medium-, and low-spontaneous firing rates, and with 80 cochlear filters that their characteristic frequencies (CFs) range is between 50-12k and equally spaced on an Equivalent Rectangular Bandwidth (ERB)-rate scale. Thirty thousand ANFs were simulated in normal hearing conditions. The APM also simulated the Acoustic Reflex (AR) and Medial OlivoCochlear Reflex (MOCR) in the pathways. The sound stimuli consisted of LJSpeech dataset and they were presented at 70 dB sound pressure level (SPL) or desired SPL when computing the ANFs' responses. The spike trains produced by the ANFs were sampled at 22,050 Hz.

[0057] In dataset pre-processing, first 20 sound files were in the test-set, and all the others were used in training.

[0058] In pre-processing, input data were summed into separate channels for each characteristic frequency, and re-binned to match the desired time-resolution, such as 22, 050 samples per second, to form auditory neurograms. The pre-processed neurograms have 240 channels and were presented at the same time-resolution of output waveforms.

[0059] The proposed model is consisted of a re-sampler network and modified vocoder models to solve the sound reconstruction tasks from auditory neurograms. If spike-sampling-per-second of input auditory neurograms is the same with the output waveforms’s ampling rate, the model will direct take the input as the input of remaining networks.

[0060] The pre-processed auditory neurogram were passed through a re-sampler network, which was implemented with two-layers of 2D convolutions (with kernel= [3, 3] , stride= [1, 1] , and padding= [1, 1] ) and leaky RELU activation function to match the sampling rate of output waveforms.

[0061] To synthesize sound from auditory neurograms, I modified Artificial Neural Network-based sound synthesizer models, e. g. DiffWave, to decode the spiking trains of ANFs’ responses. The output of re-sampler network were used as the input to the modified vocoder networks without their up-sampling networks.

[0062] The modified ANN-based sound synthesizer models of this system is diffusion-based neural vocoder. The modified ANN-based sound synthesizer models using the processed auditory neurograms with 240 input channels is implemented as a 1D convolutional neural network with 30 residual layers, each using a kernel size of 3 and dilations cycling through powers of two [1, 2, 4, …, 1024] to capture long-range temporal dependencies. Each residual block employs residual and skip channels of 64, a gated-tanh activation mechanism, and layer normalization applied to both the residual and conditioning paths. The diffusion process supports up to 200 time steps, and each step index t is embedded via a sinusoidal positional encoding, followed by two linear layers with SiLU activations to produce a time embedding added to every residual block.

[0063] Training procedure

[0064] During training, the model predicted the corresponding waveform and was optimized using an L1 or L2 loss between the predicted and target waveform. The training employed the Adam optimizer (β1=0.9, β2=0.999) with a learning rate of 1e-4 and a batch size of 16.

[0065] During inference, the model samples from Gaussian noise and performs iterative denoising over the specified number of diffusion steps 50 to generate time-domain waveforms.

[0066] Performance Results

[0067] To evaluate model performance, objective metrics such as Perceptual Evaluation of Speech Quality (PESQ) and Short-Time Objective Intelligibility (STOI) were assessed between the original and reconstructed sound waveforms. Figure 2 presents the objective metric results from different model with different input size. Consequently, the proposed models deliver high-fidelity synthesized sound.

[0068] Generalization test: Zero-shot performance

[0069] A generalization test utilizing a zero-shot approach was conducted to address potential overfitting. The unseen test sets included four stimulus types: VCTK (British speaker) , MCDC8 (Mandarin speech) , and excerpts of single musical instruments from the Medley-solos-DB dataset, for example, piano. Zero-shot performance of my proposed model were showed in Figure 3. Based on these results, the proposed system and method outperform state-of-the-art sound synthesis models.

[0070] In conclusion, this model synthesizes acoustic waveforms with high fidelity directly from neuronal spiking activity within human auditory pathways, compensating for attenuation effects caused by the Acoustic Reflex (AR) and Medial OlivoCochlear Reflex (MOCR) .

[0071] Model structure comparison with existing two-stage models

[0072] The sound reconstruction task from auditory spike trains remained a long-standing unsolved research question until my proposed models. There were some researchers tried to solve this problem using two-stage approaches. In 2015, the paper (Zai, et al., 2015) presented their two-stage model using signal processing techniques, which achieved the average PESQ scores of between 1.5 and 2.0 and the average STOI scores of between 0.6 and 0.8. In 2024, the team (Liu, et al., 2024) proposed a two-stage model with “aconventional signal processing algorithm” (Chen, &Toda, 2024) . Their results don't even have the same precise envelope of the reconstructed waveform under the Normal Hearing (NH) condition, compared to its original waveform in their paper, but my proposed model can. Unfortunately, since none of these existing models have successfully solved the problem.

[0073] My proposed model provides much better performance than the recent works. My proposed model is an end-to-end model, while the existing framework (Liu, et al., 2024) is a two-stage model. They used the signal processing-based WORLD vocoder designed for computing-power-limited devices and had the average PESQ scores of 3.398±0.124 on the LJSpeech dataset (Fang, 2021) , which provides the upper bound of their two-stage model’s  performance on the same dataset used for training my model. These two-stage model cannot compete with my proposed models using end-to-end ANN structure. Therefore, my proposed model (Liu, P. T. B., 2024) has solved the reconstruction problem, and is the SOTA on this sound reconstruction task.

[0074] Alternative Embodiments

[0075] A neurogram is recorded from auditory-related neurons or generated using auditory modeling techniques that simulates or mimics auditory nerve fibers (ANFs) and other functions in pathways using neural models, spiking neural networks, mathematical methods, biological-plausible neuronal models, or hardware devices, and represented in form with time-frequency characteristics. On the other terms, it can be called “auditory spectrogram” , neuro-gram, spike-gram, and nerve-gram.

[0076] An auditory neurogram can be generated with normal hearing or hearing conditions using APMs, and the output should be use of the paired simulated perceived sound by the specific hearing conditions.

[0077] These are the common setting that can used in APM. The APM simulated thirty thousand ANFs, which is the number of ANFs in the normal human ear, at high-, medium-, and low-spontaneous firing rates, and with 26, 80 or more cochlear filters that their characteristic frequencies (CFs) range is between 70 and 8, 000 Hz or 50-12k and equally spaced on an Equivalent Rectangular Bandwidth (ERB) -rate scale. Thirty thousand ANFs were simulated in normal hearing, hearing-impired or hearing-loss conditions. The APM also can simulate the Acoustic Reflex (AR) , and Medial OlivoCochlear Reflex (MOCR) . The APMs can include higher stages in auditory pathways, such inferior (IC) , Medial Geniculate Nucleus (MGB) , Auditory Cortices.

[0078] The following settings are in the various configurations of auditory neurogram that can use as input in my proposed system.

[0079] When preparing dataset, the Sound Pressure Leve range of the acoustic stimulus can be between 20 and 120 dB.

[0080] The use CFs of an auditory neurogram can be represented using Mel, Bark, ERB-rate, or Greenwood scales.

[0081] An auditory neurogram can be consisted of one of three spontaneous rates, namely HSR only, MSR only, or LSR only.

[0082] An auditory neurogram can be consisted of mean-rate of three spontaneous rates.

[0083] When producing an auditory neurogram, the spike trains produced by the ANFs were sampled at 16k, 22, 050 or 24k Hz.

[0084] The channels of a neurogram can be presented in order of ascending, descending, or pre-fixed randomization.

[0085] Design of re-sampler networks for different dimensions of input

[0086] Implementations of re-sampler networks can be one of the following specifications with specific dimensions of input:

[0087] 1. For the example embodiment in earlier section, taking 2D input with dimension= [80, t] , the re-sampler networks were implemented with two-layers of 2D convolutions (with kernel= [3, 3] , stride= [1, 1] , and padding= [1, 1] ) and leaky RELU activation function.

[0088] 2. To take 3D input of auditory neurograms, for example [3, 80, t] , the re-sampler networks were implemented with a layer of 3D convolutions (with kernel= [3, 3, 3] , stride= [1, 1, 1] , and padding= [0, 1, 1] ) and leaky RELU activation function. This approach can take the advantage of three SRs to extract features across three SRs using convolutions. The results using 3D input are shown in the Figure 2.

[0089] 3. If the spike-sampling-per-second of input auditory neurograms is the same with the output waveforms’ sampling rate, the model can direct take the input into remaining networks. So one of practical implementations of this system can be constructed without re-sampler networks to solve the problem of sound reconstruction from auditory spiking data, and this variant still provide almost the same quality of synthesized sound. The results in this case are shown in the Figure 2.

[0090] The provided system may be implemented by one of the following neural vocoder models, including various neural vocoder models. However, person having ordinary skill in the art can understand that the provided system does not limit to the following neural vocoder models or their variants and equivalents.

[0091] - WaveFit

[0092] - Fftnet

[0093] - Periodnet

[0094] - iSTFTnet

[0095] - DDSP

[0096] - Paranet

[0097] - Autoregressive Models

[0098] WaveNet

[0099] WaveRNN

[0100] SampleRNN

[0101] FeatherWave

[0102] - Source-filter models

[0103] LPCnet

[0104] - Diffusion models

[0105] Diffwave

[0106] WaveGrad

[0107] Prior Grad

[0108] - Glow / Flow models

[0109] WaveFlow

[0110] WaveGlow

[0111] FloWavenet

[0112] Squeezewave

[0113] Wavenode

[0114] ClariNet

[0115] - VAE

[0116] WaveVAE

[0117] - Transformer-based models

[0118] GoodBye WaveNet --A Language Model for Raw Audio with Context of 1 / 2 Million Samples (2022)

[0119] A compact transformer-based GAN vocoder (2022)

[0120] - GAN models

[0121] (Parallel) WaveGAN

[0122] HiFiGAN

[0123] (Multi-band) MelGAN

[0124] VocGAN

[0125] BigVGAN, BigVGAN-v2

Claims

1.A method for synthesizing acoustic waveforms from spiking-based time-frequency representations of the corresponding acoustic waveforms, comprising:receiving a neurogram comprising spike trains that are simulated or recorded from the neural activities in auditory-related pathways as input data;inputting the neurogram into an artificial neural network (ANN) -based sound synthesis model with or without a re-sampler network;generating, by the ANN-based sound synthesis model, a synthesized acoustic signal from the neurogram; and outputting the synthesized acoustic signal as a representation of the original acoustic stimuli.2.The method of claim 1, wherein the neurogram is recorded from auditory-related neurons or generated using auditory modeling techniques that simulates auditory nerve fibers (ANFs) and other functions in pathways using neural models, spiking neural networks, mathematical methods, biological-plausible neuronal models, or hardware devices, and represented in the form with time-frequency characteristics.3.The method of claim 1, wherein the re-sampler network us applied on the input spiking data to prepare for the ANN-based sound synthesis model.4.The method of claim 1, wherein the ANN-based sound synthesis model is a vocoder-based model without its first layers, which plays the role of up-sampling.5.A system for synthesizing acoustic stimuli from auditory-related spiking data, comprising:receiving an unprocessed auditory neurogram to form an input data;inputting the input data into a modified artificial neural network (ANN) -based sound synthesis model;generating, by the ANN-based sound synthesis model, a synthesized acoustic signal from the neurogram;and outputting the synthesized acoustic signal as a representation of the original acoustic stimuli.