Neural network with bayesian-graph retrieval augmented generation for advanced maximal entropy media compression processing
A neural network governed by PDEs with Bayesian evaluation and multi-agent verification addresses the limitations of existing networks, enhancing interpretability and accuracy, and facilitates efficient media compression.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- ZON GLOBAL IP INC
- Filing Date
- 2025-10-24
- Publication Date
- 2026-05-07
AI Technical Summary
Existing neural networks lack versatility across various modalities and suffer from limited interpretability and accuracy in generating responses.
A neural network defined by a neuromorphic field governed by partial differential equations (PDEs) with adjustable hyperparameters, combined with Bayesian evaluation and multi-agent verification for enhanced accuracy and relevance, and a system for media compression using deep learning and entropy maximization.
The solution enables neural networks to operate across diverse modalities with improved interpretability and accuracy, and achieves efficient media compression while maintaining compatibility with existing ecosystems.
Smart Images

Figure US2025052421_07052026_PF_FP_ABST
Abstract
Description
NEURAL NETWORK WITH BAYESIAN-GRAPH RETRIEVAL AUGMENTED GENERATION FOR ADVANCED MAXIMAL ENTROPY MEDIA COMPRESSION PROCESSING CROSS REFERENCES TO RELATED APPLICATIONS
[0001] This application is related to and claims priority from the following U. S. patents and patent applications. This application claims priority' to and the benefit of U. S. Patent Application No. 18 / 934,967, filed November 1, 2024, U. S. Patent Application No. 18 / 935,013, filed November 1, 2024 and U. S. Patent Application No. 18 / 935,039, filed November 1, 2024, each of which is incorporated herein by reference in its entirety.BACKGROUND OF THE INVENTION
[0002] 1. Field of the Invention
[0003] The present invention relates to neural networks, and more specifically to neural networks for compression and a continuously improving Al system.
[0004] 2. Description of the Prior Art
[0005] It is generally known in the prior art to provide neural networks and other forms of deep learning to solve various types of tasks.
[0006] Prior art patent documents include the following:
[0007] US Patent Pub. No. 2022 / 0253671 for Graph neural diffusion by inventors Chamberlain et al., filed February 7, 2022 and published August 11, 2022, discloses improved graph neural networks (GNNs) including defining a GNN architecture based on a discretized non-Euclidean diffusion partial differential equation (PDE) such that evolution of feature coordinates represents message passing layers in a GNN and evolution of positional coordinates represents graph rewiring. The GNN being based on both position and feature coordinates has their evolution derived from Beltrami flow. Tlie Beltrami flow is modeled using a Laplace-Beltrami operator, which is a generalization of the Laplace operator to functions defined on submanifolds in Euclidean space and on Riemannian manifolds. The discretization of the spatial component of the Beltrami flow offers a principled view on positional encoding and graph rewiring, whereas the discretization of the temporal component can replace GNN layers with more flexible adaptive numerical schemes. Based on this model, Beltrami Neural Diffusion (BLEND) that generalizes a broad range of GNN architectures is introduced; BLEND shows state-of-the-art performance on many benchmarks.
[0008] US Patent Pub. No. 2022 / 0092413 for Method and system for relation learning by multi-hop attention graph neural network by inventors Wang et al., filed May 24, 2021 and published March 24, 2022, discloses a system and method for completing knowledge graph. The system includes a computing device, the computing device has a processer and a storage devicestoring computer executable code. The computer executable code is configured to: provide an incomplete knowledge graph comprising a plurality of nodes and a plurality of edges, each of the edges connecting two of the plurality of nodes; calculate an attention matrix of the incomplete knowledge graph based on one-hop attention between any two of the plurality of the nodes that are connected by one of the plurality of the edges; calculate multi-head diffusion attention for any two of the plurality of nodes from the attenti on matrix; obtain updated embedding of the incomplete knowledge graph using the multi-head diffusion attention; and update the incomplete knowledge graph to obtain updated knowledge graph based on the updated embedding.
[0009] US Patent Pub. No. 2022 / 0198278 for System for continuous update of advectiondiffusion models with adversarial networks by inventors O’Donncha et al., filed December 23, 2020 and published June 23, 2022, discloses a computing device configured for automatic selection of model parameters including a processor and a memory coupled to tire processor, lire memory' stores instructions to cause the processor to perform acts including providing an initial set of model parameters and initial condition information to a model based on historical data. A model generates data based on the model parameters and tire initial condition information. After determining whether the model -generated data is similar to an observed data, updated model parameters are selected for input to the model based on the determined similarity.
[0010] US Patent Pub. No. 2024 / 0152669 for Physics-enhanced deep surrogate by inventors Pestourie et al., filed November 8, 2022 and published May 9, 2024, discloses surrogate training including receiving a parameterization of a physical system, where the physical system includes real physical components and the parameterization having corresponding target property in the physical system, lire parameterization can be input into a neural network, where the neural network generates a different dimensional parameterization based on the input parameterization. The different dimensional parameterization can be input to a physical model that approximates the physical system. The physical model can be run using the different dimensional parameterization, where the physical model generates an output solution based on the different dimensional parameterization input to the physical model. Based on the output solution and the target property, the neural netw ork can be trained to generate the different dimensional parameterizati on.
[0011] US Patent No. 11,645,356 for Deep learning for partial differential equation (PDE) based models by inventors O’Donncha et al., filed September 4, 2018 and issued May 9, 2023, discloses embodiments for deep learning for partial differential equation (PDE)-based models by a processor. A trained forecasting model and consistency constraints may be generated using a PDE -based model, a discretization of the PDE-based model, historical inputs the of the PDE-based model, and a representation of consistency constraints to generate a predictive output.
[0012] US Patent No. 10,963,540 for Physics informed learning machine by inventors Raissi et al., filed June 2, 2017 and issued March 30, 2021, discloses a method for analyzing an object includes modeling the object with a differential equation, such as a linear partial differential equation (PDE), and sampling data associated with the differential equation. Tire method uses a probability distribution device to obtain the solution to the differential equation. Ihe method eliminates use of discretization of the differential equation.
[0013] Chinese Patent Pub. No. 117521729 for Partial differential equation diffusion attenuation-based graph neural netw ork optimization method and related device, filed October 31, 2023 and published February’ 6, 2024, discloses a graph neural network optimization method based on partial differential equation diffusion attenuation and a related device, wherein the invention adopts an encoder-decoder (encoder-decoder) structure, uses an encoder to obtain characteristic embedding, and uses a decoder to decode a finally output hidden vector; the depth of tire model is controlled by the propagation time, and the time is discretized by introducing super parameters so as to obtain a model of a deep framework; in order to learn the diffusion coefficient in diffusion more effectively and increase the diffusion threshold range, the invention also applies a shared gating attention mechanism to learn the diffusion coefficient effectively according to the characteristics of the adjacent nodes, and increases the diffusion possible threshold range by selecting a proper function.
[0014] The article “Graph Neural Networks as Neural Diffusion PDEs” by authors Chamberlain and Bronstein, published July 21, 2021, concerns diffusion PDEs arising in many phy sical processes involving the transfer of “stuff1(whether energy or matter), or more abstractly, information. In image processing, one can exploit this interpretation of diffusion as linear low-pass filtering for image denoising. However, such a filter, when removing noise, also undesirably blurs transitions betw een regions of different color or brightness (“edges”). An influential insight of Pietro Perona and Jitendra Malik was to consider an adaptive diffusivity coefficient inversely dependent on the norm of the image gradient | Vx|. This w;p. diffusion is strong in “flat” regions (where | V x|~0) and weak in the presence of brightness discontinuities (where | Vx| is large). The result was a nonlinear filter capable of removing noise from the image while preserving edges.
[0015] The article “Generalized neural closure models with interpretability” by authors Gupta and Lemiusiaux, published June 30, 2023, concerns improving the predictive capability and computational cost of dynamical models is often at the heart of augmenting computational physics with machine learning (ML). However, most learning results are limited in interpretability and generalization over different computational grid resolutions, initial and boundary’ conditions, domain geometries, and physical or problem-specific parameters. In thepresent study, we simultaneously address all these challenges by developing the novel and versatile methodology of unified neural partial delay differential equations. We augment existing / low-fidelity dynamical models directly in their partial differential equation (PDE) forms with both Markovian and non-Markovian neural network (NN) closure parameterizations. The melding of the existing models with NNs in the continuous spatiotemporal space followed by numerical discretization automatically allows for the desired generalizability. Tire Markovian term is designed to enable extraction of its analytical form and thus provides interpretability. The non-Markovian terms allow accounting for inherently missing time delays needed to represent the real world. Tire flexible modeling framework provides full autonomy for the design of the unknown closure terms such as using any linear-, shallow-, or deep-NN architectures, selecting the span of the input function libraries, and using either or both Markovian and non-Markovian closure terms, all in accord with prior knowledge. We obtain adjoint PDEs in the continuous form, thus enabling direct implementation across differentiable and non-differentiable computational physics codes, different ML frameworks, and treatment of nonuniformly-spaced spatiotemporal training data. We demonstrate the new generalized neural closure models (gnCMs) framework using four sets of experiments based on adverting nonlinear waves, shocks, and ocean acidification models. Tire learned gnCMs discovermissing physics, find leading numerical error terms, discriminate among candidate functional forms in an interpretable fashion, achieve generalization, and compensate for the lack of complexity in simpler models. Finally, we analyze the computational advantages of the new framework.
[0016] The article “Autoregressive Renaissance in Neural PDE Solvers” by author Lee, published May 1, 2023, concerns recent developments in the field of neural partial differential equation (PDE) solvers having placed a strong emphasis on neural operators. However, the paper Message Passing Neural PDE Solver by Brandstetter et al. published in ICLR 2022 revisits autoregressive models and designs a message passing graph neural network that is comparable with or outperforms both the state-of-the-art Fourier Neural Operator and traditional classical PDE solvers in its generalization capabilities and performance. This blog post delves into the key contributions of this work, exploring the strategies used to address the common problem of instability in autoregressive models and the design choices of the message passing graph neural network architecture.
[0017] US Patent No. 12,001,950 for Generative adversarial network based audio restoration by inventors Zhang et al., filed March 12, 2019 and issued June 4, 2024, discloses mechanisms for implementing a generative adversarial network (GAN) based restoration system. A first neural network of a generator of the GAN based restoration system is trained to generate an artificial audio spectrogram having a target damage characteristic based on an input audiospectrogram and a target damage vector. An original audio recording spectrogram is input to the trained generator, where the original audio recording spectrogram corresponds to an original audio recording and an input target damage vector. The trained generator processes the original audio recording spectrogram to generate an artificial audio recording spectrogram having a level of damage corresponding to the input target damage vector. A spectrogram inversion module converts the artificial audio recording spectrogram to an artificial audio recording waveform output.
[0018] US Patent No. 11,514,925 for Using a predictive model to automatically enhance audio having various audio quality issues by inventors Jm et al., filed April 30, 2020 and issued November 29, 2022, discloses operations of a method including receiving a request to enhance a new source audio. Responsive to the request, the new source audio is input into a prediction model that was previously trained, draining the prediction model includes providing a generative adversarial network including the prediction model and a discriminator. Training data is obtained including tuples of source audios and target audios, each tuple including a source audio and a corresponding target audio. During training, the prediction model generates predicted audios based on the source audios. Training further includes applying a loss function to the predicted audios and the target audios, where the loss function incorporates a combination of a spectrogram loss and an adversarial loss. The prediction model is updated to optimize that loss function. After training, based on the new source audio, the prediction model generates a new predicted audio as an enhanced version of tire new source audio.
[0019] US Patent No. 11,657,828 for Method and system for speech enhancement by inventor Quillen, filed January 31, 2020 and issued May 23, 2023, discloses improving speech data quality through training a neural network for de-noising audio enhancement. One such embodiment creates simulated noisy speech data from high quality speech data. In turn, training, e.g., deep normalizing flow training, is performed on a neural network using the high quality speech data and the simulated noisy speech data to train the neural network to create de-noised speech data given noisy speech data. Performing the training includes minimizing errors in the neural network according to at least one of (i) a decoding error of an Automatic Speech Recognition (ASR) system processing current de-noised speech data results generated by the neural network during the training and (ii) spectral distance between the high quality speech data and the current de-noised speech data results generated by the neural network during the training.
[0020] US Patent Pub. No. 2024 / 0055006 for Method and apparatus for processing of audio data using a pre-configured generator by inventor Biswas, filed December 15, 2021 and published February 15, 2024, discloses a method for setting up a decoder for generating processed audio data from an audio bitstream, the decoder comprising a Generator of aGenerative Adversarial Network, GAN, for processing of the audio data, wherein the method includes the steps of (a) pre-configuring the Generator for processing of audio data with a set of parameters for the Generator, tire parameters being determined by training, at training time, the Generator using the full concatenated distribution; and (b) pre-configuring the decoder to determine, at decoding time, a truncation mode for modifying the concatenated distribution and to apply the determined truncation mode to the concatenated distribution. Described are further a method of generating processed audio data from an audio bitstream using a Generator of a Generative Adversarial Network, GAN, for processing of the audio data and a respective apparatus. Moreover, described are also respective systems and computer program products.
[0021] US Patent Pub. No. 2024 / 0203443 for Efficient frequency-based audio resampling for using neural networks by inventors Jjoshi et al., filed December 19, 2022 and published June 20, 2024, discloses systems and methods relating to the enhancement of audio, such as through machine learning-based audio super-resolution processing. An efficient resampling approach can be used for audio data received at a lower frequency than is needed for an audio enhancement neural network.
[0022] US Patent Pub. No. 2023 / 0298593 for Method and apparatus for real-time sound enhancement by inventors Ramos et al., filed May 23, 2023 and published September 21, 2023 discloses a system, computer-implemented method and apparatus for training a machine learning, ML, model to perform sound enhancement for a target user in real-time, and a method and apparatus for using tire trained ML model to perform sound enhancemen t of audio signals in realtime. Advantageously, the present techniques are suitable for implementation on resource-constrained devices that capture audio signals, such as smartphones and Internet of Things devices.
[0023] US Patent No. 10,991,379 for Data driven audio enhancement by inventors Hijazi et al., filed June 22, 2018 and issued April 27, 2021, discloses systems and methods for audio enhancement. For example, methods may include accessing audio data; determining a window of audio samples based on the audio data; inputting the window of audio samples to a classifier to obtain a classification, in which the classifier includes a neural network and the classification takes a value from a set of multiple classes of audio; selecting, based on the classification, an audio enhancement network from a set of multiple audio enhancement networks; applying the selected audio enhancement network to the window' of audio samples to obtain an enhanced audio segment, in which the selected audio enhancement network includes a neural network that has been trained using audio signals of a type associated with the classification; and storing, playing, or transmitting an enhanced audio signal based on the enhanced audio segment.
[0024] US Patent No. 10,460,747 for Frequency based audio analysis using neural networks by inventors Roblek et al,, filed May 10, 2016 and issued October 29, 2019, discloses methods, systems, and apparatus, including computer programs encoded on computer storage media, for frequency based audio analysis using neural networks. One of the methods includes training a neural network that includes a plurality of neural network layers on training data, wherein the neural network is configured to receive frequency domain features of an audio sample and to process the frequency domain features to generate a neural network output for the audio sample, wherein the neural network comprises (i) a convolutional layer that is configured to map frequency domain features to logarithmic scaled frequency domain features, wherein the convolutional layer comprises one or more convolutional layer filters, and (ii) one or more other neural network layers having respective layer parameters that are configured to process the logarithmic scaled frequency domain features to generate the neural network output.
[0025] US Patent No. 11,462,209 for Spectrogram to waveform synthesis using convolutional networks by inventors Arik et al., filed March 27, 2019 and issued October 4, 2022, discloses an efficient neural network architecture, based on transposed convolutions to achieve a high compute intensity’ and fast inference. In one or more embodiments, fortraining of the convolutional vocoder architecture, losses are used that are related to perceptual audio quality, as well as a GAN framework to guide with a critic that discerns unrealistic waveforms. While yielding a high-quality audio, embodiments of the model can achieve more than 500 times faster than real-time audio synthesis. Multi-head convolutional neural network (MCNN) embodiments for waveform synthesis from spectrograms are also disclosed. MCNN embodiments enable significantly better utilization of modem multi -core processors than commonly-used iterative algorithms like Griffin-Lim and yield very fast (more than 300x realtime) waveform synthesis. Embodiments herein yield high-quality speech synthesis, without any iterative algorithms or autoregression in computations.
[0026] US Patent No. 11,854,554 for Method and apparatus for combined learning using feature enhancement based on deep neural network and modified loss function for speaker recognition robust to noisy environments by inventors Chang et al., filed March 30, 2020 and issued December 26, 2023, discloses a transformed loss function and feature enhancement based on a deep neural network for speaker recognition that is robust to a noisy environment. The combined learning method using the transformed loss function and the feature enhancement based on the deep neural network for speaker recognition that is robust to the noisy environment, according to an embodiment, may comprise: a preprocessing step for learning to receive, as an input, a speech signal and remove a noise or reverberation component by using at least one of a beamforming algorithm and a dereverberation algorithm using the deep neural network; aspeaker embedding step for learning to classify an utterer from the speech signal, from which a noise or reverberation component has been removed, by using a speaker embedding model based on the deep neural network; and a step for, after connecting a deep neural network model included in at least one of the beamforming algorithm and the dereverberation algorithm and the speaker embedding model, for speaker embedding, based on the deep neural network, performing combined learning by using a loss function.
[0027] US Patent No. 12,020,679 for Joint audio interference reduction and frequency band compensation for videoconferencing by inventors Xu et al., filed August 3, 2023 and issued June 25, 2024, discloses a device receiving an audio signal recorded in a physical environment and applying a machine learning model onto the audio signal to generate an enhanced audio signal. The machine learning model is configured to simultaneously remove interference and distortion from the audio signal and is trained via a training process. The training process includes generating a training dataset by generating a clean audio signal and generating a noisy distorted audio signal based on the clean audio signal that includes both an interference and a distortion. The training further includes constructing the machine learning model as a generative adversarial network (GAN) model that includes a generator model and multiple discriminator models, and training the machine learning model using the training dataset to minimize a loss function defined based on the clean audio signal and the noisy distorted audio signal.
[0028] US Patent Pub. No. 2023 / 0267950 for Audio signal generation model and training method using generative adversarial network by inventors Jang et al., filed January' 13, 2023 and published August 24, 2023, discloses a generative adversarial network-based audio signal generation model for generating a high qualify’ audio signal comprising: a generator generating an audio signal with an external input; a harmonic-percussive separation model separating the generated audio signal into a harmonic component signal and a percussive component signal; and at least one discriminator evaluating whether each of the harmonic component signal and the percussive component signal is real or fake.
[0029] US Patent No. 11,562,764 for Apparatus, method or computer program for generating a bandwidth-enhanced audio signal using a neural network processor by inventors Schmidt et al., filed April 17, 2020 and issued January' 24, 2023, discloses an apparatus for generating a bandw idth enhanced audio signal from an input audio signal having an input audio signal frequency range includes: a raw signal generator configured for generating a raw signal having an enhancement frequency range, wherein the enhancement frequency range is not included in the input audio signal frequency range; a neural netw ork processor configured for generating a parametric representation for the enhancement frequency range using the input audio frequency range of the input audio signal and a trained neural network; and a raw' signal processor forprocessing the raw signal using the parametric representation for the enhancement frequency range to obtain a processed raw signal having frequency components in the enhancement frequency range, wherein the processed raw signal and tire input audio signal frequency range of the input audio signal represent the bandwidth enhanced audio signal.
[0030] US Patent Pub. No. 2023 / 0245668 for Neural network-based audio packet loss restoration method and apparatus, and system by inventors Xiao et al,, filed September 30, 2020 and published August 3, 2023, discloses an audio packet loss repairing method, device and system based on a neural network. The method comprises: obtaining an audio data packet, the audio data packet comprises a plurality of audio data frames, and the plurality of audio data frames at least comprise a plurality of voice signal frames; determining a position of a lost voice signal frame in the plurality of audio data packets to obtain position information of the lost frame, the position comprising a first preset position or a second preset position; selecting, according to the position information of the lost frame, a neural network model for repairing the lost frame, the neural network model comprising a first repairing model and a second repairing model; and sending the plurality of audio data frames to the selected neural network model so as to repair the lost voice signal frame.
[0031] US Patent Pub. No. 2024 / 0095491 for Method and system for personalized multimodal response generation through virtual agents by inventors Birru et al., filed December 1, 2023 and published March 21, 2024, discloses a method and system for multimodal response generation through a virtual agent. Tire method comprises retrieving information related to an input received by the virtual agent. The virtual agent employs an Artificial Intelligence (Al) model. The method further comprises generating a response corresponding to tire input based on the retrieved information. The m ethod may further comprise generating a plurality of prompts based on user characteristics and the input. The method may further comprises modifying the response based on tire plurality of prompts to generate a multimodal response.
[0032] US Patent No. 12,099,802 for Integration of machine learning models with dialectical logic frameworks by inventor Petrauskas, filed April 9, 2024 and issued September 24, 2024, discloses techniques for executing dialectical analyses using large language models and / or types of deep learning models. A dialectic logic engine can store and execute various programmatic processes or functions associated with applying dialectic analyses to input strings. Tire programmatic processes or functions executed by the dialectic logic engine can initiate communication exchanges with one or more generative language models to derive parameters for performing dialectic analyses and / or to derive outputs based on the parameters. In some embodiments, the dialectic logic engine also can execute functions for enforcing constraint conditions and / or eliminating bias from responses generated by the one or more generativelanguage models to improve the accuracy, precision, and quality of the parameters and / or outputs derived from the parameters. Other embodiments are disclosed herein as well.
[0033] US Patent No. 12,039,263 for Systems and methods for orchestration of parallel generative artificial intelligence pipelines by inventors Mondlock et al., filed October 24, 2023 and issued July 16, 2024, discloses improvements to generative artificial intelligence systems through the use of generative artificial intelligence pipelines to supply external information to pre-trained large language models for use in answering queries. To improve the efficiency and accuracy of large language models in responding to user queries, according to various aspects described herein, such queries may be modified and augmented with additional relevant information and may be divided into multiple queries for parallel handling, the results of which may then be combined into a response. The additional relevant information may include portions of documents or other data sets to be used in generating the response. Additional aspects may further improve resilience and flexibility by managing the generation or implementation of such modified and augmented queries.
[0034] US Patent Pub. No. 2024 / 0211973 for Technology stack modeler engine for a platform signal modeler by inventors Sandbo et al,, filed December 22, 2023 and published June 27, 2024, discloses a platform signal modeler acquiring a first technology platform signal and generates, using the first technology platform signal, a synthetic signal using natural language processing. The synthetic signal includes a first token and a second token, the tokens relating to particular technology components. The modeler determines a taxonomy binding for the tokens based on a semantic distance between the tokens and generates a co-occurrence value for the taxonomy binding. The modeler augments the synthetic signal by acquiring a second technology platform signal and determining, based on the second signal, a momentum indicium that relates to the technology component.
[0035] The article “Bayesian inference to improve quality of Retrieval Augmented Generation” by author Rao, published August 2024, discloses Retrieval Augmented Generation or RAG being the most popular pattern for modem Large Language Model or LLM applications. RAG involves taking a user query and finding relevant paragraphs of context in a large corpus typically captured in a vector database. Once the first level of search happens over a vector database, the top n chunks of relevant text are included directly in the con text and sent as prompt to the LLM. A problem with this approach is that quality of text chunks depends on effectiveness of search. There is no strong post processing after search to determine if the chunk does hold enough information to include in prompt. Also many times there may be chunks that have conflicting information on the same subject and the model has no prior experience which chunk to prioritize to make a decision. Often times, this leads to the model providing a statement thatthere are conflicting statements, and it cannot produce an answer. In this research we propose a Bayesian approach to verify the quality of text drunks from the search results. Bayes theorem tries to relate conditional probabilities of the hypothesis with evidence and prior probabilities. We propose that, finding likelihood of text chunks to give a quality answer and using prior probability of quality of text chunks can help us improve overall quality of the responses from RAG systems. We can use the LLM itself to get a likelihood of relevance of a context paragraph. For prior probability of the text chunk, we use the page number in the documents parsed. An assumption is that that paragraphs in earlier pages have a better probability of being findings and more relevant to generalizing an answer,
[0036] Tire article “Superposition Prompting: Improving and Accelerating Retrieval-Augmented Generation” by authors Merth et al., published April 10, 2024, discloses that, despite the successes of large language models (LLMs), they exhibit significant drawbacks, particularly when processing long contexts. Their inference cost scales quadratically with respect to sequence length, making it expensive for deployment in some real-world text processing applications, such as retrieval -augmented generation (RAG). Additionally, LLMs also exhibit the “distraction phenomenon,” where irrelevant context in the prompt degrades output quality. To address these drawbacks, we propose a novel RAG prompting methodology, superposition prompting, which can be directly applied to pre-trained transformer-based LLMs without the need for finetuning. At a high level, superposition prompting allows the LLM to process input documents in parallel prompt paths, discarding paths once they are deemed irrelevant. We demonstrate the capability of the method to simultaneously enhance time efficiency across a variety of question answering benchmarks using multiple pre-trained LLMs. Furthermore, the technique significantly improves accuracy when the retrieved context is large relative the context the model was trained on. For example, the approach facilitates a 93* reduction in compute time while improving accuracy by 43% on the Natural Questions-Open dataset with the MPT-7B instruction-tuned model over naive RAG.
[0037] US Patent No. 8,238,679 for Lossless video data compressor with very high data rate by inventors Rudin et al., filed June 9, 2009 and issued August 7, 2012, discloses lossless video data compression performed in real time at the data rate of incoming real time video data in a process employing a minimum number of computational steps for each video pixel. A first step is to convert each pixel 8-bit byte to a difference byte representing the difference between the pixel and its immediate predecessor in a serialized stream of the pixel bytes. Thus, each 8-bit pixel byte is subtracted from its predecessor. This step reduces the dynamic range of the data. A next step is to discard any carry bits generated in the subtraction process of two's complement arithmetic. This reduces the data by a factor of two. Finally, tire 8-bit difference pixel bytes thusproduced are subject to a maximum entropy encoding process. Such a maximum entropy encoding process may be referred to as a minimum length encoding process. One example is Huffman encoding. In such an encoding process, a code table for the entire video frame is constructed, in which a set of minimum length symbols are correlated to the set of difference pixel bytes comprising the video frame, the more frequently occurring bytes being assigned to the shorter minimum length symbols.
[0038] US Patent No. 12,015,776 for Image compression and decoding, video compression and decoding: methods and systems by inventors Besenbruch et al., filed August 4, 2023 and issued June 18, 2024, discloses a computer-implemented method for lossy image or video compression, transmission and decoding, the method including the steps of: (i) receiving an input image at a first computer system; (ii) encoding the input image using a first trained neural network, using the first computer system, to produce a latent representation; (iii) quantizing the latent representation using the first computer system to produce a quantized latent; (iv) entropyencoding the quantized latent into a bitstream, using the first computer system; (v) transmitting the bitstream to a second computer system; (vi) the second computer system entropy decoding the bitstream to produce the quantized latent; (vii) the second computer system using a second trained neural network to produce an output image from the quantized latent, wherein the output image is an approximation of the input image. Related computer-implemented methods, systems, computer-implemented training methods and computer program products,
[0039] WIPO Patent Pub. No. 2024 / 080044 for Graphical user interface for generative adversarial network music synthesizer by inventors Narita et al., filed September 7, 2023 and published April 18, 2024, discloses an information processing system that receives input sound and pitch information; extracts a timbre feature amount from the input sound; and generates information of a musical instrument sound with a pitch based on the timbre feature amount and the pitch information.
[0040] The Article “MAMGAN: Multiscale attention metric GAN for monaural speech enhancement in the time domain” by authors Guo et al., published June 30, 2023 in Applied Acoustics Vol. 209, discloses "‘In the speech enhancement (SE) task, the mismatch between the objective function used to train the SE model, and the evaluation metric will lead to the low-quality of the generated speech. Although existing studies have attempted to use the metric discriminator to learn the alternative function of evaluation metric from data to guide generator updates, the metric discriminator's simple structure cannot better approximate tire function of the evaluation metric, thus limiting tire performance of SE. This paper proposes a multiscale attention metric generative adversarial network (MAMGAN) to resolve this problem. In the metric discriminator, the attention mechanism is introduced to emphasize the meaningful features1of spatial direction and channel direction to avoid the feature loss caused by direct average pooling to better approximate the calculation of the evaluation metric and further improve SE’s performance. In addition, driven by the effectiveness of the self-attention mechanism in capturing long-term dependence, we construct a multiscale attention module (MSAM). It fully considers the multiple representations of signals, which can better model the features of long sequences. The ablation experiment verifies the effectiveness of the attention metric discriminator and the MSAM. Quantitative analysis on the Voice Bank + DEMAND dataset shows that MAMGAN outperforms various time-domain SE methods with a 3.30 perceptual evaluation of speech quality score.”SUMMARY OF THE INVENTION
[0041] Hie present invention relates to neural networks, and more specifically to neural networks for compression and a continuously improving Al system.
[0042] It is an object of this invention to provide neural networks capable of being used across a wide variety of modalities and having improved interpretability.
[0043] In one embodiment, the present invention is directed to a neural netw ork, including an input layer, an output layer, and a neuromorphic field connecting the input layer to the output layer, wherein the neuromorphic field includes a plurality of interconnected artificial neurons w'hose behaviors are defined and governed by one or more partial differential equations (PDEs).
[0044] In another embodiment, the present invention is directed to a neural network, including an input layer, an output layer, an input-to-output mapping function, and a neuromorphic field connecting the input layer to the output layer, wherein the neuromorphic field is operable to form patterns in response to different inputs and / or input types, wherein the input-to-output mapping function includes at least one adjustable partial differential equation (PDE) hyperparameter, wherein the at least one adjustable PDE hyperparameter represents complexity of the input-to-output mapping function, time scale and / or pattern weight, wherein tire neuromorphic field is defined by one or more PDEs having one or more boundary conditions and wherein the one or more PDEs includes a reaction-diffusion PDE.
[0045] In another embodiment, the present invention is directed to a method for enhancing the accuracy and relevance of generated responses, including receiving a user query', performing graph-based retrieval from a know ledge graph, at least one large language model (LLM) generating a response to the user query, performing Bayesian evaluation of the response to tire user query and determining whether the response meets a predetermined quality threshold, performing secondary'- ground truth verification on the response, verifying the response with multiple artificial intelligence (Al) agents, adjusting the response based on feedback from the multiple Al agents and delivering the response to a user device.
[0046] In another embodiment, the present invention is directed to system for compressing media content, including a media analysis module configured to perform multi -faceted analysis on input media, a manifold selection and optimization module, a deep learning model training module, an entropy maximization module, a compression application module and an encoding module configured to package the compressed media into standard format containers.
[0047] In another embodiment, the present invention is directed to a method for compressing media content, including performing multi-faceted analysis on input media, including spectral analysis, statistical analysis, perceptual analysis, and temporal-spatial correlation analysis, selecting and optimizing a dimensional manifold based on results of the multi-faceted analysis, training a deep learning model to map between an original media space and the selected dimensional manifold, applying entropy maximization techniques to a representation of the selected dimensional manifold, compressing tire media content using the trained deep learning model and entropy-maximized manifold and encoding the compressed media content into a standard format container while maintaining compatibility with existing media ecosystems.
[0048] In yet another embodiment, tire present invention is directed to a method for compressing media content, including performing multi-faceted analysis on input media, including spectral analysis, statistical analysis, perceptual analysis, and temporal-spatial correlation analysis, selecting and optimizing a dimensional manifold based on results of the multi-faceted analysis, training a deep learning model to map between an original media space and the selected dimensional manifold, computing Shannon entropy for each dimension or feature in the representation of the selected dimensional manifold, applying Independent Component Analysis (ICA) to separate statistically independent components, implementing the Principle of Maximum Entropy to optimize distribution of information across the selected dimensional manifold, developing an adaptive quantization scheme that allocates more bits to high-entropy components and compressing the media content using the trained deep learning model and entropy-maximized manifold.
[0049] These and other aspects of the present invention will become apparent to those skilled in the art after a reading of the following description of the preferred embodiment when considered with the drawings, as they support the claimed invention,BRIEF DESCRIPTION OF THE DRAWINGS
[0001] FIG. 1 is a schematic diagram of an operating system (OS) connected to a plurality of large language models (LLMs) according to one embodiment of the present invention.
[0002] FIG. 2 is a schematic diagram of core architecture of an OS configured to integrate with LLMs according to one embodiment of the present invention.
[0003] FIG. 3 is a schematic diagram for a Bayesian network-based retrieval-augmentation generation system according to one embodiment of the present invention.
[0004] FIG. 4 is a schematic diagram for a process of query' processing and response generation according to one embodiment of the present invention.
[0005] FIG. 5 is a schematic diagram of a Bayesian evaluation network according to one embodiment of the present invention.
[0006] FIG. 6 is a schematic diagram of a synthetic data feedback loop according to one embodiment of the present invention.
[0007] FIG. 7 is a schematic diagram of a multi -agent verification system according to one embodiment of the present invention.
[0008] FIG. 8 is a schematic diagram of an OS architecture integrating multiple LLMs according to one embodiment of the present invention.
[0009] FIG. 9 is a schematic diagram of intelligent virtual input / output (I / O) controls according to one embodiment of the present invention.
[0010] FIG. 10 is a schematic diagram of OS-level intelligent automation according to one embodiment of the present invention.
[0011] FIG. 11 illustrates a schematic diagram for a system for enhancing an input audio source according to one embodiment of the present invention.
[0012] FIG. 12 is a flow diagram for an audio sourcing stage of an audio restoration or enhancement process according to one embodiment of the present invention.
[0013] FIG. 13 is a flow diagram for a preprocessing and FX processing stage of an audio restoration or enhancement process according to one embodiment of the present invention.
[0014] FIG. 14 is a flow diagram for an artificial intelligence (Al) training and deep learning model stage of an audio restoration or enhancement process according to one embodiment of the present invention.
[0015] FIG. 1 is a flow diagram for an indexing and transform optimization stage of an audio restoration or enhancement process according to one embodiment of the present invention.
[0016] FIG. 16 is a flow diagram for a dimensional and environmental transform stage of an audio restoration or enhancement process according to one embodiment of the present invention.
[0017] FIG. 17 is a flow diagram for an output mode stage of an audio restoration or enhancement process according to one embodiment of the present invention.
[0018] FIG. 18 is a signal diagram for a signal formatted with Pulse Code Modulation without application of improvements provided by the present invention.
[0019] FIG. 19 is a signal diagram for a Pulse Density Modulated Signal according to one embodiment of the present invention.
[0020] FIG. 20 is a graph showing aliasing in the use of low-pass or high-pass filters for audio data.
[0021] FIG. 21 is a graph showing reflective in-band aliasing for high-pass and low-pass audio filters.
[0022] FIG. 22 is a graph showing limits for noise-constrained conversion.
[0023] FIG. 23 is a schematic diagram of a pulse density modulation process for capturing and encoding lossless audio according to one embodiment of the present invention.
[0024] FIG. 24 is a schematic diagram of an audio capture stage of a pulse density modulation process according to one embodiment of the present invention,
[0025] FIG. 25 is a schematic diagram of an analog equalization stage of a pulse density modulation process according to one embodiment of the present invention.
[0026] FIG. 26 is a schematic diagram of an analog dynamics restoration stage of a pulse density modulation process according to one embodiment of the present invention.
[0027] FIG. 27 is a schematic diagram of a target environment equalization stage of a pulse density modulation process according to one embodiment of the present invention.
[0028] FIG. 28 is a schematic diagram of a differentially modulated pulse shaping stage of a pulse density modulation process according to one embodiment of the present invention.
[0029] FIG. 29 is a schematic diagram of an analog -to-digital conversion stage of a pulse density modulation process according to one embodiment of the present invention.
[0030] FIG. 30 is a schematic diagram of a jitter correction and frame alignment stage of a pulse density modulation process according to one embodiment of the present invention.
[0031] FIG. 31 is a schematic diagram of a DSD to PCM conversion stage of a pulse density modulation process according to one embodiment of the present invention,
[0032] FIG. 32 is a schematic diagram of a digital effects processing stage of a pulse density modulation process according to one embodiment of the present invention.
[0033] FIG. 33 is a schematic diagram of a file size and bandwidth reduction processing stage of a pulse density modulation process according to one embodiment of the present invention.
[0034] FIG. 34 is a schematic diagram of a digital-to-analog conversion stage of a pulse density modulation process according to one embodiment of the present invention.
[0035] FIG. 35 is a schematic diagram of a storage function support stage of a pulse density modulation process according to one embodiment of the present invention.
[0036] FIG. 36 illustrates a schematic diagram for a method of compression according to one embodiment of the present invention.
[0037] FIG. 37 illustrates a representation of a mapping function for a neural network, including training according to one embodiment of the present invention.
[0038] FIG. 38 illustrates a schematic diagram of a system of the present invention.
[0039] FIG. 39 illustrates a schematic diagram of a system according to one embodiment of the present invention.
[0040] FIG. 40 illustrates a schematic diagram of a system utilizing both linear and nonlinear PDEs according to one embodiment of the present invention.
[0041] FIG. 41 illustrates a schematic diagram of a system for training a neural network according to one embodiment of the present invention,
[0042] FIG. 42 is a schematic diagram of a system of the present invention.DETAILED DESCRIPTION
[0043] The present invention is generally directed to neural networks, and more specifically to neuromorphic field-defined neural networks.
[0044] In one embodiment, the present invention is directed to a neural network, including an input layer, an output layer, and a neuromorphic field connecting the input layer to tire output layer, wherein the neuromorphic field includes a plurality of interconnected artificial neurons whose behaviors are defined and governed by one or more partial differential equations (PDEs).
[0045] In another embodiment, the present invention is directed to a neural network, including an input layer, an output layer, an input-to-output mapping function, and a neuromorphic field connecting the input layer to the output layer, wherein the neuromorphic field is operable to form patterns in response to different inputs and / or input types, wherein the input-to-output mapping function includes at least one adjustable partial differential equation (PDE) hyperparameter, wherein the at least one adjustable PDE hyperparameter represents complexity of the input-to-output mapping function, time scale and / or pattern weight, wherein tlie neuromorphic field is defined by one or more PDEs having one or more boundary conditions and wherein the one or more PDEs includes a reaction-diffusion PDE.
[0046] In another embodiment, the present invention is directed to a method for enhancing the accuracy and relevance of generated responses, including receiving a user query, performing graph-based retrieval from a knowledge graph, at least one large language model (LLM) generating a response to the user query, performing Bayesian evaluation of the response to the user query and determining whether the response meets a predetermined quality threshold, performing secondary ground truth verification on the response, verifying the response with multiple artificial intelligence (Al) agents, adjusting the response based on feedback from the multiple Al agents and delivering the response to a user device.
[0047] In another embodiment, the present invention is directed to system for compressing media content, including a media analysis module configured to perform multi -faceted analysis on input media, a manifold selection and optimization module, a deep learning model training module, an entropy maximization module, a compression application module and an encoding module configured to package the compressed media into standard format containers.
[0048] In another embodiment, the present invention is directed to a method for compressing media content, including performing multi-faceted analysis on input media, including spectral analysis, statistical analysis, perceptual analysis, and temporal-spatial correlation analysis, selecting and optimizing a dimensional manifold based on results of the multi-faceted analysis, training a deep learning model to map between an original media space and the selected dimensional manifold, applying entropy maximization techniques to a representation of the selected dimensional manifold, compressing tire media content using the trained deep learning model and entropy-maximized manifold and encoding the compressed media content into a standard format container while maintaining compatibility with existing media ecosystems.
[0049] In yet another embodiment, tire present invention is directed to a method for compressing media content, including performing multi-faceted analysis on input media, including spectral analysis, statistical analysis, perceptual analysis, and temporal-spatial correlation analysis, selecting and optimizing a dimensional manifold based on results of the multi-faceted analysis, training a deep learning model to map between an original media space and the selected dimensional manifold, computing Shannon entropy for each dimension or feature in the representation of the selected dimensional manifold, applying Independent Component Analysis (ICA) to separate statistically independent components, implementing the Principle of Maximum Entropy to optimize distribution of information across the selected dimensional manifold, developing an adaptive quantization scheme that allocates more bits to high-entropy components and compressing the media content using the trained deep learning model and entropy-maximized manifold.
[0050] There are numerous components and processing methods widely used in the recording and playback chain of audio that collectively affect the perceived quality and other characteristics of the sound. Every type of digital recording is based on numerous assumptions, derived from a combination of engineering approximations, trial and error methods, technological constraints and limitations, prior beliefs and available knowledge at a given time that define the extents of the ability of audio engineers to support the recording, processing, distribution, and playback of audio.
[0051] Because the collection of knowledge together with beliefs and assumptions are taught as the basis for audio engineering and related theory, these beliefs and assumptions generallydefine the accuracy and extent of the capabilities of the industry. As a result, this collective base of understanding has historically limited the ability to engineer hardware and software solutions related to audio. In its most fundamental terms, the limitations of the accuracy and extent of the collective knowledge and understanding related to audio and the processes described have always constrained the ability of the prior art to define more optimal algorithms, methods, and associated processes using traditional, non-AI-based software and related engineering methods.
[0052] The advent of artificial intelligence coupled with the evolution of digital and analog technologies available to record, transform and play audio are allowing engineers to bypass limited and otherwise imperfect knowledge and poorly supported assumptions that limit audio fidelity and processing capabilities, in favor of an Al-enabled approach built upon ground truth data supporting a foundation model derived using a combination of source disparity recognition and related methods. As evidenced over the past several years across numerous medical, gaming, and other fields, the ability of key Al architectures to derive new capabilities has resulted in entirely new levels and types of capabilities beyond what was possible via traditional human and pre-AI computing methods.
[0053] Tire process of engineering and development using Al is very’ different from traditional, non-AI software development on a fundamental level, which enables the creation of previously impossible solutions. Using Al based development, the effective algorithms and related processes become the output created by the Al itself. When ground troth data is provided as part of the training process, it enables the neural network to become representative of a “foundation model.” For the purposes of this application, ground truth data refers to reference data, which preferably includes, for the purposes of the present invention, audio information at or beyond the average human physical and perceptual limits of hearing, and a foundational model refers to a resulting Al-enabled audio algorithm that takes as input the ground truth data to perform a range of extension, enhancement and restoration of the audio, yielding a level of presence, tonal quality, dynamics and / or resulting realism that is beyond the input source quality, even where the input includes original master tapes.
[0054] As a result, the use of Al-based systems, and more specifically a level of processing power and capabilities that support the approach described herein, allows for the avoidance of traditional assumptions and beliefs in audio processing, and the resulting implicit and explicit limits of understanding associated with those assumptions and beliefs. Instead, a benchmarked standard is used based on the disparities inherent to any type of recorded music relative to reference standards by using the approach described herein.
[0055] Artificial Intelligence (Al) systems, particularly Large Language Models (LLMs), have made significant strides in generating human-like text and automating complex tasks.However, challenges remain in optimizing these systems for accuracy, efficiency, and adaptability. Traditional Retrieval -Augmented Generation (RAG) methods often rely on document-based retrieval systems, which may not capture complex interconnections between pieces of information. Additionally, managing multiple LLMs and integrating them effectively poses a significant challenge due to differences in query modalities, syntaxes, and optimal utility ranges.
[0056] The present invention introduces a method and system called Bayesian Graph-Based Retrieval-Augmented Generation with Synthetic Feedback Loop (BaG-RAG-SFL), also referred to herein as "ZON OS" for other purposes. Tills method combines Bayesian evaluation, graphbased RAG, and a synthetic data feedback loop to create a continuously improving Al system. It intelligently integrates multiple LLMs, optimizing their individual and collective performance while insulating users from the complexities and rapid advancements in Al technologies.
[0057] Referring now to the drawings in general, the illustrations are for the purpose of describing one or more preferred embodiments of the invention and are not intended to limit the invention thereto.
[0058] FIG. l is a schematic diagram of an operating system (OS) connected to a plurality of large language models (LLMs) according to one embodiment of the present invention. In one embodiment, key components of the present invention include graph-based retrievalaugmentation generation (RAG), which utilizes a knowledge graph for more complex and interconnected information retrieval. This graph-based approach surpasses traditional documentbased RAG systems by capturing relationships between data points, enabling more accurate and contextually relevant information retrieval.
[0059] In one embodiment, the key components further include Bayesian evaluation, which involves utilizing a Bayesian network to evaluate the quality and the relevance of retrieved information and generated responses from various models. Bayesian evaluation considers multiple factors such as coherence, factual accuracy, and relevance to the query. This probabilistic model allows for dynamic weighting and assessment of different aspects of the response.
[0060] In one embodiment, the key components include a secondary’ ground-truth graph, providing an additional layer of verification through a curated knowledge graph. The secondary ground-truth graph provides trusted information against which the system's outputs are checked, enhancing reliability and reducing errors.
[0061] In one embodiment, the key components include a synthetic data feedback loop, which involves generating synthetic data to further train and improve the model. In one embodiment, synthetic data is generated based on current knowledge and performance metrics.This data includes new query-response pairs, additions or modifications to the knowledge graph, and updates to Bayesian network parameters. The synthetic data is used to retrain and improve the LLMs, the graph-based RAG system, and the Bayesian evaluation network.
[0062] In one embodiment, the key components include multi-agent verification, which involves collaborating specialized artificial intelligence (Al) agents to verify and enhance outputs. Multiple specialized Al agents collaborate to verify and improve the system's outputs. These agents include fact-checkers, coherence analyzers, and relevance assessors. The agents provide feedback that informs the synthetic data generation and overall system refinement.
[0063] Tn one embodiment, the key components include intelligent virtual input / output controls, which manage input and output across various modalities and systems. The system supports multi-modal input and output, including text, audio, images, and video. It manages input and output across various systems, analogous to the frontal lobe and prefrontal cortex of tire human brain, dispatching, sequencing, and directing data as needed.
[0064] A process flow according to one embodiment of the present invention is as follows: 1. Query Processing, wherein the system receives a query and utilizes the graph-based RAG-system to retrieve relevant information from the knowledge graph; 2, Response Generation, wherein an LLM generates a response based on the retrieved information and the query; 3. Bayesian Evaluation, which evaluates the generated response for coherence, relevance, and factual accuracy; 4. Secondary Ground-Truth Verification, wherein the response is cross-checked against the secondary ground-truth graph for additional verification; 5. Multi -Agent Verification, which invol ves specialized Al agents reviewing the response and providing feedback; 6. Synthetic Data Generation, wherein synthetic data is generated to improve the system based on evaluations and feedback; and 7. Feedback Loop, wherein the synthetic data retrains components, leading to continuous improvement.
[0065] A preferred embodiment of tire present invention functions as an operating system (OS) control and automation support system, acting as a virtual user with screen input-output (I / O) control and managing multiple computers as an intelligent process automation system. The ZON OS serves as an intermediary between third-party LLMs and desired functionalities, effectively managing, dispatching, and optimizing queries and results across multiple LLMs. Key aspects of the ZON OS include: 1. Dispatcher Engine, managing optimal LLM queries based on prior training and ongoing analysis; 2. Citation Research and Analysis System, comparing, contrasting, and weighing results from multiple LLMs using Bayesian Monte Carlo analysis; 3. Asynchronous Dispatch and Results Capture System, calling external systems -with available Application Programming Interfaces (APIs) or data streams for integration; and 4. Input andOutput Capabilities, supporting rapid user interface (UI) development, multi-language I / O, and audio visual I / O, and securing integration of local documents, videos, and databases.
[0066] FIG. 2 is a schematic diagram of core architecture of an OS configured to integrate with LLMs according to one embodiment of the present invention. The present invention provides rapid LLM extensibility, privacy control over queries and data, addition of new LLMs and models, integration with external software, security enhancement, and adaptability to new Al technologies.
[0067] In a preferred embodiment, the present invention functions as an operating system control and automation support system, capable of acting as a virtual user with screen I / O control and managing multiple computers. Key functionality of this system includes: I. Natural language interface, which allows users to interact with the system through text or voice commands; 2. Sy stem Management Capabilities, which handle process management, memory allocation, file system operations, device-drivers, user accounts, security and network access; 3. Application Integration, with integration and management of various OS functionalities, utilities, and applications; 4. Screen I / O and Virtual User Capabilities, which interpret screen output and provide input / output monitoring and control using computer vision, input simulation, and audio processing; 5. Intelligent Oversight and Process Control, monitoring system resources, application performance, and user behavior to optimize processes and workflow; 6. File Management and I / O, which manages file operations, interprets natural language requests, handles file transfers, and implements smart backup systems; 7. Intelligent Task Execution, which breaks down complex tasks into sequences of actions, executing and adapting plans as needed; 8. Considerations addressing performance and resource management, security and privacy, reliability and error handling, compatibility, and user adaptation; and 9. The benefits of increased productivity, enhanced accessibility, adaptive computing, and intelligent resource management.
[0068] FIG. 3 is a schematic diagram for a Bayesian network -based retrieval -augmentation generation system according to one embodiment of the present invention. The arrow's utilized in FIG. 3 indicate the flow' of data between components. In one embodiment, components of the Bayesian Graph-based retrieval-augmentation generation w'ith synthetic feedback loop (BG-RAG-SFL) include a user interface module, a graph-based RAG system, an LLM response generator, a Bayesian evaluation netw ork, a secondary ground-truth graph, a multi-agent verification system, and a synthetic data generator.
[0069] The user interface module receives inputs for user queries. The graph-based RAG system is connected directly to the user interface module. The LLM response generator is connected directly to the graph-based RAG system, lire Bayesian evaluation network receivesdata flow from the LLM and is positioned below the LLM response generator. The secondary ground-truth graph is connected to the Bayesian evaluation network with data flowing back and forth between the two. The multi-agent verification system receives data input from the Bayesian evaluation network. The synthetic data generator receives input from the multi-agent verification system and is connected to the knowledge graph, LLM, and Bayesian netw ork to establish a feedback loop.
[0070] FIG. 4 is a schematic diagram for a process of query processing and response generation according to one embodiment of the present invention. In a first step, a user query is received and provided to a graph-based retrieval from a knowledge graph. An LLM then generates a response and Bayesian evaluation is performed forthat response. If the response meets a quality threshold, then the system delivers the response to the user, while, if it does not, secondary ground-truth verification is performed. This verification is able to include verification by multiple Al agents. The response is then adjusted based on feedback and evaluated again until it meets the quality threshold.
[0071] FIG. 5 is a schematic diagram of a Bayesian evaluation network according to one embodiment of the present invention. In one embodiment, generated LLM responses are evaluated for coherence, relevance, and factual accuracy, which impact an overall quality score.
[0072] FIG. 6 is a schematic diagram of a synthetic data feedback loop according to one embodiment of the present invention. Inputs from a Bayesian evaluation network and a multiagent feedback loop are connected to a synthetic data generator. The synthetic data generator generates new query response pairs, updates a knowledge graph, and adjusts Bayesian network parameters, which are used to inform the LLM training module and actively modify the performance of the Bayesian evaluation network.
[0073] FIG. 7 is a schematic diagram of a multi -agent verification system according to one embodiment of the present invention. Results from a fact-checker evaluating factual accuracy, a coherence analyzer agent ensuring logical consistency, and a relevance assessor agent evaluating content relevance are fed to the multi-agent coordinator to generate an evaluation, which is then able to be fed to the synthetic data generator. In one embodiment, the multi-agent coordinator is also connected to an ethical compliance agent for ensuring compliance with ethical standards.
[0074] FIG. 8 is a schematic diagram of an OS architecture integrating multiple LLMs according to one embodiment of the present invention. A ZON OS dispatcher engine is connected to multiple LLMs with bidirectional data flow. The ZON OS is also connected to a citation research module, an asynchronous dispatch system, and a results analysis engine. A user interface is connected to the dispatcher engine as well. In one embodiment, the system includes data repositories such as local knowledge bases, external APIs, and / or databases.
[0075] FIG. 9 is a schematic diagram of intelligent virtual input / output (I / O) controls according to one embodiment of the present invention, A Natural Language Interface is able to interpret user intent and plan a sequence of actions. The system determines if screen I / O is required and, if so, activates a computer vision module. If not, the system proceeds with executing commands. Screen I / O functions include screen reading, input simulation (e.g., mouse control, keyboard input, etc,), and audio processing (e.g., speech recognition, text-to-speech, etc.). The system then executes commands to perform system operations. In one embodiment, the system includes a feedback loop to monitor results and adapt as needed.
[0076] FIG. 10 is a schematic diagram of OS-level intelligent automation according to one embodiment of the present invention. In one embodiment, core components of the system include an Al operating system kernel, a system management module, and an application integration layer. Functional modules of the system include process management, memory allocation, filing system operations, device drivers, security and access control, and / or a network control module. In one embodiment, the natural language interface is connected to the Al OS kernel. The network control module is configured for communication with external devices, remote computers, and / or cloud services. In one embodiment, the Al operating system kernel includes a repository of intelligent automation scripts.
[0077] Hie present invention includes a modular, software-driven system and associated hardware-enabled methodology for improving and / or otherwise enhancing the sound quality and associated characteristics of audio to a level of acoustic realism, perceptual quality and sense of depth, tonality, dynamics and presence beyond the limits of prior art sy stems and methods, even exceeding the original master tapes.
[0078] Tire system of the present invention employs a combination of deep learning models and machine learning methods, together with a unique process for acquiring, ingesting, indexing, and applying media-related transforms. The invention enables the use of resulting output data to train a deep learning neural network and direct a modular workflow to selectively modify the audio via a novel inference-based recovery, transformation, and restoration chain. These deep learning algorithms further allow the system to enhance, adapt, and / or recover audio quality lost during the acquisition, recording, or playback processes, due to a combination of hardware limitations, artifacts, compression, and / or other sources of loss, change, and degradation.Furthermore, the system of the present invention employs a deep neural network to analyze differences between an original audio source or recording and a degraded or changed audio signal or file and, based on knowledge obtained via the training process, distinguish categories and specific types of differences from specific reference standards. This enables a novel application of both new and existing methods to be used to recover and bring the quality andnature of the audio to a level of acoustic realism, perceptual quality, sense of depth, tonality, dynamics, and presence beyond any existing method, even including original master tapes.
[0079] Tire system and method for improving and enhancing audio quality of analog and digital audio as described herein provides for an improvement in the ability to listen to and enjoy music and other audio. By utilizing the deep learning algorithms of the present invention, as well as the advanced recovery and transformation workflow, the system is able to effectively restore lost audio quality in both live and recorded audio, and in both digital and analog audio, to bring audiences closer to a non -diminished audio experience.
[0080] Tire present invention covers various uses of generative deep-learning algorithms that employ indexing, analysis, transforms, and segmentation to derive aground truth-based foundation model to recover the differences between the highest possible representative qualityaudio recorded, both analog and digitally recorded, including comparisons with bandwidth-constrained, noise-diminished, dynamic range limited, and noise-shaped files of various formats (e.g., MP3, AAC, WAV, FLAC, etc.) and of various encoding types, delivery methods, and sample rates.
[0081] Because of the modular design of the system and the directive workflow and output of the artificial intelligence module, a wide range of hardware, software and related options are able to be introduced at different stages, as explained below, supporting a virtually unlimited range of creative, restoration, transfer, and related purposes. Unlike other methods of audio modification or restoration, the system of the present invention leverages approaches that were formerly not cost or time viable prior to the current level of processing power and scalability enabled by the use of Al-based sy stems. One of ordinary' skill in the art will understand that the present invention is not intended to be limited to any particular analog or digital format, sampling rate, bandw idth, data rate, encoding type, bit depth, or variation of audio, and that variations of each parameter are able to be accepted according to the present invention.
[0082] The system is able to operate independently of the format of the input audio and the particular use case, meaning it supports applications including, but not limited to, delivery and / or playback using various means (e.g., headphones, mono playback, stereo playback, live event delivery, multi-channel delivery-, dimensionally enhanced, and extended channel formats delivered via car stereos, as well as other types, uses and environments). While a primary- use case of the present invention is for enhancing music, the system is able to be extended to optimization of other forms of audio as well, via the sequence of stages and available options as described herein. To support the extensibility to various forms of audio, the system provides for media workflow- and control options, and associated interfaces (e.g., Application Programming Interfaces (APIs)).
[0083] Furthermore, the system of the present invention also includes software-enabled methodology that leverages uniquely integrated audio hardware and related digital systems to capture and encode full spectrum lossless audio, as defined by physical and perception limits of human audiology. This approach uses a uniquely integrated Al-assisted methodology as described to bypass several long-standing limits based on beliefs and assumptions related to the frequency range, transients, phase, and related limits of human hearing, in favor of results obtained via leading-edge research in sound, neurology, perception, and related fields.
[0084] Ihe system is able to be used in isolation or in combination with other audio streaming, delivery, effects, recording, encoding or other approaches, whether identified herein or otherwise. Al is employed to support brain-computer-interface (BCI) and related brain activity monitoring and analytics, to determine physically derived perceptual human hearing limits in terms of transient, phase, frequency, harmonic content, and related factors. Current “lossless” audio formats and methods are missing over 90% of the frequency range, as well as much of the transient detail and phase accuracy necessary to be lossless, as defined by no audible signals within the limits of human hearing have been discarded, compressed, or bypassed.
[0085] Prior limits of human hearing were defined to be, at best, between 20 cycles (Hz) and 20,000 Hz using a basic pass / fail sine wave hearing test. While this is useful in a gross sense for human hearing of only sine waves, those approaches disregard the reality that virtually all sound in the real world is composed of a wide range of complex harmonic, timbral, transient, and other details. Further, virtually all hearing related tests ignore a wide range of other methods of testing and validation, including using brain pattern-based signal perception testing to ensure parity with human brain and related hearing function.
[0086] Numerous studies have begun to verify the fact that hearing extends across a much wider range of frequencies and has a much more extensive set of perceptually relevant biophysical affects. To determine the actual frequency range of human hearing, studies have been done to take such details into account, finding that human hearing extends much further when integrating those noted acoustic factors. In reality, an extended range of frequencies that are actually able to be perceived extends from 2 Hz to about 70,000 Hz. Between approximately, 2 Hz and 350 Hz, the primary part of the body able to perceive the sound is the skin or chest of a listener, while the ear is able to perceive qualities such as frequency, timbre, and intonation for sounds between approximately 350 Hz and 16,000 Hz. Between about 16,000 Hz and 70,000 Hz, the inner ear is predominant in the perception of the sound.
[0087] In addition, there are numerous other physiological and related considerations in determining how to optimally record, encode, and define analog and digital sounds. For example, the idea that humans are able to hear low frequency information solely as a frequency via ourears alone is erroneous, given the size of the tympanic membrane, which is incapable of sympathetic oscillation at frequencies much below 225 Hz. Instead, transient, harmonic, and other associated detail that provides the critical sonic information enables the ear, brain, and body to decode much of the sound enabling us to, for example, differentiate a tom-tom from a finger tap on some other surface. Further, the body acts to help us perceive audio down to a few cycles per second. As such, differential Al driven brain activity analytics are commonly employed as part of the testing to ensure definition of the actual physiological and perceptual hearing limits using complex, real world audio signals across transient, harmonics, tinibral, and other detail, rather than using common frequency based and other audiology and related testing.
[0088] Similarly, as some studies have moved away from simple sine wave data used in testing hearing sensitivity and limits, to audio test sources with a range of transient, harmonic, phase and timbrel complexity, those studies have begun to see that hearing and perception are a whole brain plus body experience, meaning that engineering and related methods need to take these factors into account in order to be reflective of real world human hearing.
[0089] Numerous other capabilities include the ability to dramatically improve tire perceived quality of the sound even when compressed. This is due to the fact that the system starts with a significantly higher resolution sonic landscape that more accurately reflects the limits of human hearing, rather than an already diminished and compromised one that does not include much of the sonic image to begin with. Among other things, this results in increased perceived quality with significantly reduced audio file sizes, along with commensurately reduced resulting bandwidths and storage requirements.
[0090] Hie unique inventive method of tire present invention employs a hybrid digital-analog format that uses a type of Pulse Density Modulation (PDM) and process that interoperates with traditional Pulse Code Modulation (PCM) systems. Because of the unique implementation described in the preferred embodiment, the system is able to bypass the requirement of a digital to analog conversion stage (DAC) and the associated jitter, sampling distortion, quantization noise, intermodulation distortion, phase distortion, and various nonlinear issues created by a DAC stage, it is able to bypass these issues because the system enables the ability to output the digital equivalent of an analog signal, hence the labeling of a hybrid digital -analog format,
[0091] FIG. 11 illustrates a schematic diagram for a system for enhancing an input audio source according to one embodiment of the present invention. As shown in FIG. 11, there are six main stages of the system of the present invention, defined as stages 1 through 6. Beginning with the input source, and proceeding to the final output, the following stages define the steps of a system able to be used for various types of input source and for various types of audio processing. The system is able to be used for two general purposes: (1) training and transferlearning, and (2) inferenced transcoding and processing operations. In both modes, the same stages are implemented, although the specific contents of each stage shift, as described below.
[0092] The first step of the system is sourcing. While any analog or digital audio source type, format, sample rate, or bit-depth is able to be used for training or inference, having a range of relative quality levels facilitates the training by enabling the deep learning neural network to derive differential, associative, and other patterns inherent to each source ty pe (and the related quality levels of each source type) to establish pattern recognition and discrimination within a neural network model. The sourcing for tire system is able to derive from any analog or digital source, whether locally provided as a digital file, introduced as an analog source, streamed or otherwise, and tire subsequent stages of the system process the sourced data further. Examples of comparative sources able to be acquired and organized according to the present invention are shown as items 1A-1E in FIG. 11, but one of ordinary' skill in the art will understand that other types of audio data are also able to be used in place of those shown.
[0093] In one embodiment, a Direct Stream Digital source, or other high quality audio source, as indicated in 1 A, is used as a reference, or ground truth source. Within the set of data constituting the ground truth source, the system is able to include a unique set of audio sourcing across many types of audio examples, including, but not limited to, a range of: musical instrument types, musical styles, ambient / room interactions, compression levels, dynamics, timbres, overtones, harmonics, and / or other types of audio source types. The use of ground truth source data allows for the system to train on audio data, including content below the noise floor that is eliminated in prior art systems. In one embodiment, the audio data includes information at least 1 dB below the noise floor. In one embodiment, the audio data includes information at least 5 dB below the noise floor. In one embodiment, the audio data includes information at least 10 dB below' the noise floor. In one embodiment, the audio data includes information at least 25 dB below' the noise floor. In one embodiment, the audio data includes information at least 50 dB below' the noise floor. In one embodiment, the audio data includes information at least 70 dB below the noise floor.
[0094] In one preferred embodiment, the system is able to use, for example, sources including one or more pristine sources of analog and digital recordings, such as high sample rate DSD audio recording acquired using high quality gear. Together with any limits, improvements, and / or enhancements made in the acquisition and recording process as described herein, the source in 1A effectively provides an exemplary' upper quality' limit for the available source examples. As described herein, there are a range of improvements made via artificial intelligence module and other related subcomponents of the system, which further elevate the capabilities of this source type and thereby the invention. Idris improvement is possible due to Al-identifiedinherent noise and related patterns even in high quality audio data, together with human perceptual adaptations where such perception is possible,
[0095] In one embodiment, source IB includes the most widely available high quality' source type used in recording, and is often referred to as lossless, “CD quality” or by the specific sampling rate and bit depth commonly used, such as 44.1 kHz at 16 bits, 96 kHz at 24 bits or another similar specification. As with 1A described above, source IB is able to be supplied across a wide range of source types.
[0096] Fidelity relative to the original source material is diminished proceeding from 1C (e.g., MP3 at 320 kbps), to ID (e.g,, Apple AAC at 128 kbps), and ultimately to IE (e.g,, MP3 at 64 kbps). While the source types depicted in FIG. 11 represent one preferred embodiment of the present invention, the actual specific levels of quality and source types are not limited to these, just as the specific audio source types and the ways in which they differ are also able to vary. These variations and options are useful in training, transfer learning and other adaptations for different purposes.
[0097] Hie next stage, stage 2, is a pre-processing and FX processing stage, including data aggregation and augmentation. In the training mode of operation, stage 2 is where the system employs standard types of audio processing variations, including compression, equalization, and / or other processes, and then turns the pre-processed and post-processed examples of these transformed audio files into digital spectrographs able to be used for Al training in stage 3. Stage 2 provides for representative training examples in the formats most efficiently usable for Al training.
[0098] Hie Al module utilized and trained in stage 3 is able to include supervised learning models (3A), unsupervised learning models (3B), and / or semi-supervised learning models (3C), In one embodiment, the machine learning module utilizes grokking to understand the source data provided and to therefore train the model, hi this stage, the system both trains the deep learning models, and secondarily derives the abilities to: (a) segment audio data by character, quality / fidelity, genre, compression rate, styles and other attributes, enabling derivation of further training, workflow' processing and output control options; and (b) create workflow and related indexes to be used for determining setings and step related variations to be used in the later stages for various restoration and / or transforms and / or effects as described herein. One of ordinary skill in the art will understand that the term “segmentation” in this case is not limited to its use in the prior art as meaning dividing the audio data into particular tracks or segments, but includes grouping multiple sources of audio data by particular trai ts or qualities, including those mentioned above. Further, this is able to be used to extend an API for other operational anddeployment purposes, such as a Platform as a Sendee, to enter a transfer learning mode (e.g., for other markets and industries), and / or other uses,
[0099] Tire system is also able to provide Al-enabled direct processing transforms to be used to enhance, extend, or otherwise modify audio files directly for an intended purpose. This is based on applying the associative and differential factors across the audio data types to a range of transform options as described herein. Providing enough examples of the right types to enable the Al in stage 3 to derive weights, biases, activation settings, and associated transforms for the deep learning model to be used in stage 4 is essential.
[0100] Stage 4 is the indexing and transform optimization stage, enabling a user to selectively employ the information and capabilities derived from the earlier stages to set and employ the necessary transforms. Standard interface dashboards and related controls enable user selective choices, which are able to be intelligently automated via a scripting API. Specifically, the API is able to receive user selection to leverage a prior input and processing for remastering an audio file more optimally for a particular format or modality (e.g., DOLBY ATMOS), or recover information lost as a result of, for example, compression or other factors. In summary, this stage provides for specific deployment that affects how the audio is transformed, and thus the form and qualities of the final output.
[0101] In step 4A, the system is able to employ Al-derived segmentation analysis, deriving and determining which subsequent transform options and settings best suit the input audio given its state, style, or other characteristics, and given the desired deliver}' modality (e.g., live, studio, headphones, etc.). In step 4B, the artificial intelligence module of the system is able to choose whether to apply differential processing to the audio to achieve a restoration, improvement, modification, or creative function, beyond any mastering grade restoration. Transforms able to be applied by the artificial intelligence module in the system of the present invention include, but are not limited to, recovering information of reduced quality audio, removing unwanted acoustics or noise, and / or other functions. In step 4C, automated Al inferencing is able to be employed to automatically achieve a selected objective, based on inherent limits of the source material in comparison to patterned references inherent in the trained neural network. In step 4C, due to the inherent support for transfer learning, the system is also able to use differential examples to creatively direct style, level, equalization, or other transformations.
[0102] In stage 5, the system selectively employs one or more transforms (e.g., analog (5 A), digital (5B) or dimensional (5C) transforms) for the audio, based on the creative or specific usage or other objectives. In one embodiment, it is at this stage where the system is able to employ transforms suitable for specific formats (e.g., DOLBY ATMOS) or have a glue-pass (i.e., a purpose-directed compression step) execute a compression or other function. Stage 5 providesthe necessary controls to apply temporal-spatial, encoding / transcoding, and channel related bridging and transforms to interface with any real-world application, while providing mitigation and enhancement support for environmental, output / playback devices, and other environmental factors. Together with the segmentation and related indexing enabled in stage 4, and associated transform control options in stage 5, this collectively enables flexible output interfacing that constitutes an important benefit of the present invention,
[0103] Stage 6 is selectively employed in one of two primary modes. The first mode generates real-world based output used to optimize the training of the Al for a particular purpose. This stage uniquely enables the Al to employ generative processes and other means for deep levels of realism or other desired effects. Unlike prior approaches that used synthetic means to introduce ambience, dynamics, channelization, and immersion and other factors, humans are extremely sensitive to even minor relative quantization, equalization, phase, and other differences, which destroy a sense of accuracy and realism. The application of this stage, together with the use of a 1A reference standard, lesser quality examples of 1B-1E, and associated constraints, ensures that the described levels of fundamental realism, fidelity and creative objectives are supported,
[0104] Tire second mode of operation is to apply desired optimization for a given target purpose, such as for binaural audio (6A), for stereo mastering (6B), for style / sound / timbral character purposes such as by impulse response generation (6C), for ATMOS or other multichannel purposes (6D), for live event purposes such as transfer control (6E), and for other targeted purposes.
[0105] While the referenced I / O types listed in stage 6 as part of tire preferred embodiment noted herein are supportive of the puipose of this invention, it must be noted that this invention is very specifically designed to be modularly adaptive, such that oilier types of I / O, even ones not related to audio / music, are easily able to be inserted within the architecture of the present invention. In fact, the architecture of this invention is very’ specifically architected to inherently support such options. Therefore, the diagram shown in FIG. 11 should not be read as limiting with regard to the output forms and modalities able to be supported by the present invention.
[0106] Uris capability is able to be used to enable the Al and supporting subsystems and phases to optimize and support a wide range of interfacing with other software or hardware for other purposes. It is an inherent part of this design to be able to selectively leverage other analog, digital and related hardware, and software together with the core system.
[0107] FIGS. 12-17 illustrate sequential flow diagrams for select stages described in the foregoing.
[0108] FIG. 12 is a flow diagram for an audio sourcing stage of an audio restoration or enhancement process according to one embodiment of the present invention. FIG. 12 shows typical levels and types of example training data for enabling the model’s training based on the differential factors between the levels of accuracy and fidelity between the ground truth (1A) Reference Ground Truth and each level of truncation, compression, artifact-induction, and related factors for sources IB through IE. Utilizing a wide variety of levels of example training types and levels helps to make the model of the present invention more effective and robust.
[0109] FIG. 13 is a flow diagram for a preprocessing and FX processing stage of an audio restoration or enhancement process, according to one embodiment of the present invention. FIG, 13 diagrams how the system of tire present invention preprocesses input / source audio data to create additional levels of synthetic training media / data, by transforming synthetically created ground truth training media / data into various levels of bandwidth-degraded, compression-degraded, sampling rate and format-degraded training examples.
[0110] FIG. 14 is a flow diagram for an artificial intelligence (Al) training and deep learning model stage of an audio restoration or enhancement process according to one embodiment of the present invention, FIG, 14 is a depiction of the training and deriving of optimized training results by the present system. The system leverages an encoder-decoder architecture, including using a Swin Transformer in a U-Net style architecture, and employs a multi-scale transformer architecture with a series of transformer blocks that are applied at different levels of the feature hierarchy. The system therefore leverages a transformer-based attention head that extracts scaleindependent and related feature details from the media / data examples as defined.
[0111] The system trains the network to map from low-quality inputs to high-quality outputs by employing a curriculum learning approach: starting with easier restoration examples using supervised learning of structured example data, and gradually increasing complexity as the system moves to substantially unsupervised learning of largely unstructured data. The system combines multiple loss terms, including, by way of example and not limitation: a) element-wi se loss (e.g., LI or L2) for defining overall structure based on the highest quality data / media option examples; b) perceptual loss using Al models capable of processing spectrographic image and similar options (e.g., Vision Transformer (ViT) and its scale-independent Al network variants such as Pyramid Vision Transformer (PVT)) to capture features at various levels); and c) adversarial loss (generative adversarial network (GAN)-based) to identify and map to highest perceptual quality media / data with the highest fidelity details.
[0112] The model is trained progressively on different levels of degradation and limitations, starting with mild degradations (e.g., CD-Quality) and gradually introducing progressively more severe degradations as indicated in FIG. 12 (e.g., low bit rate MP3).3
[0113] The system implements extensive data synthesis and augmentation to increase the model's generalization ability and provide a greater range and number of examples, including randomized bandwidth truncations, phase shifts, and related transforms, at varying levels of degradation. The model is modular, such that the system is able to support future training options including and beyond grokking.
[0114] The system pre-trains on a large set of diverse audio-spectrogram conversions (e.g., MEL spectrograms, etc.) before doing any fine-tuned training, leveraging principles of transfer learning.
[0115] To evaluate the model, the system uses both quantitative metrics (PSNR, identified / restored bandwidth, compression, and dynamics) and qualitative assessments based on typical Human Reinforcement Learning feedback. Additional metrics able to be used include, but are not limited to, signal-to-noise ratio (SNR) (indicating level of desired signal relative to background noise, with higher values indicating better quality), total harmonic distortion (THD) (quantifying the presence of harmonic distortion in the signal, with lower values indicating less distortion and higher fidelity), perceptual evaluation of audio quality (PEAQ) (based on an ITU-R BS, 1387 standard for objective measurement of perceived audio quality with a score from 0 (poor) to 5 (excellent)), mean opinion score (MOS) (a subjective measure with listeners rating audio quality on a scale of 1 to 5), frequency response (measuring how well the system reproduces different frequencies, which is ideally flat across a spectrum of 2 Hz to 100 kHz), intermodulation distortion (IMD) (measuring distortion caused by interaction between different frequencies, with lower values indicating better fidelity), dynamic range (i.e., the ratio between the loudest and quietest sounds in the audio, with higher values usually indicating better quality), spectral flatness (measuring how noise-like or tone-like a signal is in comparison to ground truth data, which is useful for accessing the presence of unwanted tonal components and phase anomalies), cepstral distance (measuring the difference between two audio signals in the cepstral acoustic domain, with smaller distances indicating higher similarity and, typically, better fidelity), perceptual evaluation of speech / vocal quality (PESQ) (an ITU-T standard for assessing speech quality with scores from -0.5 to 4.5, with higher scores indicating better quality), perceptual objective listening quality analysis (POLQA) (i.e,, un updated version of PESQ able to be used for super wideband audio and evaluated between 1 and 5), articulation index (Al) or speech / vocal intelligibility index (SII) (measuring the intelligibility of speech in the presence of noise, with scores from 0 to 1, and with higher values indicating better intelligibility), modulation transfer function (MTF) (assessing how well a system preserves amplitude modulations across frequencies, which is important for maintaining clarity and definition in complex audio), noise criteria (NC) or noise rating (NR) curves (used to assess background noise levels in differentenvironments, with lower numbers indicating quieter environments), loudness (e.g., measured in loudness units relative to full scale (LUFS), which is useful for ensuring consistent loudness across different audio materials), short-time objective intelligibility (STOI) (measuring intelligibility of speech signals in noise conditions, with scores from 0 to 1 and with higher values indicating better intelligibility), binaural room impulse response (BRIR) metrics (i.e., various metrics derived from BRIR measurements to assess spatial audio quality including interaural cross correlation (IACC) and early decay time (EDT)), spectral centroid (indicating where the “center of mass” of the spectrum is located for assessing brightness or dullness of a sound), weighted spectral slope (WSS) (measuring the difference in spectral slopes between original and processed speech, with lower values indicating higher similarity), and log likelihood ratio (LLR) (comparing differences between linear predictive coding (LPC) coefficients of the original and processed speech, with lower values indicating higher similarity).
[0116] The system is able to train the model using degradation type / level as an additional input to the transformers taken in sequence and in-parallel, and apply a cascaded approach where the output is iteratively refined through multiple stages as described in FIG. 15.
[0117] FIG. 15 is a flow diagram for an indexing and transform optimization stage of an audio restoration or enhancement process according to one embodiment of the present invention. FIG. 15 further delineates the processing and related handling of the media based on a segmented set of styles, delivery / output formats, and related options. The system also allows for selective control of the system based on directed input and selection of options by human and / or AI-automated processes. For example, in one embodiment, the indexing and transform optimization stage includes a first Al-based segmentation analysis substage, in which optimal transform options are determined based on audio state, style, and delivery mode. This segmentation analysis then leads to an Al-based differential processing substage, in which the Al is applied for restoration, improvement, modification, or for creative functions (e.g., recovering information). After the differential processing, the audio data is then able to be put through an automated AI-based inferencing substage, in which objectives are achieved automatically, with support for transfer learning for other creative transformations. This stage is able to operate based on user-directed transformations for particular known types of conversion (e.g., remasters for DOLBY ATMOS, compression, recovery', etc.).
[0118] FIG. 16 is a flow diagram for a dimensional and environmental transform stage of an audio restoration or enhancement process according to one embodiment of the present invention. FIG. 16 provides additional sets of important transforms able to be applied to support the range of ty pical analog and digital formats and standards, such as immersive, multi-channel, stereo, and related real-world applications and usage that must be handled. These transforms are able toinclude analog, digital and dimensional transforms, with controls for each transform, including temporal-spatial, encoding / transcoding, and channel bridging. Additional specific transforms are also able to be applied, including but not limited to those associated with particular formats (e.g., DOLBY ATMOS), a compression “glue pass,” and / or environmental factors mitigation. The transformed data is able to then pass through a flexible output interface for a real-world application to produce a final transformed audio output.
[0119] FIG. 17 is a flow diagram for an output mode stage of an audio restoration or enhancement process according to one embodiment of the present invention. FIG. 17 diagrams the modular approach of the present invention for supporting currently defined standards and supported uses for various use cases. The modular approach also allows the system to support integration of new or otherwise additional formats. FIG. 17 illustrates that Stage 6 of the audio restoration or enhancement process provides for support for other optional use cases, such as gaming, immersive realism such as in mixed reality, hyper-realism, BCI-enabled deep realism, and other such applications. Finally, tire architecture supports feedback into the system for progressive optimization based on Al-driven processes.
[0120] Turning the attention now to the PDM method of lossless processing capability of the system of the present invention, FIG. 18 depicts a signal format traditionally used via Pulse Code Modulation (PCM) without any improvements enabled by the present system. Alternatively, FIG.19 shows a version of the spectrum with PDM sampled at 5000,000 Hz as part of the present system, showing the spectrum devoid of aliasing, phase, distortion, and other artifacts.
[0121] Without the ability to eliminate phase, quantization, transient, frequency, and sampling-related artifacts and distortion, which is enabled by the present invention, the benefits and capabilities described herein would not be possible.
[0122] FIG. 20 is a graph showing aliasing in the use of low-pass or high-pass filters for audio data. As shown in FIG. 20, traditional PCM causes aliasing and therefore artifacts appear above the Nyquist frequency. While many systems employing PCM assume 20 Hz to 20 kHz to be the range of perceptible audio and design filters accordingly, as previously noted, such assumptions are faulty and therefore the artifacts induced in such PCM methods are perceptible. Furthermore, as shown in FIG. 21, even below the Nyquist frequency (i.e., below 20 kHz), reflected in-band aliasing occurs, affecting perceptible qualities of the sound, such as frequency, timbre, imaging, and other qualities. The ability to eliminate the need for low pass and high pass filtering in ground truth training and reference samples, as well as in the output generation workflow allows for the elimination of reflected in-band aliasing under the Nyquist frequency.
[0123] In addition to bypassing the frequency related effects, and the resulting tonal, timbral and intonation-related impact on the sound, the system therefore also enables the elimination ofphase anomalous effects that all prior and currently available methods introduce to analog and digitized audio. The result of this is that imaging, dimensional characteristics, and the localization and positional representations of the original emitting elements will not be impacted during new recording and output. Also, the ability to apply the Al-driven semantic pattern identification of phase effects allows the system to eliminate them from existing recordings, resulting in the first and only system capable of such recovery’ and restoration.
[0124] FIG. 23 is a schematic diagram of a pulse density modulation process for capturing and encoding lossless audio according to one embodiment of the present invention. In embodiments of the present invention there are 12 stages to the system and method, defined as stages 1 through 12 identified in FIG. 23 as 1-12. One of ordinary' skill in the art will understand that the stages provided with reference to FIG. 23 are distinct from those described in FIG. 11, as the processes depicted in the diagrams of each figure are distinct. In embodiments of the present invention stages 1-8, along with one or more of stages 9- 12 are required. Which of the stages 9-12 are chosen depends on the specific output format and usage that is required as described below.
[0125] Operational Sequence of System and Method
[0126] As shown in FIG. 23 and more particularly in FIG. 24, stage 1 is tire capture phase, which requires a minimum of the following configuration and enabling capabilities. A diversity microphone set of two (2) omni directional and two (2) cardioid microphones capable of being equalized to - 1.5 dB and +1.5 dB, between <=5 Hz and >=65 kHz, and having a self-noise of no more than 15 dBA is used. First, the set is established for diversity capture using a stereo pair of cardioid pattern microphones as appropriate for the given distance, application, and recording purpose. Secondarily, for the same source, the microphone set includes a head related transfer function (HRTF)-configured omni-directional microphones complying with the above specification.
[0127] The defined two or more sets of microphones enable the invention to be able to differentiate betyveen the room conditions and related frequency, phase, delay, and other responses and the sound source itself. While other numbers of sets of microphones are able to be used to enable a diversity recording capability’, to be faithful to the extent and capabilities of human hearing and perception, the microphones need to, at a minimum, have the specifications and capabilities defined herein. Furthermore, while the microphone configuration described supports the necessary’ capture requirements as described, other configurations that support these requirements are also able to be used if available.
[0128] As shown in FIG. 25, stage 2 is an analog equalization stage using analog audio equalization hardyvare or similarly purposed devices appropriate to enable capture at 90V orgreater rail to rail Direct Current (DC) voltage, to ensure support for a dynamic range of at least 145 dB at a slew rate that supports fundamental and partial harmonics through the entire capture range of 65kHz or greater. The purpose is to ensure that the bandwidth is ideally at least 65 kHz, and that the inherent signal-to-noise ratio (SNR) and associated dynamic range supported is commensurate with Stage 6 and its requirements as described below. The goal of the equalization at this stage is to provide support for what is referred to in the industry as a near field Fletcher Munson or other desired equalization curve that is compliant with the final audio preference and related requirements, if any, prior to the final playback method, medium, or environment related requirements as further described below.
[0129] As shown in FIG. 26, stage 3 is an analog dynamics restoration stage via an analog audio compressor, or similarly purposed devices that do analog dynamics restoration appropriate to ensure capture at 90V or greater rail to rail DC voltage, to provide support for an SNR enabling a dynamic range of at least 145 dB at a slew rate that supports fundamental and partial harmonics through the entire capture range of 65kHz or greater. The purpose is to ensure that the bandwidth is at least 65 kHz and that the inherent signal to noise ratio and associated dynamic range supported is commensurate with the analogous limits of Stage 6 as described below.
[0130] As shown in FIG. 27, stage 4 uses either the same or a different equalizer, or another audio device that is able to serve the same purpose to enable analog audio equalization necessary for target playback medium and / or environment. In contrast to stage 2, stage 4 provides the optimal support for a particular target recording medium, purpose or environment. For example, if the final output of this stage 4 is to be optimal for mastering to vinyl, a Recording Industry Association (RIAA) or comparable curve is able to be implemented at this stage to allow for required results compatible with the usage requirements. For example, an RIAA curve is important to ensure that the resulting audio is able to be properly used for mastering and creating vinyl records.
[0131] Stage 4 supports the creation of one or more target environment compliant equalizations, the output of which then may continue to stage 5 for differential modulation and related transform as described. Stage 4 is able to be bypassed by having the output of stage 3 proceed directly to stage 5, The puipose of stage 4 is to provide a flexible option to support the specific requirements most optimally, rather than leaving such optimization to others at some point in the future. Further, the fact that, as part of the system, Stage 4 supports 100 kHz or greater bandwidth, with a slew rate, SNR, and associated dynamic range at or above the perceptual and related limits of human hearing in the analog domain, allows the invention to mitigate any digital jitter, noise, and related factors that reduce the audible accuracy when compared to what a person would hear if at the original location of the audio event.
[0132] As shown in FIG. 28, stage 5 employs differentially modulated pulse shaping using hardware or software that reduces aliasing effects throughout the subsequent band-limited stages and reducing temporal and spectral artifacts that otherwise are introduced in the analog-to-digital and related conversion processing. The pulse shaping is able to be based on any standard or purpose-specific nonlinear or linear-based implementation, including triangle, square wave, fractal, or sine wave based. As is commonly known by many in the field, each approach has its own benefits and related considerations, in terms of optimal usage cases and constraints. As such, a unique aspect of the system is that it has an architecture and signal and processing flow to support any one or more of them, while obtaining maximal benefit of their use case, due to the unique full spectrum and related perceptual human hearing factors driven capabilities.
[0133] As shown in FIG. 29, stage 6 is the analog to digital conversion process, which is able to employ, for example, the Direct Stream Digital (DSD) audio format-based encoding at an ideal, though not fixed, sampling rate of 11.2 MHz or above. The sampling rate is chosen in the analog to digital (A / D) conversion stage to maintain conversion limits at or above the physical and perceptual limits of human hearing as indicated in FIG. 22. While other existing standard or proprietary’ methods are able to be substituted for this in stage 6, such methods must be able to meet or exceed the noise, dynamics and other limits supported in stage 1-4 and optimized via stage 5. Standard Pulse Code Modulation (PCM)-based methods are not capable of supporting this level of audio performance and associated requirements, due to phase and aliasing anomalies and artifacts, anomalies and noise introduced via PCM functional requirements.
[0134] Stage 9 is an adaptive stage that serves to optimize for specific environments and usage. Stage 10 is designed to provide the highest available quality at the lowest possible bandwidth and file size as described herein. Stage 11 is a required stage for playback or storage of the unique audio in analog format. The last optional stage, stage 12, supports analog or digital storage in one or more of the four modes described herein.
[0135] The following are the core hardware and software components and requirements within the selected Direct Stream Digital (DSD) conversion: Decimation Filter: Sample rate conversion, optimally at a ratio of 16: 1 or greater, such as from 11.2 MHz to a typical range of higher-end PCM bandwidths; Quantization: To a higher bit depth of at least 16:1, from 1 -bit samples to 16 bit or higher PCM samples, to support an increase in the dynamic range and allow for complex digital routing and signal processing; Resampling: Frame boundary aligned sample rates, reducing any previously audible jitter, and enabling digital math to be executed as required with support of solid frame boundary alignment and tracking; Normalization & Dithering:Options such as TPDF or other dither are able to be used for lower quantized sample depth PCM rates of 16 bits or less; and Encoding: the data into a fully compliant PCM, Digital extremeDefinition (DXD) or other digitally editable format, capable of standard digital processing, such as equalization, delay, reverb, compression, or other processing commonly, although not always, implemented via plugin, firmware, or other code.
[0136] Preferably, the DSD data that constitutes the ground truth data in one embodiment of the present invention is not processed by any high pass filter or low pass filter (with the possible exception of only a DC offset filter), such that no data is lost during the preprocessing of the data and the otherw ise-filtered portions of the source audio are still able to be used to train the audio restoration and enhancement system as described herein.
[0137] As shown in FIG. 30, stage 7 uses software or an embedded hardware or related approach to provide the jitter-corrected frame boundary and related alignment necessary- to optimally support any possible data and associated sampling rate and related conversion to PCM as desired. A frame-integrator pulse function conversion using an adapted pulse function as described above supports a unique lossless PCM conversion in subsequent stages. This is necessitated because there is a difference between the natural integer multiples of PCM frequencies versus DSD, as well as to allow for clock cycle jitter that must be corrected for. This effectively mitigates the truncation or other irrecoverably discarded information that represents perceptually or otherwise audible detail (i.e., “lossy”) decimation that are otherwise necessary and part of all current implementations.
[0138] As shown in FIG. 31, stage 8 defines the use of a Weiss or other DSD to PCM conversion function, with appropriate Triangular Probability Density Function (TPDF) or other dither pattern as most appropriate to the output and intended use as would generally be known byaudio and music recording professionals. The purpose of this is to enable support for standards-based digital audio tools and methods, rather than having to introduce new, proprietary or other non-standard approaches that potentially confound usage and adoption.
[0139] As shown in FIG. 32, stage 9 is the first of four optional stages of the overall system shown in FIG. 23. Stage 9 uses software or its hardware -based equivalent for digital effects processing, such as digital delay, reverb, equalization, compression, deverb, or any other effects processing. The purpose of this is to provide for a user or other requirements of those using this type of system for typical music or related audio purposes.
[0140] As shown in FIG. 33, stage 10 uses software or a hardware -based equivalent that provides optional support for down-sampling of any standard or other types of resampling often employed in music and other audio professions, along with any associated dithering or requisite noise profiling desired. This is useful to reduce the requisite file sizes and associated bandwidth, while minimizing any audible artifacts. While the core nature of this system supports true perceptually lossless audio, it also enables support for “lossy” (i.e., non-full spectrum andotherwise reduced audible detail based) implementation. However, even in lossy use cases, the perceptually and otherwise relevant characteristics provides distinct, audible enhancement and associated fidelity compared to other approaches, meaning there is more perceptually, and otherwise relevant audible information provided by the system, even after lossy or lossless compression, than any existing prior art system.
[0141] As shown in FIG. 34, stage 11 provides optional support for analog output for playback or other purpose. Any standard format is supported, and support of stage 4 is able to be combined to facilitate and otherwise optimize for this stage.
[0142] As shown in FIG. 35, stage 12 provides storage function support for any of the following 4 types of storage: 1. Raw, directly from stage 8 with none of the processing of stages 9-11 employed; 2. Playback-adapted storage, employing reduced size and bandwidth requirements; 3. Reduced size storage, using stage 10 processing as defined herein; and 4. Pure analog storage, employing stage 11 analog output conversion.
[0143] As defined herein, the system employs an advanced hardware-enabled capture approach, together with uniquely integrated and configured audio capture hardware and software that drives the conversion sequence and supporting algorithms as described herein. Further, leveraging differential Al enabled analysis, the system in tegrates the results of Brain Computer Interface (BCI) or other comparably sourced data to optimize for actual perception and related results in those areas as noted. The result is a new and unique process to capture / 'record, represent, convert, and store sound to and from a digital or analog source or medium, which dramatically reduces size, bandwidth, and related overhead at any level of perceived audio quality. Analog components are selected and configured in a manner designed to avoid any need for low pass filtering and noise profiling w ithin the boundaries of human h earing.
[0144] Together with the Al-assisted approach described herein, the system avoids all aliasing and phase skew back into the perceptual and physiological ranges of human hearing. Hie result of this is the first acoustically lossless inventive method to capture, store and transfer audible information for human hearing. The resulting capabilities go beyond simply improving sound to effectively match the limits of human hearing and perception. The audio output of the system is also able to be used to provide a reference audio standard for training Al via, among other things, providing a universal reference standard, or ground truth.
[0145] Hie audio processing techniques described herein are able to be used for a variety of purposes for improving the field of the art, including but not limited to those described belowz
[0146] Fidelity Enhancement:
[0147] The Al-driven fidelity enhancement capabilities of the present invention represent a large improvement in audio quality enhancement not possible in prior art systems and methodsowing to the use of the unique generative Al and training of the present invention. By leveraging advanced machine learning algorithms and training approaches, the system is able to analyze audio content across multiple dimensions (e.g., frequency spectrum, temporal characteristics, and spatial attributes) to identify and correct imperfections that detract from the original artistic intent as defined by actual ground truth perfection.
[0148] The system employs a unique neural network trained on vast custom datasets of ultra-high through low quality audio, allowing the system to recognize and rectify issues, such as frequency issues, phase anomalies, and various types of modulation, aliasing and other distortion types, resulting in a dramatically clearer, more detailed, and more engaging listening experience across all types of audio content.
[0149] For music and sound in a range of applications, the system therefore brings out the nuances of instruments and vocals that are generally masked in the original recording. For speech, it ensures every word is crisp and intelligible, even in challenging acoustic environments. The end result is audio that approaches the previously unobtainable ground truth (i.e., a perfect representation of the sound as it was intended to be heard).
[0150] Bandwidth Optimization:
[0151] The bandwidth optimization technology of the present invention represents a significant improvement for streaming sendees, and content delivery networks. By employing a unique set of Al training and optimization methods, the system is able to intelligently analyze and adapt audio con ten t to make optimal use of available bandwidth wi thout compromising on lossless files standards, or perceptually defined quality.
[0152] The system works by identifying all perceptually and audibly relevant information in the audio signal and prioritizing its optimization based on transmission and data rate constraints. The significant amount of noise, compression artifacts, and aliasing related masking elements that often account for 50% or more of tire size of many recordings are eliminated. The actual information is intelligently compressed and recast into the desired standard format using the advanced Al models of the present invention, allowing for significant reductions in data usage -generally up to 50% or more - while maintaining, and in many cases improving, the perceived audio quality.
[0153] Unlike present compression means, the lossless audio produced by the present invention stays not only lossless, but is also able to remain in the same standard lossless formats, with the same being true with lossy formats like MP3 and others. This avoids the need to distribute new types of players, encoder / decoders and other technologies, enabling immediate usability and global deployment. Moreover, the system is able to adapt in real-time to changing network conditions, ensuring a consistent, high-quality listening experience even in challengingconnectivity scenarios. This not only enhances user satisfaction but also reduces infrastructure costs for service providers.
[0154] Imaging Correction:
[0155] Imaging correction capabilities of the system of the present invention improves the spatial perception of audio, particularly for stereo and multi-channel content. Using advanced Al algorithms, the system is able to identify and correct issues in the stereo field or surround sound image, resulting in a more immersive and realistic audio experience.
[0156] The system analyzes the phase relationships between channels, corrects phase and intermodulation anomal ies, and perfects the separation and placement of audio elements within the soundstage based on ground truth training and related definitions, resulting in a wider, deeper, and more precisely defined spatial image and associated sound stage, bringing new life to everything from classic stereo recordings to modem surround sound mixes. Unlike prior approaches, this works with both traditional speakers as well as headphones and in-ear monitors.
[0157] For stereo content, this means a more expansive and engaging soundstage, with instruments and vocals precisely placed and clearly separated. In surround sound applications, it ensures that each channel contributes accurately to the overall immersive experience, enhancing the sense of being "there" in the acoustic space.
[0158] Noise Reduction:
[0159] The Al-powered noise reduction capabilities of the present invention provide a notable improvement in audio cleanup and restoration. Unlike traditional noise reduction methods that often introduce artifacts or affect the quality of the desired signal, the system of the present invention uses advanced machine learning to intelligently separate noise from the primary audio content.
[0160] The Al model is trained on a vast array of noise types - from compression and digital encoding artifacts and background hum, to intermittent types of spurious noise, allowing the system to identify’ and remove these unwanted elements with unprecedented accuracy.Additionally, the system is able to adapt to novel noise profiles in real-time, making it effective even in unpredictable acoustic environments.
[0161] The result is clean, clear audio that preserves all the details and dynamics of the original signal. This technology is particularly valuable in applications ranging from audio restoration of historical recordings to real-time noise cancellation in telecommunication systems and hearing aids.
[0162] Dynamic Range Optimization:
[0163] The dynamic range optimization capabilities of the present invention represent a paradigm shift in audio dynamics. Using sophisticated Al algorithms, the system analyzes the4dynamic structure of audio content and intelligently adjusts it to suit different playback scenarios and devices based on a range of ground truth examples beyond current recording methods and approach, all while preserving the original artistic intent.
[0164] The system goes beyond simple compression or expansion by understanding the contextual importance of dynamic changes, preserving impactful transients and dramatic silences where such elements are crucial to the content, while subtly adjusting less critical variations to ensure clarity across different listening environments.
[0165] This intelligent approach ensures that audio remains impactful on high-end audio systems, while still being fully enjoyable on mobile devices or in noisy environments, which is particularly valuable for broadcast applications, streaming services, and in-car audio systems, where maintaining audio quality across a wide range of listening conditions is crucial.
[0166] Spectral Balance Correction:
[0167] The spectral balance correction capabilities of the system of the present invention utilize the Al to achieve ground truth-perfected tonal balance in any form of audio content. By analyzing the frequency content of audio in relation to vast databases of beyond master-quality references, the system identifies and corrects spectral imbalances that detract from the natural and pleasing quality of the sound.
[0168] The Al does not simply apply broad, one-size-fits-all equalization. Instead, the system understands the spectral relationships within the audio, preserving the unique character of instruments and voices while correcting problematic resonances, harshness, or dullness, resulting in audio that sounds natural and balanced across all playback systems. The system is therefore invaluable in mastering applications, broadcast environments, and consumer devices, ensuring that audio always sounds its best, regardless of the original production quality' or the playback system.
[0169] Transient Enhancement:
[0170] The transient enhancement capabilities of the present invention provide a higher level of clarity and impact to audio content. Leveraging advanced Al algorithms, the system identifies and enhances transient audio events (i.e., split-second bursts of sound that characterize percussive elements like the attack of a drum hit or the pluck of a guitar string). Furthermore, by using an extensive amount of custom-created ground truth examples in training, the system is able to define and restore the sonic character based on frequency vs. phase over time and related partial harmonics relationships, resulting in a ground truth defined level of temporally accurate transient accuracy.
[0171] By intelligently optimizing these transients without negatively affecting the underlying sustained sounds by disregarding their temporal context, the system is able todramatically improve the perceived clarity and definition of the audio. This process does not only make the input audio louder, but also helps to reveal the subtle details that make the audio sound physically present.
[0172] The system is particularly effective in music production, live sound reinforcement, and audio post-production for film and TV, as it is able to provide additional character to flat or dull recordings, enhance the impact of sound effects, and ensure that every’ nuance of a performance is clearly audible.
[0173] Mono to Stereo Conversion:
[0174] The mono to stereo conversion capabilities of the system of the present invention provide an improvement beyond traditional up-mixing techniques, which is impossible prior to the Al-enabled system and technique of the present invention. Using advanced Al models trained on vast libraries of ground truth defined stereo content and other related custom audio data, the system is able to transform mono recordings into real, spatially accurate stereo soundscapes.
[0175] The system analyzes the spectral and temporal characteristics of the mono signal to intelligently reconstruct audio elements across the dimensional stereo field. This process does not add artificial reverb or delay, rather creating a ground truth-defined, real-sounding stereo imaging that respects the original character of the audio while restoring innate width, depth, and immersion that was collapsed in mono source material. The system therefore has particular use in remastering historical recordings, enhancing mono content for modern stereo playback systems, and improving the listener experience for any mono source material, with the results often rivalling true stereo recordings in their spatial quality and realism.
[0176] Stereo to Surround Sound Up-mixing:
[0177] The stereo to surround sound up-mixing capabilities of the present invention take two-channel audio to new dimensional levels or realism and presence. Powered by advanced Al algorithms, the system analyzes stereo content and intelligently distributes it across multiple channels to create a uniquely accurate immersive surround sound experience.
[0178] Unlike traditional up-mixing methods that typically result in artificial, phase-incoherent surround fields, the Al-enabled system understands the spatial cues inherent in the stereo mix. The system is able to identify individual elements within the mix and localize and distribute them naturally in the surround field based on ground truth training examples, creating a sense of envelopment that respects the original stereo image, and expanding it into three-dimensional space.
[0179] The system has particular use for home theater systems, broadcasting, and remastering applications, allowing vast libraries of stereo content to be experienced in rich,immersive surround sound, dramatically enhancing the listening experience without requiring access to original multi -track recordings.
[0180] Legacy Format to Immersive Audio Conversion:
[0181] The legacy format to immersive audio conversion capability of the system of the present invention bridges the gap between traditional audio formats and cutting-edge immersive audio experiences. Using state-of-the-art Al training and physics informed and optimized approaches, the system transforms content from any legacy format (e.g., mono, stereo, or traditional surround) into fully immersive audio experiences compatible with formats such as DOLBY ATMOS, SONY 360 REALITY AUDIO, as well as other current and future standards.
[0182] The Al does not only distribute audio to more channels, but rather understands the spatial relationships within the original audio and extrapolates them to create a ground truth accurate, phase and frequency-coherent three-dimensional soundscape. Individual elements within the mix are able to be identified and placed as discrete objects in 3D space, allowing for a level of immersion previously impossible with legacy content.
[0183] The system provides for additional opportunities for content owners, allowing entire back catalogs to be remastered for immersive audio playback using fully Al-automated generative transforms trained on custom-created ground truth libraries. This also provides a benefit in broadcast and streaming applications, enabling the delivery of immersive audio experiences even when only legacy format masters are available.
[0184] Adaptive Format Transcoding:
[0185] The adaptive format transcoding capability of the system represents the cutting edge of audio format conversion. Powered by sophisticated Al algorithms and unique ground truth reference constraints, the system dynamically converts audio between various formats and standards, optimizing the output based on the target playback system and environmental conditions.
[0186] The Al-based system does not merely perform a straight conversion, but also understands the strengths and limitations of each format and adapts the audio and associated requirements accordingly. For instance, when converting from a high-channel-count format to one with fewer channels, the system intelligently down-mixes in a way that preserves spatial cues and maintains the overall balance of the mix, considering phase vs frequency and the interplay of the format with those and related constraints.
[0187] Moreover, the system is able to be set to adapt in real-time to changing playback conditions. In a smart home environment, for example, the system is able to seamlessly adjust the audio format as a listener moves between rooms with different speaker setups. This ensures the best possible listening experience across all devices and environments.
[0188] Dialogue Intelligibility Enhancement:
[0189] The dialogue intelligibility enhancement capability of the system addresses one of the most common complaints in modem audio content, namely unclear or hard-to-hear dialogue. Using advanced Al models trained on vast datasets of clear speech, the system is able to identify and enhance dialogue within complex audio mixes without affecting other elements of the soundtrack,
[0190] The system goes beyond simple frequency boosting or compression and understands the characteristics of human speech and perceptual hearing factors and limitations, and separates it from background music, sound effects, and ambient noise. The system then enhances the clarity and prominence of the dialogue in a way that sounds natural and preserves the overall balance of the mix.
[0191] The system provides a benefit in broadcast, streaming, and home theater applications. It ensures that dialogue is always clear and intelligible, regardless of the viewing environment or playback system, dramatically enhancing the viewer experience for all types of content.
[0192] Audio Restoration:
[0193] The audio restoration capabilities of the system of the present invention represent an improvement in the ability to recover and enhance degraded audio recordings. Leveraging powerful Al algorithms, ground truth data, and associated training methods, the system is able to analyze damaged or low -quality audio and reconstruct it to a level of quality that often surpasses the original recording, while maintaining frequency vs phase, format-specific and other key constraints while doing so. Without this unique set of Al capabilities that govern the process, a significant amount of articulation and realism previously had to be sacrificed. Similarly, the ground truth reference sources that are trained on enable a level of perfect standards reference that did not exist before, and therefore were not able to be applied as restoration and optimization constraints to any process.
[0194] The system is trained on a vast array of audio imperfections - from the wow and flutter of old tape recordings to a range of digital artifacts in CDs and other digital audio formats. This is able to identify hundreds of primary, secondary, and other issues, and not only remove them but also reconstruct the sample level and temporally defined audio that should have been there, thereby going far beyond traditional noise reduction or editing techniques.
[0195] The system is therefore particularly valuable for archivists, music labels, and anyone dealing with historical audio content and is able to breathe new life into recordings that were previously considered beyond repair, preserving audio heritage for future generations.
[0196] Personalized Audio Optimization:
[0197] The personalized audio optimization capabilities of the system of the present invention bring a new level of customization to the listening experience. Using generative and related machine learning approaches, coupled with unique ground truth training and reference data sets defined within the requirements of human hearing and perception using a broad frequency range, the system is able to analyze a listener's preferences, hearing capabilities, and even current environment to dynamically adjust audio content for optimal delivery.
[0198] The system is able to produce a personalized hearing profile for each user, understanding their frequency, phase and related sensitivities and limitations, dynamic range profile and preferences, and includes subjective tastes in aspects such as frequency accentuation / amelioration, timbral characteristics, and imaging vs soundstage characteristics. It is then able to apply these constraints to any audio content in real-time, ensuring that everything sounds its best for that specific listener.
[0199] Moreover, the Al is able to adapt to changing conditions. If the listener moves from a quiet room to a noisy environment, or a room with a different damping profile for instance, the system automatically adjusts to maintain intelligibility and enjoyment based on the listening device and criteria. The system has applications ranging from personal audio devices to car sound systems and home theaters, ensuring the first ground truth defined listening across virtually any situation.
[0200] Acoustic Environment Compensation:
[0201] The acoustic environment compensation capability of the system of the present invention brings studio-quality sound to any listening environment. Using advanced Al algorithms and custom ground truth defined training and reference / constraint data, the system analyzes the acoustic characteristics of a space in the context of a massi ve set of interrelated constraints that were impossible to consider prior to these Al enabled methods, and apply realtime corrections to the audio signal, effectively neutralizing the negative impacts of the room.
[0202] The system goes beyond traditional room correction system s by, first, not just adjusting frequency response, but also understanding complex room interactions, early reflections, and resonances, partial harmonics vs. listener perception interactions and preferences, and applying corrections that make the room 'disappear' acoustically as much as desired. Tire result is a listening experience that is as close to the original studio mix as possible, or even leverages ground truth references to go beyond that level of perfected sound, regardless of the actual physical space. Further, most systems employ frequency equalization as a primary goal. In contrast, the present system addresses inter-related factors such as frequency vs. phase, and perception vs. playback device nonlinearities, while ensuring intonation, partial harmonics and other key elements are maintained or recovered based on ground truth.
[0203] The system has applications ranging from home audio and home theaters to professional studio environments and ensures consistent, high-quality audio playback across different rooms and spaces, which is particularly valuable for professionals who need to work in various environments.
[0204] Future Format Adaptation:
[0205] The future format adaptation capabilities of the system of the present invention allow for future -proofing audio content and systems. Using highly flexible Al models, the system is able to learn and adapt to new audio formats and standards as they emerge, ensuring that today's content and hardware investments remain viable well into the future.
[0206] As new audio formats are developed, the system is able to be quickly trained to understand and work with these formats without requiring a complete overhaul, meaning that content created or processed with the system today is easily able to be adapted for the playback systems of tomorrow. Because the heavy lifting is done prior to playback, the approach enables existing and future playback hardware and related devices to continue to function. No special playback hardware, software or related decoding elements are required. However, rather than being locked-into a given hardware set, playback chain, or formats / standards, this approach is able to address the strengths, capabilities and weaknesses of new formats and related options using the same Al architecture.
[0207] For content creators and distributors, this means their archives remain perpetually relevant. For hardware manufacturers, it offers the potential for devices that are able to be updated to support new formats long after purchase.
[0208] FIG. 36 illustrates a schematic diagram for a method of compression according to one embodiment of the present invention. The compression method of the present invention includes a plurality of steps. First, media analysis is used to process the data. This approach provides the initial processing supporting the subsequent steps, allowing them to work with tire most perceptually and statistically significant features of the media as defined within hyperparameters and related constraints as defined herein. It also provides the foundational underpinnings for efficient compression by identifying redundancies and perceptual irrelevancies.
[0209] In one embodiment, the media analysis includes spectral analysis. In one embodiment, for audio data, a Short-Time Fourier Transform (STFT) with overlapping windows (e.g., an average of 35-45% overlap) is utilized. In one embodiment, Hamming and Hann window functions are utilized. In one embodiment, for video data, a 3D Fourier Transform is used on groups of frames (e.g., up to 16 frames) to capture temporal frequencies. In one embodiment, wavelet transforms (e.g., Daubechies wavelets) are used for multi-resolution analysis, supporting optimized localization of frequency content in time and space.
[0210] Statistical analysis is also able to be used on the data. In one embodiment, first-order statistics, such as mean, variance, skewness, and kurtosis for each frequency band, are computed. In one embodiment, second order statistics, such as autocorrelation and cross-correlation between frequency bands, are performed. In one embodiment, principal component analysis (PCA) is applied to identify the most-significant components in the frequency domain.
[0211] The media analysis is also able to include perceptual analysis. For example, in one embodiment, for audio data, psychoacoustic models based on critical bands and masking effects are implemented. In one embodiment, mel spectrograms and scaling for frequency mapping are used. In another embodiment, for video data, visual saliency models (e.g., deep learning-based saliency detection) identify perceptually important regions. In one embodiment, Just Noticeable Difference (JND) thresholding models are used to determine perceptual thresholds for different media components.
[0212] Die media analysis is also able to include temporal and spatial correlation analysis. In one embodiment, the temporal and spatial correlation analysis includes computing autocorrelation functions for different time / space lags to identify' periodic patterns, while forming compressed encoded spatial mappings. In one embodiment, the analysis includes implementing motion estimation techniques (e.g., block matching, optical flow, etc.) for video to capture temporal dependencies. In one embodiment, the analysis includes applying texture analysis methods (e.g., Gray Level Co-occurrence matrix) to capture spatial paterns in images or video frames.
[0213] After media analysis is performed, dimensional manifold selection takes place. The careful selection and optimization of dimensional manifolds enables high compression of accurate representations of the media. This step is crucial for achieving high compression ratios while preserving the essential structure and quality of the original content.
[0214] In one embodiment, for audio data, time-frequency manifolds are employed based on Gabor frames or wavelet packets, depending on the media types. This process implements adaptive time-frequency representations including matching pursuit and basis pursuit. In one embodiment, for video data, spatiotemporal manifolds are selected using 3D wavelet transforms and curvelet transforms, which supports motion-compensated temporal filtering for efficient represen tation of motion, while reducing redundancies and artifacts. For audio and video, selection of non-linear manifolds such as diffusion maps and Laplacian eigenmaps supports complex data structures inherent in the spatiotemporal complexity of media.
[0215] In one embodiment, the dimensional manifold selection includes dimensionality estimation, applying techniques based on maximum likelihood estimation methods to estimatethe intrinsic dimensionality of the data. This approach enables the selective use of fractal dimension analysis to represent the complexity of the data across different scales.
[0216] In one embodiment, the dimensional manifold selection includes manifold learning and optimization. In one embodiment, the selective transform optimization methodology supports manifold learning via Isomap, Locally Linear Embedding (LLE) and t-SNE to learn the structure of the data. In one embodiment, the present system supports Riemannian optimization techniques to fine-tune the manifold parameters, minimizing distortion while maximizing compactness. Finally, in one embodiment, sparse encoding, dropouts and other appropriate regularization methods are used to prevent overfitting and ensure smooth manifold structures,
[0217] In one embodiment, the dimensional manifold selection includes multi-scale analysis. In one embodiment, the multi-scale analysis includes implementation of multi-resolution analysis using wavelet packet decomposition and multi-scale singular value decomposition (SVD) In one embodiment, the use of this technique coupled with Uniform Manifold Approximation and Projection (UMAP) allows the system to optimize scale selection processes based on balancing local and global feature representation.
[0218] After dimensional manifold selection takes place, deep learning model training is performed. The deep learning model serves as the core engine for the compression process, learning to map between the original media space and the compact manifold representation. Its effectiveness directly enables the high compression ratios and the quality of the reconstructed media.
[0219] In one embodiment, the present invention includes a neural network architecture design employing an encoder-decoder architecture using variants of autoencoders as described herein (e.g., variational autoencoders, adversarial autoencoders, etc.). This approach incorporates self-attention and cross-attention architecture to capture long-range dependencies in the data. In one embodiment, the deep learning model training implements residual connections and skip connections to facilitate gradient flow and preserve fine-grained details, while avoiding overfitting during fine-tuning and related training processing. This method leverages adapted activation functions optimized for the specific manifold structure.
[0220] In one embodiment, the deep learning model training includes loss function formulation, which leverages a multi-term loss function incorporating reconstruction error, perceptual loss, and manifold consistency terms. This technique implements adaptive weighting of loss terms based on characteristics of the input data, w here 180-degree phase-rotated null tests are employed within the loss constraint calculations during training. This incorporates regularization terms to encourage sparsity and prevent overfitting.
[0221] In one embodiment, the training process leverages curriculum learning, starting with simple patterns and gradually increasing complexity. In one embodiment, the process leverages advanced optimization algorithms, such as Adam and RMSprop-based optimization functions with learning rate scheduling. This training process supports batch normalization, layer normalization, and weight normalization to stabilize training.
[0222] In one embodiment, the training includes fine-tuning and adaptation, supporting transferring learning techniques to adapt initial pre-trained models to specific types of media formats and content genres. In one embodiment, the training supports few-shot learning methods to quickly adapt to new, unseen types of media with minimal additional training.
[0223] In one embodiment, the compression includes entropy maximization. The use of entropy maximization ensures that the compressed representation retains critical informative aspects of the original media that are critical to the level of encoding. This step is crucial for achieving compression efficiency while maintaining the ability to reconstruct high-quality media from the compressed data.
[0224] In one embodiment, the entropy analysis includes computing Shannon entropy across different dimensionally-encoded features within the manifold representation. In one embodiment, the entropy analysis involves calculating mutual information between different components to identify redundancies, where, for two random variables X and Y, MI(X; Y) = H(X) - H(X|Y) = H(Y) - H(Y|X), where H(X) is the entropy of X, and H(X|Y) is the conditional entropy of X given Y. In one embodiment, higher-order entropy measures (e.g., Renyi entropy) are used to provide more accurate representation and comprehensive analysis of the information content.
[0225] In one embodiment, the entropy maximization includes information decomposition. In one embodiment, independent component analysis (ICA) is used to separate statistically independent components of the data. In one embodiment, non-negative matrix factorization (NMF) is used for parts-based decomposition of the data. In one embodiment, tensor decomposition is employed for higher-dimensional data structures.
[0226] The entropy maximization algorithms implement the principle of maximum entropy to optimize distribution of information across the manifold. In one embodiment, sparse coding techniques are applied to present the data using a minimal number of active components. In one embodiment, particular entropy maximization algorithms are applied based on the selected manifold structure.
[0227] hi one embodiment, the present invention implements a rate-distortion optimization (RDO) to balance between compression ratio and reconstruction quality. In one embodiment, the present system applies adaptive quantization schemes to allocate more bits to high-entropycomponents. In one embodiment, the present invention implements perceptual bit allocation, giving priority to perceptually significant components identified in the analysis phase.
[0228] Compression Application
[0229] Tliis step applies the developed compression scheme to actual media, transforming it into a highly compressed representation. The effectiveness of this stage directly determines the final compression ratio and the quality of the compressed media.
[0230] In one embodiment, the compression system of the application implements adaptive noise reduction techniques (e.g., wavelet denoising, non-local means filtering, etc.) to clean input media. In one embodiment, normalization procedures are applied to standardize input ranges across different media types and sources. In one embodiment, for video data, color space transformations (e.g., RGB to YCbCr) are performed to separate luminance and chrominance information, thereby increasing dimensional complexity.
[0231] In one embodiment, the trained deep learning model is then used to transform the media into an optimized manifold representation, hi one embodiment, batch processing techniques are used for efficient handling of large media files. In one embodiment, error handling and recovering mechanisms are used to optimize loss differential and related issues during transformation.
[0232] In one embodiment, vector quantization is implemented for groups of related manifold coordinates. In one embodiment, adaptive quantization schemes are applied to adjust quantization levels based on local entropy and perceptual importance. In one embodiment, non-uniform quantization techniques are applied to better match the distribution of manifold coordinates.
[0233] Tn one embodiment, entropy coding is implemented with arithmetic coding and range coding techniques. In one embodiment, the present system involves applying context-adaptive coding schemes that exploit inherent local patterns in the quantized data. In one embodiment, the present invention implements adaptive probability’ models that update based on observed symbols frequencies, providing predictive encoding optimization.
[0234] After compression, the data is encoding into a standard format. In one embodiment, the system applies a flexible, hierarchical data structure to represent the compressed data, manifold parameters, and model configuration. In one embodiment, the system leverages versioning mechanisms to ensure forward and backward compatibility as the compression technique evolves, lire system creates metadata structures to store essential information about the compression process and original media characteristics, supporting additional types of workflow and process automation.
[0235] In one embodiment, for audio, the system supports methods to embed the compressed data within standard containers like waveform audio file format (WAV) or Free Lossless Audio Codec (FLAG), utilizing custom chunks or metadata fields. In one embodiment, for video, the system is able to encapsulate the compressed data within containers such as MP4 or MKV, supporting private data streams and custom metadata tracks and associated workflows. This ensures compliance with specifications of chosen container formats, including proper header structures and stream synchronization.
[0236] Ihe present system is able to be used to generate both lossless and lossy versions compatible with standard formats. Tire decoder portion of the trained model is able to generate a reconstruction with lossless and lossy optimized constraints. In one embodiment, the present invention supports sequenced application of conventional lossy compression (e.g., MP3, AAC, H.264, HEVC, etc.) for reconstruction. In one embodiment, a lossless version sen es as a “base layer” in the final output, which is able to be lossless or lossy when transcoded for final output.
[0237] Ihe dimensional manifold representation of the present invention inherently eliminates redundancy by capturing dimensionally reduced, encoded structure of media, and the entropy maximization step helps separate signal from noise. Low-entropy components have higher correlation to noise and redundant information and these components are more selectively quantized out at higher rates. The deep learning model learns to focus on perceptually important features, further reducing non-essential information.
[0238] Ihe full system and process of the present invention allow for loss reconstruction, storing the difference between the original media and the lossy reconstruction and encoding this difference as described herein. This forms the “differential decoding layer” of the final output.
[0239] Hie system of the present invention also enables scalable compression for standard players able to play the standard lossy versions. The advanced decoder layer enables both modes for lossless and lossy playback. Intermediate quality levels are also able to be achieved by differentially decoding via the enhancement layer to the desired output format.
[0240] By combining these techniques, the system and method achieve high compression ratios while maintaining compatibility with standard formats and allowing for lossless reconstruction when needed. The system and method significantly reduce redundancy and nonsignal elements through the use of dimensional manifolds, and entropy maximization leveraging deep learning, while providing flexibility in output format and quality.
[0241] Ihe current technology of utilizing graphics processing units (GP U s) for Al spends 99% of the energy on training utilizing fully connected hidden layers. However, this process is both wasteful and does not accurately mimic the way physical brains process information.
[0242] The present invention changes the system of hidden layers to utilizing partial differential equations (PDEs) as a sort of pre-pre-training for the system, which helps in mimicking the nonlinearity seen in physical white mater and grey matter, thereby more effectively mimicking biological systems. The majority of the mapping in this case is from inputs to outputs with correlated mappings defined. The system utilizes PDEs that are both nonlinear and bounded to prevent chaotic dynamics, with boundary’ conditions keeping the model in localized trajectories and pilot functions used for initial training. The pilot function utilized by the present invention allows the system to ensure that it references the mapping equations, with the pilot function serving as the primary foundation for the system.
[0243] Tire use of nonlinearities for activation functions in the present invention provides for increased efficiency when combined with linear functions. The system does not require the use of washing or backpropagation to tweak weights and biases, which is a highly inefficient process. Limited pre-pre-training is utilized to create texture within the system, edges, contours, and to create nonlinear complexity. A pattern emerges over time from input to output with tire texture held constant, hi this way, the present system does not require training of hidden layers and instead uses mapping function relationships to avoid propagation of the weights and biases and deal with the volume of modifications.
[0244] The sy stem of the present invention is advantageous over prior art systems in various manners. For example, the present invention demonstrates a high degree of adaptability’, as the neuromorphic field is able to form complex patterns in response to different inputs or input types. This also allows for high versatility of the system, as the system design allowed it to be applied to various different domains, from robotics and loT devices to complex industrial systems, without significant changes in the architecture of the system. The system also exhibits a high degree of fault Tolerance, as the distributed nature of the neuromorphic field helps the system remain operational even if parts of it fail or are compromised, ensuring high reliability. The ability of the system to process and filter out noisy data further contributes to the robustness of the system, especially in environments where sensor data is imperfect or incomplete.
[0245] The system is naturally equipped to handle time-dependent data, which enables a continuous representation that allow’S for infinite-dimensional ACNs. Tire continuous nature of the ACN makes the system scalable, and capable of handling large amoun ts of data, w hich means that, as the complexity of tasks increases, the system scales the neuromorphic field to accommodate more data points and interactions without a significant drop in performance. The ability of the neuromorphic field to evolve based on incoming data allows the system to continuously improve its performance without the need for frequent manual updates. The systemof the present invention is able to learn and utilize spatial relationships in data, and offer context-aware processing, which enhances and evolves predictive analytic capabilities,
[0246] Tire PDE used in the present invention encapsulates complex dynamics with relatively few parameters to process complex inputs in a distributed and efficient manner, while the PDE Solver enables continuous evolution of the field, allowing the system to adapt to changing inputs without requiring extensive retraining, and while avoiding the need to train hidden layers. Thus, the PDE-based system allow s for fine-grained control over the system’s behavior, leading to more accurate and reliable outputs.
[0247] An additional advantage to the present invention includes the ability- to adapt to more specialized applications, including embodiments based on doped-photonic and hybrid photonic components that allow the system to process a w ider range of inputs, including light-based signals, and operate at higher processing and pow er efficiencies, enhancing versatility, and providing a better input match for biological, analog and related systems for high-precision, realtime diagnostics. The system also has the potential for even greater interpretability through the integration of Physics Informed Neural Network (PINNs) input and related support.
[0248] Neuromorphic Field
[0249] A neuromorphic field, as defined herein, is a continuous tensor-definable field able to be implemented using current means, as interconnected Artificial Cognitive Nodes’ activity in a continuous space represented by a field of partial differential equations (PDEs). The neuromorphic field, or dynamic grid, is designed to emulate the interconnectedness and adaptability found in natural systems where information is distributed and evolves overtime, and supports both traditional discrete ANNs as well as continuous signal propagation and data representations.
[0250] Artificial Cognitive Nodes (ACNs) interact with each other based on their inputs and both statically defined and dynamically sequenced states. These ACNs behave as a highdimensional, continuously propagating dynamic system, yvhere each node is governed by PDEs that represent and simulate the spatiotemporal dynamics of complex inputs (e.g., continuous analog, time sequenced biologically-driven, and / or digitally defined signals), thereby supporting a range of discrete statically, as well as dynamically, defined input types.
[0251] One goal of the system is to achieve a level of efficiency, dynamism and adaptability akin to biological systems, for the purpose of proactively adapting and optimizing sets of data within changing conditions and parameters to push the capabilities of artificial systems, by allowing the architecture to replace the explicit need for deep network training with a system supporting fully differentiable, bounded field propagation. Therefore, the present invention, in utilizing the neuromorphic field aims to improve current paradigms of predictive analytics,behavioral modeling, enriched contextual and spatial awareness, adaptive real-time responses, response optimization, enhanced operational capacities.
[0252] Tire neuromorphic field in this system is a complex, deterministic subsystem that remains unchanged throughout the training and operation of the network, which is a continuous, deterministically-bounded function over a spatial domain (Subfield; Resonant ACNs) and time. The neuromorphic field is able to be conceptualized as a vast, multidimensional field of interconnected ACNs, governed by a set of partial differential equations (PDEs) that define their behavior, interactions, and external influences.
[0253] Hie neuromorphic field is described by a system of reaction-diffusion PDEs, such with respect to equation 1 below, where Uj(x, t) is the activation of neuron type i at position x and time t, D ', is the diffusion coefficient for neuron type i, ft is a nonlinear function representing the interactions between neurons, and n is the number of different neuron types in the neuromorphic field.
[0254] (Equation 1) ∂u_i / ∂t = D_i∇²u_i + f(u_1,..., u_n, x, t)
[0255] The neuromorphic field includes multiple types of Artificial Cognitive Nodes. While the field is continuous, it exhibits localized regions of activity analogous to “Subfields,” which contain clustered resonant “nodes” that interact through spatial coupling terms in the governing PDEs. Each node has its own characteristics.
[0256] For example, excitatory' nodes tend to increase the activation of neighboring nodes. Excitatory nodes are responsible for propagating electrical signals to other nodes, which increases likelihood of the subsequent and neighboring nodes being activated, which overall represents and promotes action potential. Hie activation of nodes and propagation of signals across the network facilitates foiward transmission or information and activation patterns,
[0257] Inhibitory' nodes tend to decrease the activation of neighboring nodes. Inhibitory' nodes serve to control and refine signaling, which helps to prevent overexcitation and helps to maintain neural circuit stability by reducing activity of non-relevant nodes. This is helpful for feedback mechanisms, which prevent runaway reactions, ensure noise resistance, and support system stability.
[0258] Modulatory nodes, or dimensional weights, influence the strength and effectiveness of electrical signal transmission. Modulatory' nodes help to adjust the overall responsiveness of nodal circuit activity to inputs, and helps to simulate learning or other adaptation mechanisms by adjusting tire strength of connections based on reinforcement signals received. This mimics longterm potentiation, modulating responses to inputs over longer timescales. Modulatory' nodes interact through various mechanisms such as local interactions (i.e., nodes that influence immediate neighbors), learning (i.e., strength of connection based on correlation of pre- and post-nodal activation or deactivation), long-range connections (i.e., nodes influencing distant parts of the neuromorphic field), and electrical diffusion (i.e., neurotransmitter-like function of electrical signals that diffuse through the neuromorphic field, influencing node behavior). Analogies for the functioning of modulatory nodes include those grounded in photonics, such as “light that flows together, grows together” or general signals, such as “signals that sync together, link together.”
[0259] Other types of nodes include oscillatory’ nodes, which exhibit periodic behavior, creating rhythmic patterns in the neuromorphic field, and adaptive nodes, which change behavior based on the recent history of their activation.
[0260] Tire neuromorphic field has specific boundary' conditions that define its behavior at the edges. These boundary conditions are able to include periodic boundaries, where the neuromorphic field wraps around, creating a toroidal topology, reflective boundaries, where node activations are reflected at the edges, absorbing boundaries, where node activations decay at the edges, Robin boundaries, which are influenced by both the state of the neuromorphic field and its rate of change and which allows for a controlled interaction with the environment factoring into how the system interfaces with its surroundings, and, finally, mixed boundaries, applying different types of boundary’ conditions at different parts of the boundary’ or for different variables or equations within a multi -field or multi-modal system.
[0261] In one embodiment, the neuromorphic field is pretrained with a carefully well-defined pattern of activati on that ensures rich dynamics, allowing the neuromorphic field to exhibit complex, deterministic, non -chaotic behavior. The field also demonstrates stability such that the overall activity of the neuromorphic field remains bounded. The field also exhibits sensitivity, allowing small perturbations to lead to significant, but deterministic changes. Hie pretraining also provides for quick convergence, where the system reaches optimal performance more quickly, reducing the time and resources spent during the training phase, while both avoiding local minima (preventing the system from getting stuck in suboptimal local minima) and preventing divergence (preventing the system from exhibiting runaway behaviors where errors or activations escalate uncontrollably).
[0262] The behavior of the neuromorphic field is able to be analogized to ripples on a water surface, where information propagates as waves, interacts, and creates complex patterns, providing an intuitive understanding of the system's dynamics. To provide an analogy to help clarify the purpose of the neuromorphic field, it is metaphorically analogous on some level to propagating signals within a musical instrument, such as a drum, guitar, or violin or other instrument. When an input is introduced like one or more strings are plucked, it creates a complex pattern of energetic activations (signals or “signatures”) that propagate through theneuromorphic field. These signatures interact with each other and tire underlying dynamics of the neuromorphic field, creating a rich, time-varying pattern of activity.
[0263] For systems using light signals, the analogy is literal, as the photonic signals involve the propagation of light waves, which are able to be directed, reflected, and modulated within the system, much like water waves are reflected off barriers or pass through different mediums.
[0264] Input Layer
[0265] The input layer of the neural network included in the present invention is responsible for capturing and processing the incoming electrical signals, bio-signatures, photonic data, and other multi-modal types (enabling multi-modal processing and the integration of diverse input sources), then maps these signals into a format compatible with the neuromorphic field as described above, ensuring that the data is effectively processed by the PDE-based system. The input layer converts discrete and time-sampled input data selectively into continuous and / or discrete signals based on the pilot function, as described below, and injects these signals into the neuromorphic field. The signals are then able to be transformed into time vs. frequency -based inputs to be used as input to output mapped across the neuromorphic field.
[0266] Input-to-Output Mapping Function
[0267] Tire input-to-output mapping function is the core component of the system that exhibits learning. The mapping function takes the input signals and the sampled outputs from the neuromorphic field to produce the final output of the network.
[0268] In order to process the inputs, the mapping function is able to perform various tasks including input encoding, which converts raw input data into a format suitable for the neuromorphic field, spatial mapping, which determines how inputs are spatially distributed across the input nodes, enhancing data quality, which reduces noise especially with precision applications, temporal encoding for time series-data and encoding temporal aspects of the input, real-time data handling, input scaling, which normalizes input values to match the operating range of the neuromorphic field, and source integration, allowing data to come from a variety of sources such as sensors, digital feeds, user inputs, or databases. Integrating these diverse data streams in a cohesive manner is essential.
[0269] Tire mapping function induces such interactions with the neuromorphic field as input injection (i.e., applying the encoded input to the appropriate locations in the neuromorphic field) with multi-modal signal integration in case multiple inputs affect the field simultaneously, PDE solving (i.e., running the PDE Solver for a specified time interval) with the field evolving according to its governing PDEs and processing input information through the dynamics of the field, state sampling (i.e., sampling the state of the neuromorphic field at predetermined locations and times), learning and adaptive responses, where feedback from the output and the real-worldeffects of the system’s actions are used to fine-tune the field’s response to similar future inputs, enhancing the system’s learning and adaptation capabilities, signal propagation (i.e., inputs causing signals to propagate throughout the field, like waves spreading across a surface), diffusion mechanisms (i.e., to help in spreading the influence of inputs across a wider area, optimizing effectiveness of tasks that require a global perspective from local inputs), and nodal sy nchronization across different parts of the field, leading to coherent behaviors in response to inputs, mimicking how biological brains achieve functional synchronization between distant regions.
[0270] Tire mapping function utilizes feature extraction to process the sampled states to extract relevant features and functional mapping of the high-dimensional data from the field to a lower-dimensional output space that accurately reflects the results of computations, as well as increases fidelity of relevance to the task at hand. The mapping function is able to take advantage of temporal integration to combine samples from different time points.
[0271] Processes such as dimensionality reduction, nonlinear transformation, data synthesis (e.g., summarizing data trends and anomalies into a diagnostic report), and output scaling further assist in the mapping function generating output based on the input data. These techniques are able to be optimized for speed, with techniques streamlining the data pipeline, optimizing algorithms for speed, and / or utilizing hardware acceleration to enhance real-time responsiveness.
[0272] The mapping function can be represented according to equation 2 below, where x is the input, P(x) is the result of applying x to the neuromorphic field, S is the sampling function, and M is the learnable mapping function, and where M is able to be implemented as a neural network, a polynomial function, or any other suitable approximator.
[0273] (Equation 2) y = M(x, S(P(x)))
[0274] In order to train the mapping function, in a forward pass, an input is applied to the neuromorphic field, the neuromorphic field state is sampled, and the output is computed using the current mapping function. Differences between the produced output and the desired output are then calculated. Backpropagation is used, wherein gradients of the loss are computed with respect to the parameters of the mapping function, and then optimization techniques (e.g., Adam, RMSprop, etc.) are used to minimize the loss function. These optimization techniques update parameters of the mapping function and adjust weights and biases in the model to minimize error. This training process is able to iterate for a plurality of epochs over an entire training dataset.
[0275] Advantages
[0276] The mapping function of the present invention allows for several advantages of the present invention. First, it provides for separation of concerns, with complex dynamics handled by the fixed neuromorphic field, and learning focusing on the mapping function. The system alsoimproves noise handling by using the neuromorphic field to filter out irrelevant or spurious data to focus on relevant patterns, and efficiency, with the neuromorphic field providing complex transformations.
[0277] Hie mapping function is more easily able to be interpreted independently of the neuromorphic field dynamics, allowing for improved transparency and is more adaptable to new tasks, as only the mapping function needs to be retrained, as opposed to the neuromorphic field. The mapping function is able to continuously learn and improve over time based on new data and experiences, without needing explicit reprogramming.
[0278] Additionally, the system provides for improved modeling and predicting future states based on historical data is enhanced by the system's dynamic learning and adaptability, usefill in predictive maintenance, and other predictive analytics applications.
[0279] This enhanced system leverages the complex, deterministic dynamics of the neuromorphic field to create a rich representation of the input, while concentrating all learning in the input-to-output mapping function. This approach combines the power of PDEs for creating complex transformations with the flexibility and trainability of traditional machine learning models,
[0280] Further, entirely new functionality and modes of operation is able to be deployed, enabled, upgraded on the fly by switching between different mapping functions, which are leveraged via enabling different Pilot functions as defined herein, or other such means,
[0281] Output Layer
[0282] The output layer of the neural network of the present invention serves as the interface between tire processed data within the system and the external environment where this data is utilized. The output layer is responsible for converting the mapping function logits of the inputs across the neuromorphic field into accessible, practical, understandable, useful, and actionable outputs as needed for a given training, inference and related purposes. The output layer translates complex neural computations into intelligence and insight for taking autonomous actions based on input to output mapping of the PDE driven computations. Tire output layer also includes feedback mechanisms where the outputs are used to refine and retrain the system such that there it supports Pilot function enabled modes as described herein.
[0283] Tire present invention is able to sample the input to output mapping across the neuromorphic field at specific points in time vs frequency to produce the output. Insights are able to be pooled together from relevant “subfields” (resonant signatures of ACNs) found within the neuromorphic field. For clarity, the neuromorphic field represents the dynamic system’s continuous representation and ever-evolving interpretations of responses to inputs, so the data process is highly dimensional and interconnected. Therefore, the output layer works to synthesizethis data into coherent subfields as a function of the inputs (relevant resonant ACNs) for discrete outputs. For example, in the application of diagnostic, preventative or prevision healthcare, the system is able to interpret the input to output mapping functions across the field of various physiological data streams to output comprehensive health status, precision diagnostics, or precision treatments in a highly energy and processing-efficient manner.
[0284] The output layer is further able to apply a transformation to match the desired output format and time-sampling function. This allows the neural netw ork of the present invention to translate complex patters of mapped nodal activity into digital signals, text, commands, or visual outputs, depending on the application. Tire transformation applied ensures compatibility and usability of the data produced by the input to output mapping across the neuromorphic field via correctly matching outputs to interface with other digital systems, databases, format or related requirements.
[0285] PDE Solver
[0286] The PDE Solver helps to compute the evolution of the neuromorphic field at any point in time, as well as overtime as determined via the Pilot function. The accuracy and timeliness of the computations of the PDE Solver directly impact the effectiveness of the spatiotemporal dynamic insights of the subsystem, which is the result of evaluating interactions between the input and output over time. The PDE Solver does this by numerically solving the governing PDE of the neuromorphic field. In addition, the PDE Solver is responsible for defining the required multi-dimensional manifold mapping as a function of the inputs to the outputs. The PDE Solver updates the mapped input-output state of the neuromorphic field at discrete points in time, as w ell as providing function of time (F(t)) transform mapping.
[0287] Training Module
[0288] The training module of the present invention optimizes the system's input to output mapping parameters to improve performance on given tasks. The system of the present invention uses adapted Pilot Function -defined static and continuous fields and their representative PDEs, allowing the system to learn both the field dynamics and the input / output transformations, which enables the system to handle the complexities of training a system. To avoid introducing instability into the learning process, the system of the present invention employs well-defined regularization, scaled dot product attenuation, and gating structures to ensure stability of the fields mapping functions. The system optimizes the input-to-output mapping function based on machine learning methods as defined herein through one of many possible machine learning network configurations, and adjusts the PDE hyperparameters and sampling points based on the training. In one embodiment, hyperparameters are chosen for a given PDE-driven field based on pre-training such that the field presents a rich set of deterministic mapping of inputs to outputs.
[0289] Mathematical Formulation
[0290] Hie neuromorphic field is described by a reaction -diffusion PDE, as shown below as Equation 3, where u is a function representing the ACN activation and represents the state of the neuromorphic field at position x and time t, D is the diffusion coefficient and reflects signal propagation, speed, and interaction strength, V2indicates spatial diffusion (meaning DV2u models the spread or dispersion of the state across a function of time), f is a nonlinear function representing ACN dynamics (representing local interactions and behavior of inputs-to-outputs vs. subfields), meaning f(u, x, t) nonlinear interactions with the neuromorphic field depending on the state, position, and time that later translates into actionable intelligence, and 0 are learnable parameters:
[0291] (Equation 3) cti / ct = DVhi + f(u, x, t, 0)
[0292] By solving above equation, the system is able to account for time-dependent changes in the field, enabling real-time adaptation. Furthermore, Equation 3 helps to recognize complex patterns and simulate future scenarios.
[0293] Training Process
[0294] The training process for the neural network of the present invention begins by initializing the PDE parameters and sampling points randomly to optimize tire field’s responsiveness to new inputs. For each training example, the system converts the inputs to continuous signals, injects the signals into the neuromorphic field, solves the PDE for a fixed time interval, samples the neuromorphic field to produce one or more outputs, compares the one or more outputs with one or more desired results, computes the gradient of the loss with respect to PDE parameters and sampling points, and updates the parameters using an optimization algorithm (e.g., Adam, AdaGrad, etc.). This process is able to be repeated until convergence or for a fixed number of epochs, such that the training process is iterative and includes feedback mechanisms that ensure continuous improvement of the system’s overall performance, while accounting for the continuous nature and evolution of the field.
[0295] Operational Process
[0296] The process employed by the system of the present invention begins with receiving input data and preprocessing the data to normalize and pipeline the data into the system. The input is then converted to continuous signals, which are injected into the neuromorphic field. The system solves the PDE (e.g., Equation 3 above) for a fixed time interval or until equilibrium achieved as appropriate to application. Tire neuromoi'phic field is then sampled at learned sampling points, and as defined by the Pilot Function, after which an output transformation is applied to produce a final result.
[0297] Addressing Challenges and Considerations
[0298] Solving PDE's numerically is often computationally intensive, but much less so than typical Deep Neural Networks and associated Fully Connected and related hidden layers.Ensuring numerical stability of the PDE solver is crucial, which is mitigated through the incorporation of well-defined regularization, scaled dot product attenuation, and gating structures as described herein. The dynamics of the neuromorphic field are sometimes difficult to interpret, represent, and visualize, which is why the present invention employs Pilot Function-defined input-output mapping functions rather than trying to explicitly solve for the PDE's comprising tire Neuromorphic Field.
[0299] Potential Applications
[0300] Tire system of the present invention has application with regard to numerous different fields as well as having different applications with those fields. Examples of uses of the present invention include time series prediction, multi-scale analysis and forecasting of complex systems, music recommendation systems based on real-time user’s mood, location, and activity, dynamic soundscapes, predictive health monitoring, personalized medicine, personalized learning paths, performance prediction, predictive maintenance, spatial-temporal data analysis, continuous control and planning in high-dimensional state spaces with application in robotics, interactive concert experiences that adjust lighting, sound, and visuals to adapt to environment responding to audience movement patterns or sonic / bio-signals (frequency and energy levels), patient movement and behavior analysis, epidemiological tracking, pollution tracking and climate change analysis, control systems for continuous processes, automated music production, robotic surgery, continuous glucose monitoring systems, market conditions and risk management, automated effort-grading learning system adjusting levels of exercises or lessons based on realtime performance data, static solutions, image classification, speech recognition, machine translation, sentiment analysis, handwriting recognition, anomaly detection, time series forecasting, recommendation systems, playing board games, facial recognition, natural language processing with context-dependent relationships, Al songwriting assistants, voice -activated music systems, customer behavior analysis, smart cities, patient interaction system, medical documentation automation, Al tutors, personalized assistants and adaptive learning platforms, language learning systems, computer vision tasks involving motion or temporal evolution, music video generation, live motion capture for performance taking real-time movements and synchronization with visual effects or music to increase engagement, advanced analysis and interpretation of multi-modal medical data, Al-powered diagnostics, rehabilitation monitoring, and / or other applications.
[0301] This system represents a novel approach to neural network design, leveraging the power of PDEs to create a flexible, adaptive architecture that bypasses the major training andinference inefficiencies of Deep Learning Neural Networks and their huge number of parameters that must be trained. In addition to providing three or more orders of magnitude of time, power and rel ted efficiency gains, enabled by bypassing the majority of overhead inherent to DL / NN training, the system offers unique advantages in handling of discrete, continuous and spatial-temporal data.
[0302] Exemplary System
[0303] The subsystem includes a neuromorphic field evolving in space and time, coupled with input encoding and output decoding functions. This structure allows for processing of diverse datatypes within a unified continuous framework:
[0304] Consider a large, recurrent neural network that serves as a neuromorphic field of neurons as referenced herein. This network is designed to mimic the complex, non-linear behavior of a physical material. The neurons in this network are interconnected in a way that allow s for rich, dynamic signal propagation as further described.
[0305] Multiple input sources introduce signals into this neuromorphic field. These inputs are able to be concurrent (activating different parts of the network simultaneously) or sequential (occurring in a specific order over time). Each input creates a unique pattern of activations that propagates through the network, interacting with existing patterns in complex, but deterministically defined, w ays meaning that given the same set or sequence of input as a function of time, the same outputs will be defined based on bounded constraints as described.
[0306] At various points in the network, the network includes output nodes. These nodes sample the state of the network at specific times or over specific time intervals. Die activations of these output nodes form raw output data. To interpret this raw7output data and produce meaningful results, the system employs one or more mapping functions, which model different discrete or time-sequenced behaviors. Diese functions take as input both the original input signals and the sampled output data and then produce the final output of the system.
[0307] Die mapping functions are the only trainable part of this system, They have adjustable hyperparameters that control how the functions interpret the relationship between inputs and outputs. In one embodiment, the hyperparameters include, for example, the complexity of the function, the time scale, and / or the weight given to different types of patterns in the network.
[0308] Finally, while the currently defined preferred embodiment is designed to support existing Al means of representing the system, such as software-enabled machine learning using traditional silicon-based computers, the nature of the system is both modularly and systematically defined such that it is also able to be implemented via other organic, photonic or other means that provide for the functional input, output and nonlinear and related subsystem propagation.
[0309] Let the state of the neuromorphic field be described by a vector function u(t) G RN, where N is tire number of neurons in the field, T is a time constant, W 6 RNxNis the recurrent weight matrix, Win E RNxis the input weight matrix, gif t) G R is the i-th input signal, and b G RNis a bias vector. The evolution of u is governed by a non-linear differential equation (Equation 4 shown below).
[0310] (Equation 4) du / dt = -U / T + tanh(Wu + WinVigi(t) + b)
[0311] Each input i is represented by a time-dependent function g-(t). For a point input at time ti, the function is able to be described as shown below in Equation 5, where A; is the amplitude of the input, and 5 is the Dirac delta function.
[0312] (Equation 5) gj(t) = A;5(t - ti)
[0313] Similarly, the output is able to be represented by a time-dependent function. Let {yj(t)}j «iKbe a set of K output nodes. Each node samples the neuromorphic field state as shown in equation 6 below, where WouL jE RlxNis the output weight vector for the jthoutput node, and T]i(t) represents measurement noise.
[0314] (Equation 6) y,(t) = Wout;u(t) + ifyt)
[0315] Die mapping function M in this embodiment takes as input the original input signals {gi}i iMand the sampled outputs {yj}j iK, and produces the final output z, according to Equation 7 shown below, where 0 represents the hyperparameters of the mapping function,
[0316] (Equation 7) z = M({gi];=iM, {yj[j=iK; 0)
[0317] For a time-dependent mapping, a variation of Equation 7, designated as Equation 8 below, is able to be used.
[0318] (Equation 8) z(t) = M({gi(r)}i iM, {yj(r)}j-iK, T G [0, t]; 0)
[0319] In this one preferred embodiment, the hyperparameters 0 are able to include K (a parameter controlling the complexity of M (e.g., degree of a polynomial, number of layers in a neural network)), T (A time scale parameter determining how7far back in time M considers), and {w7i};-iM, {vj}j-iK(Weight parameters for different inputs and outputs),
[0320] Die training process involves finding the optimal hyperparameters 0* that minimize some loss function L according to Equation 9 below7, where ztargetis the desired output for a given set of inputs.
[0321] (Equation 9) 0* = argmi.nf) (z, ztarget)
[0322] Die neuromorphic field is initialized with random or other weights that satisfy certain conditions. In one embodiment, conditions to be satisfied include, but are not limited to spectral radius (i.e., the largest absolute eigenvalue of W is set to a value slightly less than 1 to ensure stable dynamics), sparsity, where W is typically sparse, with only a small fraction of possibleconnections realized, and input scaling, where win is scaled to control the strength of input perturbations.
[0323] Thi s initialization ensures that the neuromorphic field exhibits rich, non-linear dynamics without becoming chaotic or unstable. This mathematical framework provides a rigorous description of a system where a complex neuromorphic field transforms inputs, and trainable mapping functions interpret these transformations to produce meaningful outputs. The separation of the fixed, non-linear dynamics of the neuromorphic field from the trainable mapping function allows for a flexible system able to adapt to various tasks without changing its underlying network structure,
[0324] FIG. 37 illustrates a representation of a mapping function for a neural network, including training according to one embodiment of the present invention. The process flow in FIG. 37 goes from left to right, w ith an input being provided to an eigenfield and the eigenfield providing information to a mapping function, which then provides an output. In one embodiment, the eigenfield is a neuromorphic substrate. Advantageously, the mapping function and training process of the present invention require less power and is more efficient than traditional neural network mapping and training. Mapping is operable to include processes, including synaptic pruning, to eliminate extra synapses and mapping between hidden layers of a neural network, including forward mapping and backward mapping. In one embodiment, the mapping function and training process include a defined set of partial differential equations and pre-pre-traming as described herein.
[0325] FIG. 38 illustrates a schematic diagram of a system of the present invention. Multiple inputs are operable to be received by the eigenfield or a neuromorphic substrate, and the eigenfield or neuromorphic substrate is operable to create multiple outputs. In one non-limiting example, the eigenfield or neuromorphic substrate is operable to receive three inputs and produce two outputs. The inputs are operable to include a plurality of different types or formats of inputs and the outputs are operable to include a plurality of different types or formats of outputs.
[0326] Example
[0327] In one example, the present invention is able to be used for image recognition, particularly for recognition of hand-drawn numerals utilizing a hybrid system that combines elements of convolutional neural networks w ith the neuromorphic field and associated pre-biased and trained PDEs that provide rich, deterministic nonlinearities to support input-to-output mapped training.
[0328] The primary' goal of the invention in this case is to bypass the need for traditional deep learning related training of hidden lay ers in neural networks, thereby offering significantly faster training times and more efficient computation, as well as other advantages.
[0329] Ihe key innovation lies in replacing the hidden layers of a traditional neural networks with a neuromorphic field governed by PDEs, analogous to the human brains basis formed early on in life, but with key differences and resulting utilities. This approach leverages the power of both neural networks and differential equations to create a more flexible and potentially more interpretable model for function approximation and representation tasks, thereby enabling this neural network-based system to serve as a universal function approximator, while bypassing the need for time, cost, and energy inefficient training and optimization of hidden layers inherent to traditional deep learning neural network architectures.
[0330] As shown in FIG. 39, the input layers in this example are derived from the initial layers of a CNN and are responsible for initial feature extraction from input images of handwritten digits. The neuromorphic field replaces the hidden layers of a traditional CNN and utilizes PDEs and related architecture of the neuromorphic field to process and transform the extracted features. Tire output layers are then derived from the final layers of a CNN, producing final classification probabilities.
[0331] In the input processing stage, the hand-drawn numeral image is prepared for feature extraction. Tills hand-drawn numeral image (e.g,, 28x28 pixels in greyscale) is passed through initial convolutional and pooling layers of a CNN which are configured using standard techniques. Feature maps are extracted from the input layers, representing low-level features such as edges, comers, and textures in the image. Therefore, the input layer transforms raw pixel data into a meaningful set of features that is able to then be transformed by the neuromorphic field and differentially mapped to an output across the field, leveraging proven effectiveness of CNNs in extracting relevant features while handing off primitive elements to the neuromorphic field.
[0332] Tire neuromorphic field stage represents a core portion of this particular invention. The neuromorphic field is pre-defined and pre-trained to induce deterministically bounded nonlinearities across the field using a set of PDEs. This creates sets of linear and nonlinear related PDEs that support input to output transformation across the field, such that the inputoutput mapping is defined by the training using the handwritten numeral data. The PDEs employed represent archetypical patterns designed to provide a rich set of deterministically bounded transformations for the input data. The neuromorphic field is used to map the extracted features from the input processing stage across the field, involving transforming the feature maps into transformed elements. Nonlinear transformations are applied and solved across the PDEs to process the information based on standards sets of differentially mapped transform equations, which involves numerical integration of the PDEs over a specified time interval. A diagram of this process is shown in FIG. 40.
[0333] In this way, the neuromorphic field replaces the hidden layers of a traditional deep learning neural network, while the PDEs introduce a continuous physics-inspired model able to effectively capture complex patterns and relationships in the data for increased flexibility, efficiency, and interpretability compared to hidden layer approaches.
[0334] The output generation stage produces the final classification result, mapping the processed information from the neuromorphic field to the input of the CNN’s output layers. The output layers of the CNN are used to generate the final classification probabilities for numerical digits (e.g., 0-9). In one embodiment, the layers involve one or more fully connected layers followed by a softmax activation,
[0335] Tire mapping optimization training process in this case differs from traditional neural network training as, instead of using backpropagation through multiple hidden layers, the PDEs and neuromorphic field parameters are adjusted based on error between predicted and actual outputs. This adjustment process involves techniques from both machine learning and numerical analysis of PDEs and is iterated until the desired accuracy is achieved or until a maximum number of iterations is reached. Uris process is shown in the diagram of FIG. 41.
[0336] By bypassing traditional hidden layer weight updates, the system is able to achieve faster training times, especially compared to other deep networks. Tire system also has improved interpretability due to the use of PDEs, as the PDEs have clear physical and mathematical meanings and are fully differentiable and deterministically bounded. Further, by combining the strengths of CNNs in feature extraction and the expressive power of PDEs, the system is more flexible and able to capture more complex patterns in the data. Finally, the continuous nature of the neuromorphic field offers improved scalability to high-dimensional data compared to discrete neural network layers.
[0337] FIG. 42 is a schematic diagram of an embodiment of the invention illustrating a computer system, generally described as 800, having a netw ork 810, a plurality of computing devices 820, 830, 840, a server 850, and a database 870.
[0338] The server 850 is constructed, configured, and coupled to enable communication over a network 810 with a plurality of computing devices 820, 830, 840. Die server 850 includes a processing unit 851 with an operating system 852. The operating system 852 enables the server 850 to communicate through network 810 with the remote, distributed user devices. Database 870 is operable to house an operating system 872, memory 874, and programs 876.
[0339] hi one embodiment of the invention, the system 800 includes a network 810 for distributed communication via a wireless communication antenna 812 and processing by at least one mobile communication computing device 830. Alternatively, wireless and wired communication and connectivity between devices and components described herein includewireless network communication such as WI-FI, WORLDWIDE INTERO PERAB ILITY FOR MICROWAVE ACCESS (WIMAX), Radio Frequency (RF) communication including RF identification (RFID), NEAR FIELD COMMUNICATION (NFC), BLUETOOTH including BLUETOOTH LOW ENERGY (BLE), ZIGBEE, Infrared (IR) communication, cellular communication, satellite communication, Universal Serial Bus (USB), Ethernet communications, communication via fiber-optic cables, coaxial cables, twisted pair cables, and / or any other type of wireless or wired communication. In another embodiment of the invention, the system 800 is a virtualized computing system capable of executing any or all aspects of software and / or application components presented herein on the computing devices 820, 830, 840, In certain aspects, the computer system 800 is operable to be implemented using hardware or a combination of software and hardware, either in a dedicated computing device, or integrated into another entity, or distributed across multiple entities or computing devices.
[0340] By way of example, and not limi tation, the computing devices 820, 830, 840 are intended to represent various forms of electronic devices including at least a processor and a memory, such as a server, blade server, mainframe, mobile phone, personal digital assistant (PDA), smartphone, desktop computer, netbook computer, tablet computer, workstation, laptop, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be exemplary only, and are not meant to limit implementations of the invention described and / or claimed in the present application.
[0341] In one embodiment, the computing device 820 includes components such as a processor 860, a system memory 862 having a random access memory (RAM) 864 and a readonly memory (ROM) 866, and a system bus 868 that couples the memory 862 to the processor 860. In another embodiment, the computing device 830 is operable to additionally include components such as a storage device 890 for storing the operating system 892 and one or more application programs 894, a network interface unit 896, and / or an input / output controller 898. Each of the components is operable to be coupled to each other through at least one bus 868. The input / output controller 898 is operable to receive and process input from, or provide output to, a number of other devices 899, including, but not limited to, alphanumeric input devices, mice, electronic styluses, display units, touch screens, gaming controllers, joy sticks, touch pads, signal generation devices (e.g., speakers), augmented reality / virtual reality (AR / VR) devices (e.g., AR / VR headsets), or printers.
[0342] By way of example, and not limitation, the processor 860 is operable to be a general-purpose microprocessor (e.g., a central processing unit (CPU)), a graphics processing unit (GPU), a microcontroller, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), a Programmable Logic Device(PLD), a controller, a state machine, gated or transistor logic, discrete hardware components, or any other suitable entity' or combinations thereof that can perform calculations, process instructions for execution, and / or other manipulations of information.
[0343] In another implementation, shown as 840 in FIG. 42, multiple processors 860 and / or multiple buses 868 are operable to be used, as appropriate, along with multiple memories 862 of multiple types (e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core).
[0344] Also, multiple computing devices are operable to be connected, with each device providing portions of the necessary’ operations (e.g., a server bank, a group of blade servers, or a multi -processor system). Alternatively, some steps or methods are operable to be performed by circuitry that is specific to a given function.
[0345] According to various embodiments, the computer system 800 is operable to operate in a networked environment using logical connections to local and / or remote computing devices 820, 830, 840 through a network 810. A computing device 830 is operable to connect to a network 810 through a network interface unit 896 connected to a bus 868. Computing devices are operable to communicate communication media through wired networks, direct-wired connections or wirelessly, such as acoustic, RF, or infrared, through an antenna 897 in communication with the network antenna 812 and the network interface unit 896, which are operable to include digital signal processing circuitry' when necessary’. The network interface unit 896 is operable to provide for communications under various modes or protocols.
[0346] In one or more exemplary aspects, the instructions are operable to be implemented in hardware, software, firmware, or any combinations thereof. A computer readable medium is operable to provide volatile or non-volatile storage for one or more sets of instructions, such as operating sy stems, data structures, program modules, applications, or other data embodying any one or more of the methodologies or functions described herein, lire computer readable medium is operable to include the memory’ 862, the processor 860, and / or the storage media 890 and is operable to be a single medium or multiple media (e.g., a centralized or distributed computer system) that store the one or more sets of instructions 900. N on-transitory computer readable media includes all computer readable media, with the sole exception being a transitory, propagating signal per se. The instructions 900 are further operable to be transmitted or received over the network 810 via the network interface unit 896 as communication media, which is operable to include a modulated data signal such as a carrier wave or other transport mechanism and includes any delivery' media. The term “modulated data signal” means a signal that has one or more of its characteristics changed or set in a manner as to encode information in the signal.
[0347] Storage devices 890 and memoty 862 include, but are not limited to, volatile and nonvolatile media such as cache, RAM, ROM, EPROM, EEPROM, FLASH memory, or other solid state memory technology; discs (e.g., digital versatile discs (DVD), HD-DVD, BLU-RAY, compact disc (CD), or CD-ROM) or other optical storage; magnetic cassettes, magnetic tape, magnetic disk storage, floppy disks, or other magnetic storage devices; or any other medium that can be used to store the computer readable instructions and which can be accessed by the computer system 800.
[0348] In one embodiment, the computer system 800 is within a cloud-based network. In one embodiment, the server 850 is a designated physical server for distributed computing devices 820, 830, and 840. In one embodiment, the server 850 is a cloud-based server platform. In one embodiment, the cloud-based server platform hosts serverless functions for distributed computing devices 820, 830, and 840.
[0349] In another embodiment, the computer system 800 is within an edge computing network. The server 850 is an edge server, and the database 870 is an edge database. The edge server 850 and the edge database 870 are part of an edge computing platform. In one embodiment, the edge server 850 and the edge database 870 are designated to distributed computing devices 820, 830, and 840. In one embodiment, the edge server 850 and the edge database 870 are not designated for distributed computing devices 820, 830, and 840. The distributed computing devices 820, 830, and 840 connect to an edge server in the edge computing network based on proximity, availability, latency, bandwidth, and / or other factors.
[0350] It is also contemplated that the computer system 800 is operable to not include all of the components shown in FIG. 42, is operable to include other components that are not explicitly shown in FIG. 42, or is operable to utilize an architecture completely different than that shown in FIG. 42. The various illustrative logical blocks, modules, elements, circuits, and algorithms described in connection with the embodiments disclosed herein are operable to be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application (e.g., arranged in a different order or partitioned in a different way), but such implementation decisions should not be interpreted as causing a departure from the scope of the present invention.
[0351] Certain modifications and improvements will occur to those skilled in the art upon a reading of the foregoing description. The above-mentioned examples are provided to serve thepurpose of clarifying the aspects of the invention and it will be apparent to one skilled in the art that they do not serve to limit the scope of the invention. All modifications and improvements have been deleted herein for the sake of conciseness and readability but are properly within the scope of the present invention.7:
Claims
1. CLAIMS2.The invention claimed is:
1. A neural network, comprising:4.an input layer;5.an output layer; and6.a neuromorphic field connecting the input layer to the output layer;7.wherein the neuromorphic field includes a plurality of interconnected artificial neurons whose behaviors are defined and governed by one or more partial differential equations (PDEs).
2. The neural network of claim 1, wherein the output layer is operable to apply a transformation to achieve a desired output format.
3. The neural network of claim 1, wherein the one or more PDEs include a reaction-diffusion PDE4. The neural network of claim 1, wherein a PDE Solver computes evolution of the neuromorphic field over time as determined via a Pilot signal.
5. The neural network of claim 1, wherein the neural network is trained using well-defined regularization, scaled dot product attenuation, and gating structures.
6. The neural network of claim 1, wherein the input layer receives input data, normalizes the input data, converts the input data into continuous signals, and injects the continuous signals into the neuromorphic field.
7. The neural network of claim 1, wherein the plurality of interconnected artificial neurons includes one or more excitatory nodes, one or more inhibitory nodes, one or more oscillatory nodes, one or more adaptive nodes, and / or one or more modulatory nodes,8. The neural network of claim 1, wherein the neuromorphic field includes one or more boundary conditions, including one or more periodic boundaries, one or more reflective boundaries, one or more Robin boundaries, and / or one or more absorbing boundaries.
9. The neural network of claim 1, wherein a training process of the neural network includes backpropagation.
10. The neural network of claim 1, wherein the neural network is utilized for audio enhancement and / or audio restoration.
11. The neural network of claim 1, further comprising an input-to-output mapping function including at least one adjustable PDE hyperparameter.
12. The neural network of claim 11, wherein the input-to-output mapping function is operable to perform synaptic pruning.
13. The neural network of claim 1, wherein the neuromorphic field is operable to be sampled at learned sampling points as defined by a Pilot function.
14. A neural network, comprising:20.an input layer;21.an output layer;22.an input-to-output mapping function; and23.a neuromorphic field connecting the input layer to the output layer;24.wherein the neuromorphic field is operable to form patterns in response to different inputs and / or input types;25.wherein the input-to-output mapping function includes at least one adjustable partial differential equation (PDE) hyperparameter;26.wherein the at least one adjustable PDE hyperparameter represents complexity of the input-to-output mapping function, time scale and / or pattern weight;27.wherein the neuromorphic field is defined by one or more PDEs having one or more boundary conditions; and28.wherein the one or more PDEs includes a reaction-diffusion PDE.
15. The neural network of claim 14, wherein the output layer is operable to apply a transformation to achieve a desired output format.
16. The neural network of claim 14, wherein a PDE Solver computes evolution of the neuromorphic field over time as determined via a Pilot function.
17. The neural network of claim 15, wherein the neuromorphic field is operable to be sampled at learned sampling points as defined by the Pilot function.
18. The neural network of claim 14, wherein the input-to-output mapping function is operable to perform synaptic pruning.
19. A method for enhancing the accuracy and relevance of generated responses, comprising: receiving a user query;34.performing graph-based retrieval from a knowledge graph;35.at least one large language model (LLM) generating a response to the user query; performing Bayesian evaluation of the response to the user query' and determining whether the response meets a predetermined quality threshold;36.performing secondary ground truth verification on the response;37.verifying the response with multiple artificial intelligence (Al) agents;38.wherein the multiple Al agents include a fact-checker, a coherence analyzer, a relevance assessor agent, and an ethical compliance agent;39.adjusting the response based on feedback from the multiple Al agents; and delivering the response to a user device.
20. The method of claim 19, further comprising a synthetic data generator providing feedback to the knowledge graph, to the at least one LLM, and / or to inform the Bayesian evaluation.
21. The method of claim 19, wherein the at least one LLM includes a plurality of LLMs.
22. The method of claim 19, wherein the predetermined quality threshold is based on coherence, relevance, and factual accuracy of the response.
23. The method of claim 19, wherein the multiple Al agents include at least three Al agents.
24. The method of claim 19, further comprising iteratively repeating one or more steps of the method.
25. A system for compressing media content, comprising:45.a media analysis module configured to perform multi-faceted analysis on input media; a manifold selection and optimization module;46.a deep learning model training module;47.an entropy maximization module;48.a compression application module; and49.an encoding module configured to package the compressed media into standard format containers.
26. The system of claim 25, wherein tlie multi-faceted analysis includes spectral analysis, statistical analysis, perceptual analysis, and temporal-spatial correlation analysis.
27. The system of claim 26, wherein the spectral analysis includes application of a Short-Time Fourier Transform (STFT) with overlapping windows for audio content, employment of a 3D Fourier Transform on groups of frames for video content; and implementation of a Wavelet Transform for multi-resolution analysis of both audio and video content,28. A method for compressing media content, comprising:53.performing multi-faceted analysis on input media, including spectral analysis, statistical analysis, perceptual analysis, and temporal-spatial correlation analysis; selecting and optimizing a dimensional manifold based on results of the multi-faceted analysis;54.training a deep learning model to map between an original media space and the selected dimensional manifold;55.applying entropy maximization techniques to a representation of the selected dimensional manifold;56.compressing the media content using the trained deep learning model and entropy- maximized manifold; and encoding the compressed media content into a standard format container while maintaining compatibility with existing media ecosystems.
29. The method of claim 28, wherein the spectral analysis comprises:58.applying Short-Time Fourier Transform (STFT) with overlapping windows for audio content;59.employing a 3D Fourier Transform on groups of frames for video content; and implementing a Wavelet Transform for multi -resolution analysis of both audio and video content.
30. The method of claim 28, wherein the perceptual analysis comprises:61.implementing psychoacoustic models based on critical bands and masking effects for audio content;62.applying visual saliency models to identify perceptually important regions in video content; and63.incorporating Just Noticeable Difference (JND) models to determine perceptual thresholds for different media components.
31. The method of claim 28, wherein selecting and optimizing the dimensional manifold comprises:65.estimating an intrinsic dimensionality of the input media using techniques including false nearest neighbors algorithm or maximum likelihood estimation;66.implementing manifold learning techniques including Isomap, Locally Linear Embedding (LLE), or t-SNE to learn a structure of the input media; and67.applying Riemannian optimization techniques to fine-tune manifold parameters.
32. The method of claim 28, wherein training the deep learning model comprises:69.designing an encoder-decoder architecture with attention mechanisms; incorporating residual connections and skip connections to facilitate gradient flow and preserve fine-grained details;70.implementing a multi-term loss function incorporating reconstruction error, perceptual loss, and manifold consistency terms; and71.applying curriculum learning, starting with simple patterns and gradually increasing complexity during the training.
33. The method of claim 28, wherein applying entropy maximization techniques comprises: computing Shannon entropy for each dimension or feature in the representation of the selected dimensional manifold;73.applying Independent Component Analysis (ICA) to separate statistically independent components; implementing the Principle of Maximum Entropy to optimize distribution of information across the selected dimensional manifold; and74.developing an adaptive quantization scheme that allocates more bits to high-entropy components.
34. The method of claim 28, wherein compressing the media content comprises:76.preprocessing the input media using adaptive noise reduction techniques;77.applying the trained deep learning model to transform the input media into an optimized manifold representation;78.implementing vector quantization for groups of related manifold coordinates; and applying context-adaptive coding schemes that exploit local patterns in the quantized media content.
35. The method of claim 28, wherein encoding the compressed media into a standard format container comprises:80.developing a flexible, hierarchical data structure to represent the compressed media content, manifold parameters, and model configuration;81.implementing versioning mechanisms to ensure forward and backward compatibility; and encapsulating the compressed media content within standard containers while ensuring compliance with specifications of chosen container formats.
36. A method for compressing media content, comprising:83.performing multi-faceted analysis on input media, including spectral analysis, statistical analysis, perceptual analysis, and temporal-spatial correlation analysis; selecting and optimizing a dimensional manifold based on results of the multi-faceted analysis;84.training a deep learning model to map between an original media space and the selected dimensional manifold;85.computing Shannon entropy for each dimension or feature in the representation of the selected dimensional manifold;86.applying Independent Component Analysis (ICA) to separate statistically independent components;87.implementing the Principle of Maximum Entropy to optimize distribution of information across the selected dimensional manifold;88.developing an adaptive quantization scheme that allocates more bits to high-entropy components; and89.compressing the media content using the trained deep learning model and entropy- maximized manifold.
37. The method of claim 36, wherein the spectral analysis comprises:90.applying Short-Time Fourier Transform (STFT) with overlapping windows for audio content;91.employing a 3D Fourier Transform on groups of frames for video content; and implementing a Wavelet Transform for multi -resolution analysis of both audio and video content.
38. The method of claim 36, wherein the perceptual analysis comprises:93.implementing psychoacoustic models based on critical bands and masking effects for audio content;94.applying visual saliency models to identify perceptually important regions in video content; and95.incorporating Just Noticeable Difference (JND) models to determine perceptual thresholds for different media components.
Citation Information
Patent Citations
Multiscale deep-learning method for training data-compression systems
EP4303774A1
Bayesian graph-based retrieval-augmented generation with synthetic feedback loop (BG-RAG-SFL)
US12437213B2
Platform for Gathering Real-Time Analysis
US20170099200A1
Bias scheme for single-device synaptic element
US20220180165A1
Graph neural diffusion
US20220253671A1