A computer-implemented method for generating a multimodal machine learning model, and related computer-implemented methods

By integrating audio-spatial representations with modality training data, the method addresses the spatial correspondence issue in multimodal learning, achieving accurate and efficient generation of synchronized audiovisual content.

WO2026049667A1PCT designated stage Publication Date: 2026-03-05MONAVA AB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/SE2025/050768
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-27
Filing Date
2025-08-27
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Current multimodal machine learning models fail to preserve precise spatial correspondence between audio and visual data due to separate processing of audio and visual streams, leading to loss of fine-grained spatial-frequency information and requiring extensive computational resources.

Method used

A method that integrates audio-spatial representations with modality training data to create a shared joint embedding space, enabling spatially-coherent multimodal content generation through cross-modal spatial constraints, without manual feature engineering or spatial annotation.

Benefits of technology

Enables accurate spatial predictions and generation of synchronized audiovisual content by learning direct spatial relationships between audio and visual data, reducing computational requirements and enhancing model coherence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SE2025050768_05032026_PF_FP_ABST
    Figure SE2025050768_05032026_PF_FP_ABST
Patent Text Reader

Abstract

A method for enhancing spatial understanding of machine learning models for at least increasing the generation capabilities of audio and image (video) models is described. The method includes collection of audio spatial data, preparing and processing of the audio spatial data, creating datasets with training inputs and outputs by some combination of the audio spatial data with other modalities, adding either encoder and-or decoder for audio spatial data to the model, training the machine learnings model using the dataset and optionally removing the audio spatial encoder and decoder. Some alternatives include the possibility to combine images with audio spatial frequency images or combine video with frequency audio-spatial temporal data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] A computer-implemented method for generating a multimodal machine learning model, and related computer-implemented methods

[0002] Field

[0003] The present inventive concept relates to a computer-implemented method for generating a multimodal machine learning model, and related computer-implemented methods.

[0004] Background Multimodal machine learning models process audio and visual data to generate synchronized content across different modalities. Current approaches typically treat audio as temporal waveforms or frequency spectrograms without preserving explicit spatial information about sound source locations within the scene. This results in generated audiovisual content that lacks precise spatial correspondence between sound sources and their visual counterparts.

[0005] Existing multimodal training methodologies rely on learning implicit spatial relationships from large datasets, requiring extensive computational resources and training data to achieve spatial coherence. These approaches process audio and visual streams separately before combining them at higher abstraction levels, potentially losing fine-grained spatial-frequency information present in the original acoustic scene.

[0006] Contemporary audio processing for machine learning applications focuses primarily on temporal and frequency characteristics while spatial attributes are either ignored or processed through separate localization algorithms that operate independently from the main multimodal training pipeline. This separation limits the model's ability to learn direct spa- tial correspondences between acoustic events and their visual manifestations during the training process.

[0007] It is therefore an objective to address the above short-comings.

[0008] Summary According to a first aspect of the disclosure, a computer-implemented method for generating a multimodal machine learning model comprises obtaining audio training data representing a spatial distribution of audio frequencies at multiple locations within a scene, obtaining modality training data associated with the scene, generating a training dataset by combining the audio training data with the modality training data, and training a machine learning model using the training dataset to generate a multimodal machine learning model.

[0009] The present inventive disclosure provides for a transformative approach to machine learning that fundamentally enhances how models understand and represent spatial rela- tionships by introducing frequency-specific audio-spatial representations that are crosstrained with other modalities, rather than merely processing audio and visual data as separate, unconnected streams. This approach maps explicit spatial audio and various modalities into a shared joint embedding space, thereby enabling generation of spatially-coherent multimodal content and accurate spatial predictions through learned cross-modal spatial constraints rather than post-hoc alignment techniques. It may be noted that the training process as disclosed herein may be performed automatically and continuously without requiring manual feature engineering, spatial annotation, or explicit correspondence labelling between modalities. For instance, the audio-spatial representations may be automatically generated through beamforming algorithms and integrated into the training pipeline without manual spatial calibration or alignment. Hence, there is enabled a selfsupervised spatial learning paradigm that operates independently of annotated spatial ground truth data, learning spatial relationships directly from the natural correspondence between audio propagation physics and visual scene geometry.

[0010] Obtaining audio training data representing a spatial distribution of audio frequencies pro- vides explicit spatial localization of frequency content across the scene without requiring manual annotation of sound source locations, thereby eliminating the ambiguity inherent in conventional audio processing that treats sound as a single-point or non-spatial signal, and enabling the model to learn the fundamental relationship between frequency characteristics and spatial position. Obtaining modality training data associated with the scene enables the system to establish correspondence between the audio-spatial representation and other perceptual modalities, providing multiple views of the same spatial reality that reinforce and validate the learned spatial relationships through cross-modal consistency constraints during training.

[0011] Generating a training dataset by combining the audio training data with the modality train- ing data creates aligned multimodal pairs that preserve both temporal and spatial cor- respondence, enabling the model to learn not just what sounds and looks correlate, but precisely where in space these correlations occur, thereby establishing a shared spatial coordinate system across modalities.

[0012] Training a machine learning model using the training dataset enables the transfer of spatial constraints from the audio domain to other modalities through gradient descent optimization, whereby the explicit spatial structure in the audio data acts as a teaching signal that guides the model to develop spatial ly-aware representations even in modalities that lack inherent spatial organization.

[0013] This dual-representation architecture combining frequency-specific spatial audio data with conventional modalities addresses the limitations of purely appearance-based or contextbased multimodal learning by providing an explicit spatial scaffold that grounds abstract features in physical space.

[0014] The term "spatial distribution of audio frequencies" may relate to a multi-dimensional representation where each spatial location within a scene is associated with a frequency spectrum or spectral features, creating a spatial map of frequency content that explicitly encodes both what sounds are present and where they originate, including but not limited to beamformed audio grids, acoustic heatmaps, or spatial spectrograms.

[0015] The term "modality training data" may relate to any modality that provides complementary information to the scene, including but not limited to visual imagery, video sequences, depth maps, thermal imaging, radar returns, text descriptions, or previously learned latent representations from other models. In some instances, the modality training data may relate to spatial audio.

[0016] The term "multimodal machine learning model" may relate to a neural network architecture capable of processing, understanding, and generating data across multiple modalities through shared representations, including but not limited to transformer-based models, joint embedding architectures, cross-attention networks, or diffusion models with multimodal conditioning. It may be said that the multimodal machine learning model may relate to a generative neural network.

[0017] The term "scene" may relate to any physical or virtual environment where audio and other modal signals can be captured or simulated, encompassing both the spatial extent and temporal duration of the recording or generation, including indoor spaces, outdoor environments, or synthetic scenarios.

[0018] The phrase "modality training data associated with the scene" encompasses data hav- ing temporal, spatial, semantic, or learned correspondence with the audio training data, where strict scene-level alignment is not required. Modality training data may be captured simultaneously from the same physical scene as the audio data, may share semantic content while originating from different scene(s), may belong to the same category or domain as represented in the audio data, or may be associated through learned relationships discovered during training.

[0019] The phrase "combining the audio training data with the modality training data" may relate to the process of creating unified training examples where audio-spatial representations and other modalities are brought together in a structured format that preserves their spatial and temporal correspondence, including but not limited to paired associations, concatenated features, or synchronized multi-stream inputs that enable the machine learning model to learn cross-modal relationships.

[0020] In this context, "combining" may refer to temporal synchronization where audio-spatial data and modality data from the same time instant are paired together, spatial registration where coordinate systems are aligned such that location (x, y) in the audio-spatial representation corresponds to the same physical location in other modalities, feature-level concatenation where tensors from different modalities are merged into multi-channel inputs, or joint embedding formation where separate encoders process each modality before combining their outputs in a shared latent space. This combining process preserves the inherent spatial structure of the audio data while establishing learnable relationships with other modalities, enabling the model to discover cross-modal spatial constraints through the natural correspondence between how sound propagates through space and how that space appears in other sensory modalities, thereby facilitating the transfer of explicit spatial knowledge from the audio domain to other modalities where spatial relationships might be implicit or ambiguous.

[0021] Optionally in some examples, the modality training data comprises audio data, visual data, text data, coordinate data, radar data, radio frequency, vibration data, depth maps, latent representations, or thermal data, which enhances the model's ability to learn spatial relationships from diverse data types. It may further be expressed that the modality training data enables cross-modal knowledge transfer where learned representations from one modality enhance performance in another modality, including non-spatial applications. For example, pairing acoustic recordings of a cat with generic images of cats enables the model to learn acoustic classification of cats through the visual domain, where well- established image classifiers can contribute similar classification capabilities in the acous- tic domain without requiring extensive acoustic training data. This cross-modal transfer allows models to leverage existing high-performance classifiers from data-rich modalities to improve performance in data-sparse modalities, extending beyond spatial relationships to include semantic and categorical knowledge transfer across different sensory domains.

[0022] Optionally in some examples, the audio training data comprises temporal variations of the audio frequencies at multiple locations, which allows the model to capture dynamic spatial changes over time.

[0023] In this context, temporal variations may refer to the evolution of frequency content at each spatial location across time, creating a four-dimensional representation (x, y, z, frequency, time) that captures both static and dynamic aspects of the acoustic scene, including mov- ing sound sources, changing acoustic conditions, or temporal patterns in spatial audio events.

[0024] Optionally in some examples, the audio training data is generated using a machine learning model configured to produce multi-channel audio data or spatial-frequency audio data, which provides richer spatial information for training. Optionally in some examples, obtaining audio training data comprises positioning sound sensor(s), simultaneously recording audio signal(s), defining spatial location(s), calculating time delays, applying beamforming processing, and generating frequency represen- tationss, which enables precise spatial localization of sound sources.

[0025] In this context, beamforming processing may refer to signal processing techniques that computationally focus the microphone array's sensitivity toward specific spatial location, including but not limited to delay-and-sum beamforming, adaptive beamforming algorithms, or neural beamforming networks that learn optimal spatial filtering strategies.

[0026] Optionally in some examples, applying beamforming processing comprises neural beamforming, delay-and-sum beamforming, adaptive beamforming, or MUSIC algorithm pro- cessing, which improves the accuracy and efficiency of spatial feature extraction. Applying beamforming processing may also comprise deconvolution beamforming.

[0027] Optionally in some examples, training the machine learning model is based on a transformer, a convolutional neural network, a recurrent neural network, a diffusion model, a generative adversarial network, a joint-embedding predictive architecture, a deep neural network, or a generative neural network, which provides flexibility in model architecture to optimize spatial understanding.

[0028] Optionally in some examples, training the machine learning model comprises adding an encoder unit to process audio training data, training the model with the encoder unit, and optionally removing the encoder unit.

[0029] An advantage of the removable encoder architecture is deployment flexibility. This reduces computational overhead and memory requirements at inference time.

[0030] In this context, the encoder unit may refer to a specialized neural network component designed to transform audio-spatial representations into latent embeddings compatible with the model's shared representation space, including but not limited to 3D convolutional networks, vision transformers adapted for spatial-frequency grids, or custom architectures that preserve spatial structure while reducing dimensionality.

[0031] Optionally in some examples, the audio training data is organized as a two-dimensional or a three-dimensional frequency-spatial representation, which provides a structured format for the model to learn complex spatial relationships.

[0032] Hence, the audio training data may be organized as a frequency-spatial representation having two or more dimensions.

[0033] In this context, two-dimensional representation may refer to a Frequency-Audio-Spatial (FAS) format where spatial locations are arranged in a 2D grid with frequency spectra at each point, analogous to a multi-channel image where channels represent frequency bins rather than color.

[0034] In this context, three-dimensional representation may refer to either a 3D spatial grid with frequency information or a 2D spatial grid with an additional time dimension (FAST), cre- ating volumetric data structures that can be processed by 3D convolutional networks or other volumetric architectures. For video applications, a four-dimensional representation (FASTT) may be employed where each video timestep contains a temporal buffer of audio-spatial data, wherein at a given video frame time the model accesses a limited historical window of spatial audio information rather than the entire temporal sequence. This FASTT representation enables the model to maintain temporal context while limiting computational complexity by constraining the temporal buffer size, allowing the model to process dynamic spatial audio changes at each video frame without requiring access to the complete temporal history.

[0035] According to a second aspect of the disclosure, a computer-implemented method for gen- erating modality output data comprises providing a multimodal machine learning model according to the above, providing modality input data to the model, and generating modality output data using the model. This method allows for the creation of diverse output data from a trained model. Optionally in some examples, the modality output data comprises any of audio data, visual data, text data, coordinate data, radar data, radio frequency data, vibration data, depth maps, latent representations, or thermal data. This enables the trained model to generate diverse output types across multiple modalities, allowing applications ranging from spa- tial audio generation and visual content creation to sensor data synthesis and coordinate prediction tasks.

[0036] Optionally in some examples, the modality input data comprises any of audio data, visual data, text data, coordinate data, radar data, radio frequency data, vibration data, depth maps, latent representations, or thermal data. This input flexibility allows the model to process heterogeneous data sources and perform cross-modal inference tasks, enabling applications such as generating spatial audio from visual inputs or predicting coordinates from textual descriptions.

[0037] Optionally in some examples, the step of generating modality output data comprises adding at least one decoder unit configured to generate the modality output data, and optionally removing the decoder unit after generating the output data. This modular decoder approach provides architectural flexibility for deployment optimization, enabling task-specific decoders to be temporarily added for particular generation tasks while maintaining a compact core model for general inference operations.

[0038] According to a third aspect of the disclosure, a computer-implemented method for classi- tying an audio source using spatial audio data comprises providing a multimodal machine learning model according to the above, providing input spatial audio data associated with an audio source to the model to generate a spatial audio representation, and classifying the audio source based on the spatial audio representation. This enables accurate identification of audio sources based on their spatial characteristics. It may alternatively be said that the method encompasses performing various machine learning tasks on the spatial audio representation, wherein the tasks may comprise classification, regression, clustering, dimensionality reduction, generation, reinforcement learning, or retrieval, thereby enabling diverse analytical and generative applications beyond audio source identification.

[0039] It should in this context be noted that advantages of each of the aspects of the present inventive concept are equally applicable to each other. For instance, the advantages associated with the first aspect regarding enhanced spatial understanding through audio- spatial representations are equally applicable to the second and third aspects.

[0040] It may also be said that the present inventive concept provides a system, process and method for creating and utilizing audio-spatially tuned models to at least achieve and learn a more comprehensive understanding of the relationships between at least one other modality and at least audio-spatial data. The proposed method involves acquiring and processing audio-spatial data, combining it with at least one modality such as visual and audio data to create, add encoders and decoders for in some embodiments FAST or at least audio-spatial data or in some embodiments a combined representation of at least audio-spatial data, and train a generative multimodal machine learning model and in some embodiments remove the added encoders and decoders, to then arrive at a model that can be inferenced with superior audio-spatial capabilities even when the audio-spatial either encoders and decoders are no longer present or when data has been stripped of audio-spatial dimensions.

[0041] 1. The method of using two-dimensional frequency based audio-spatial features to crosstrain with other modalities of neural networks to improve an audio-spatially improved model; wherein said method uses AS, FAS, FAST, FASTT or some other equivalent combination of such using generation, algorithms or combinations. 2. The process of using audio-spatial features to cross-train with other modalities of neural networks to derive an audio-spatially improved model; wherein said product was derived using AS, FAS, FAST, FASTT or some other equivalent combination of such using generation, algorithms or combinations.

[0042] 3. The process of using audio-spatial recording with said processing unit or physical ele- ment producing a recording, generated encoding or latent encoding that is applicable for cross-training neural networks, or derive the same data through generation, algorithms or combination of audio-spectral with another modalities into a third modality. a. Using beamforming b. Using other means such as MUSIC. 4. The method of using frequency based audio-spatial features with convolutional neural networks to classify objects spatially using acoustics. a. Where in item 4, it can be extended with

[0043] 5. The method of using two-dimensional frequency based audio-spatial features with convolutional neural networks to classify objects spatially using acoustics. 6. with other modalities of neural networks to improve an audio-spatially improved model; wherein said method uses AS, FAS, FAST, FASTT or some other equivalent combination of such using generation, algorithms or combinations. Brief Description of the Drawings

[0044] Figure 1. The method and system.

[0045] Figure 2. Illustration of how cross-learning occurs.

[0046] Figure 3. Example of combining different modalities dataset to align data. Figure 4. The joint embedding approach.

[0047] Figure 5. Recording using multiple micro-arrays and processing using beamforming highlighting how delays in the right part of the picture could be summed if imagining the left side shows how sound propagates with a delay.

[0048] Figure 6. Beamformed illustration of noise source and car source (x,y, power). Figure 7. Overlayed figured 6 on a real image. Car is referenced forth as 1, noise source as 2, bottom right corner is referenced as number 3.

[0049] Figure 8. Three Logmel Spectogram of three points in figure 6 and 7, first is car, second is noise source, and third is bottom right corner (Frequency, (3x) Time illustration).

[0050] Figure 9. The FAS representation, note: aggregated and contoured local maximum illus- tration of frequencies of one instance in time of figure 6 and 7. The figure is an Audio spectrum along every beamformed “pixel” point. (Frequency, x,y). Y is height, X is width as if they were images, but instead of colors, it is audio frequency. The color is relative magnitude in frequency band normalized on each frequency across x,y.

[0051] Figure 10. Mapping the relationship of the figure 7, 8, 9. Figure 11. The FAST representation. Note: Combining figure 8 and 9 is 4D representation (Frequency, Time, X,Y). FAST shows the historic data in space. FASTT shows FAST for each timestep.

[0052] Figure 12. Example of physical position of microphones.

[0053] Figure 13. Preparing FAS or FAST data. Figure 14. Simpler way of arriving at FAST by first deriving FAS.

[0054] Figure 15. Preparing FAS or FAST data.

[0055] Detailed Description The detailed description set forth below provides information and examples of the disclosed technology with sufficient detail to enable those skilled in the art to practice the disclosure.

[0056] Figure 1 shows a method and system for generating a multimodal machine learning model. The system includes components such as audio training data, modality training data, and a processing unit. The audio training data provides spatial information of audio signal, including location and audio features. The modality training data provides additional information, such as visual data, text data, coordinate data, video data, radar data, or thermal data. The processing unit trains the multimodal machine learning model using a training dataset that combines the audio training data and modality training data.

[0057] Figure 2 shows an illustration of how cross-learning occurs in the multimodal machine learning model. The model integrates audio training data and modality training data to learn spatial relationships from multiple perspectives. The training inputs combine these datasets to allow the model to extract and understand spatial features, improving its spatial understanding.

[0058] Figure 3 shows an example of combining different modalities dataset to align data. The audio training data is aligned with modality training data, such as images

[0014] or videos

[0015] , to ensure temporal and spatial synchronization. This alignment process enhances the comprehensiveness of the training dataset. Figure 4 shows a joint embedding approach used in the multimodal machine learning model. The approach combines spatial features from audio training data and modality training data to create a unified representation. This facilitates the integration of diverse spatial features, such as direction, distance, and motion, into the model's learning process.

[0059] Figure 5 shows a process of recording using multiple microphones commonly arranged as an array and processing using beamforming. The figure highlights how delays in the audio signal are calculated and summed to focus on specific spatial location. The sound sensor captures audio signal and spatial information

[0114] , which are processed to generate frequency representation for each spatial location.

[0060] Figure 6 shows a beamformed illustration of a noise source and a car source. The figure represents spatial data in terms of azimuth and elevation coordinates with signal strength indicating summed audio power across all frequencies for each spatial location.

[0061] Figure 7 shows an overlay of the beamformed data from Figure 6 on a real image. The car is referenced as 1, the noise source as 2, and the bottom right corner as 3. This overlay demonstrates the spatial alignment between audio data and visual data, which is a aspect of the multimodal training process.

[0062] Figure 8 shows three Logmel Spectrograms corresponding to an arbitrarily picked refer- ence point and two points in Figures 6 and 7. The first spectrogram represents the car, the second represents the noise source, and the third represents the bottom right corner. Each spectrogram illustrates frequency and time variations for the respective spatial location.

[0063] Figure 9 shows a FAS (Frequency-Audio-Spatial) representation. The figure illustrates aggregated and contoured local maximum frequencies for one instance in time from Figures 6 and 7. The representation is an audio spectrum along every beamformed pixel point, where x and y represent spatial dimensions, and the color indicates the relative magnitude in frequency bands.

[0064] Figure 10 shows a mapping of relationships among Figures 7, 8, and 9. The figure demon- strates how spatial, temporal, and frequency data are interconnected to provide a comprehensive representation of the scene.

[0065] Figure 11 shows a FAST (Frequency-Audio-Spatial-Time) representation. This representation combines the FAS data from Figure 9 with temporal data from Figure 8 to create a 4D representation (Frequency, Time, X, Y). The FAST representation captures historical data in space, while the FASTT representation extends this to include a relative rolling window FAST representation for each timestep.

[0066] Figure 12 shows an example of the physical positioning of microphones used to capture audio training data.

[0067] Figure 13 shows a process of preparing FAS or FAST data. Figure 14 shows a method for deriving FAST data by first generating audio from each spatial cell using inverse fourier transform after the delay and sum to then process and acquire the spectograms.

[0068] Figure 15 shows another example of preparing FAS or FAST data.

[0069] Hence, it may be said that there is provided a system, process and method for training a model to improve the model audio-spatial understanding and generation (or diffusion).

[0070] In some embodiments audio-spatial data can be derived from using beamforming by first a process of establishing microphone physical coordinates and calculating beamformed spatial delays of the microphones used during recording, and in some embodiments use additional delays for adjusting phase and frequency delays to adjust for how sound and frequencies propagate through air, with an angle map that beamforms such that the values can be projected as physical space locations, and further adjust how sound and frequency arriving from a certain source to this beamformed angle map. In some embodiments the delays can be calculated by someone skilled in the art of utilizing beamforming and are established by calculating the speed of sound with respect to air pressure and temperature, and in some embodiments a constant for the speed of sound can be used. As sound travels through air, different frequencies have different phases as when they arrive on the microphone and in some embodiments adjusting for these can be done by calculating phase delays for the frequencies. A person skilled in the art can make variations and avoid using certain steps such while the result still falls under an acceptable result level. A person skilled in the art understands and can make variations such that the order and exact structure of calculations of which this system, method and process is done, in such cases the invention is the overall usage of the data representation for the specific task of training models to learn audio-spatial representations using the AS, FAS, FAST or a combined data type of the mentioned data and other modalities. AS stands for audio- spatial, FAS for frequency audio-spatial, and FAST for frequency audio spatial temporal. The letter A, can be thought of as a buffer of audio, where the S stands for Spatial, and can be thought of x,y, and optionally z. With F, we arrive at a audio (frequency) spectrum.

[0071] T stands for temporal and if F and T are combined then a Mel Spectogram is a representation that is used, the FAST is then a Spatial Mel Spectogram, where it is also understood that other formats and calculations can be used such as MFCC, Spectogram, Mel Spec- tograms, Chromagrams, Zero Crossing rate, RMS energy and others to resemble FAST, where in such case, this invention in some embodiments represents FAST as the spatial frequency through time, and in some embodiments. In some other embodiments, there is a FASTT representation where a model for each instance in time such as during building video generation models the data would contain an additional time aspect such that for each time step, the model can generalize on what the current state of sound is at the current time step, where for example if a video is 60 seconds long, the FAST part would be represented by some buffer size smaller than the video length for all the T timesteps. The recording of audio for this system is recognized as optional where the person skilled in the art using this invention can derive both audio, AS-data, FAS-data or FAST-data, FASTT, or combined data through a generative or diffusion model, by the mere fact that the model in some embodiments is trained to predict some variant of mentioned data, in such cases, the deriving of this data is part of the invention. Bucketing or aggregating AS, FAS, FAST, FASTT is also considered as part of this invention which can be done to decrease computational load where in some embodiment applying a mel spectogram is one such aggregation method. The usage of which data to use will depend on the use cases, FAS data can be used for continuous sound as it is smaller and more compact and can be further compacted by aggregating. AS-data can be used for compression related tasks. And FAST or FASTT data can be used where the audios historical sequence is important such as for the model to derive if a sword swing has just occurred (close in time spatial sound artifact) or is about to happen (no previous spatial sound artifact or longer interval since last swing sounds spatial artifact).

[0072] The invention might refer to data such as FAST, AS, FASTT, Audio-spatial data. And sometimes, FAS or FAST might be the general name name for all above data, especially when used without further context or obvious narrowing down for certain choice. This is done because the actual needed data will depend on what task is at hand and where AS data is also referenced as such to avoid circumvention by creating an inferior but patentwise bypassable solution. The best-mode of the paper is FAS and FAST, where FAS is better for classification tasks and FAST is better for generative tasks. Audio-spatial data is referred to as the general invention of using audio-spatial pixel or 3d point projected data to derive models that are audio spatially and temporally aware. Where the best-mode is to keep the spatial information in two dimensions.

[0073] The data can be processed with many algorithms for the one skilled in the art of acous- tics. Beamforming is the simplest and most effective for this use case, but in some embodiments adaptive beamforming can be applied and other algorithms such as MUSIC, or other known in the art.

[0074] The FAST data can be prepared by following this process:

[0075] 1. Load recordings from microphone array recordings. OR generated recordings. 2. Gather physical coordinates of microphones, [microphone, x,y[,z]]

[0076] 3. Calculate an angle map. a. Create an angle map; Height, Width 2-dimensional array (or tensor). b. Define your x,y angle starting point c. Define your x,y angle radius d. Calculate Phi, Theta, and optionally Psi and Map the angles onto the angle map. e. Convert to radians if applicable 4. Calculate the delays a. Create point coordinates for x, z and optionally z. Where the point coordinate shape is [dim, pixel_x, pixel_y] b. set distance constant to 50 c. example: point_coordinates[“x”, :] = distance * torch. sin(theta) * torch. cos(phi) d. then take the physical coordinates and for each dimension take the delta of it with the point coordinate across all point coordinates x and y and optionally z and square them. Then sum it in the dim dimension and square it to get the distances. e. calculate the delays with based on the sample_rate * (distances - distance) / speed of sound

[0077] 5. Calculate the phase delays a. e (delays*pi*li* ((-2 * k / n) + 1))) b. Where delays is height width and microphones. c. Where n is sample size. d. k is a integer array from 0 to n, and then fft shifted (without fourier transforming the integer array) such that the numbers go from 512 to 1024 and then 0 to 512. e. This creates a x,y, frequency, microphone datashape

[0078] 6. Start beamforming a. get the buffer sample of length n, optionally multiply it with banning window of length n and apply a fast fourier transform. b. multiply the phase delays with the fourier transformed buffer to create the frequency angle map. Further apply Mean on the microphone dimension. c. optionally enrich by getting the frequency magnitudes by taking the frequency_an- gle_map multiplied by the conjugate of the same frequency_angle_map. d. This has now created FAS.

[0079] 7. Optionally if optimizing for least artifact and simplicity, FAST can be created by a. get the buffer sample of length n * time_samples and apply a fast fourier transform. b. multiply the phase delays with the fourier transformed buffer to create the frequency angle map. c. inverse the fourier transform and sum on the microphone dimension. d. The shape then is x,y, audio wave. Across all x,y, apply a mel spectogram or another algorithm to arrive at some FAST representation. e. FASTT can be made by offsetting the buffer by a time-frame and performing above for each timestep where the buffer is offset by the time-frame.

[0080] 8. The result can now be saved in a dataset and further combined with other modalities, further explained in the next section.

[0081] There exists many multimodal training strategies. This patent does not claim invention on the training methods. This patent claims the invention of utilizing AS, FAS, FAST, FASTT representations to derive the data, and improve the models and the inference of such enriched models. The training regiment is in great detail by utilizing existing methods such as JEPA Joint Embedding Predictive Architecture and ImageBind. It is understood by the person skilled in the art that the invention can work across many different architectures and the approach is generally applicable. For instance, encoders and decoders, can be transformers with or without attention, with or without multi-head attention, but also with simple multi-layer- perceptrons, convolutions, diffusion, adapted LSTMs and as long there is a latent space representation that can be expressed in models, the system, method and process is applicable. In the figure 4, encoders are shown as what is in some embodiments a MLP or a convolution layer, and in some other embodiments the decoder is a MLP or a deconvolution layer, it is understood by the skilled in the art, that transform er architecture refers to en- coders and decoders slightly different, in where this invention classifies the initial modules in a transformer decoder as the joint embedding, or latent space aggregator, where the initial elements of transformers concatenate the user input encoded embedding received from the encoder with the model output embedding from the previous next-predicted token and in some embodiments of transformers, there is only a decoder where there is in some embodiments only a embedding input or if it is recurrent the next output is directly fed into the decoder again, where yet again this paper sees the deeper layers i.e the multi head attention or attention mechanisms of the decoder transformer are first decoding such latent space, hence the decoder aspect as traditionally referred to in transformer is seen as joint embedding in this invention and where the transformers later parts are the actual decoder. In modern transformers there is no longer any encoder however, the joint embedding or if the embedding is a single representation, the input of transformer could already be considered encoded, where transformers decoder mere previous intended use was to gather encoded user input as embeddings, and then reuse embeddings from the output stream for it’s previous next token, for the skilled in the art can understand that to a large extent the output of a transformer could first be decoded into some modality for instance, text, to then go into the encoder without having such a decoder. As such, this invention asserts that feature / latent spaces, encoders and decoders, are relative to the actual model architecture used. It is also understood by the skilled in the art that, the effect reached with the model can be arrived to by many methods, such as data can be split and further refined to no longer be in the same shape or format, for instance the time component is an iterator and so could other parts of the data or datareaders be structured . It is also understood by the skilled in the art the invention can be attempted to be cir- cumvented by using a single modal model encoder to generate a AS, FAS, FAST encoded representation, in such case, this is considered part of the invention, and skilled in the art can generate audio recordings that mimic the gathering of the invention and skilled in the art are also able to algorithmically simulate this same data, it is also considered part of this invention if a model has been trained to generate direct embeddings from a model that has been trained on this type of data or generated data or algorithmic data, the invention claims these as this method discloses the most efficient way to gather this type of data, whereas generated and algorithmic data likely cause artificial artifacts and have shortcomings of sorts where the bestmode is to record data directly as part of using the invented system, method and process. The dataset can be in some embodiments be built by choosing a set modality that is the anchor point. In imagebind, images are bounded with other modalities, for instance Image with Audio (I, A), Image with Text (I ,T), and Image with Depth (I , D). In such cases, training on the bounded modalities in the dataset creates indirect correlation such that training on a dataset of (11, Al) (12, T2) (13, D3), improves the recall of the model tasked to textually describe some aspect of depth based on only 12 where the depth learnings are learned by the joint embedding. In some embodiments joint embeddings are used and in other embodiments, it can be adapted by those skilled in the art to involve close to any architecture containing a latent space or by those skilled in the art understood to be equivalent to latent spaces from this context into the context of the other particular architecture. For instance, in convolutional neural nets there is a latent space at different scales as the image gets convoluted, and some architectures combine different levels of the convolution latent space to derive a latent space, in LSTMs, the internal memory state is a latent space where the inputs to the neuron can be adjusted to accept different encodings, latent embeddings or decoded such into a new latent embedding.

[0082] The invention in some embodiments adjusts or arrives at a model that to the losest term resembles a encoder and-or a decoder specifically for AS, FAS, FAST or FASTT data.

[0083] Adding such is trivial for the one skilled in the art, but for guidance, the FAS can be interpreted by a simple convolution with a filter for each frequency with the output channels being the same as the target joint embedding space count, and if more advanced solution is sought after a Vision transformer can be used. In some embodiments the decoder can be added by the one skilled in the art and for guidance a deconvolution decoder layer can learn to extract from the latent space to in some embodiments predict the exact first datapoint.

[0084] In some embodiments, AdamW can be used as an gradient optimizer and the exact parameters will vary on exact datasets and combinations made, and alignments used and the target purpose of the model and for the skilled in the art, the exact optimizer and hyperparameter is dependant on the use case. In some embodiments, a contrastive loss function can be used with either a fixed or variable temperature depending on the use case. Further the skilled in the art, will understand how to use the ImageBind architecture to incorporate FAST or it’s derivatives, where in some embodiments FAS can be encoded and treated as a simple image, meaning that in some embodiments a transformer and other embido- ments diffusion, and in third embodiments, deconvolutions or MLP decoders, or for the one skilled in the art, some way of generating an output based on the latent space can reproduce it.

[0085] The invention, by training the model on having bounded training sets, or, a combined dataset such as in some embodiments the effective shape of (x,y, f (+image color layers), t) or as mentioned above, the same effect through using a decoder only approach; or some generative model, will start transferring the audio-spatial constraints through in some embodiments gradient descent, and in other embodiments by feed-forward networks through methods like activity-based perturbation, or in some embodiments through some physi- cal process for deriving gradient descent, feed forward, or a machine learning method, like thermal-surface-based gradients or electrical propagation gradients. In some embodiments the training can include some inputs being bounded, and in other embodiments the outputs are bounded, and in other embodiments more than two modalities can be bounded, and in some training inputs or outputs can be single targets to train the joint em- bedding. In some embodiments, the joint embedding is a sum, average, concatenation, or similar aggregating function and in some embodiments is not directly tied to the same layers or neurons with other for the skilled in the art conceptually equivalent encoders and decoders, but where it conceptually achieves the same function. Encoders is a neural network function of reducing inputs into a smaller feature space, and decoding is used in this patent to describe the upscaling of the features or the extraction of features into an output modality using for instance multihead attention and transformers for next-token prediction.

[0086] This patent claims that the system uses some variation of FAST data and its variations derived through mentioned methods used to cross train models. Many have previously failed to arrive at this solution and it is not obvious that this connection exists and it is evident that its industrial value is high.

[0087] This method can be further applied for classification for object detection. By simply adjusting an existing yolov8 model and simply changing the input channels of the yolov8 model to be, in some embidoments just 512, or 1024 channels and adding our data. Cross train- ing is optional in this phase. This works by the convolutional layer interpreting each layer as a color, and in a convolutional layer, each output channel is considered a feature, where each feature has a kernel for each color or in this case frequency, creating a 3x3 kernel per channel per feature, creating the ability to learn directly as if this was a normal convld on a mel spectogram. 1 Enhancing Spatial Understanding Of Machine Learning Models Method Details

[0088] The method for generating a multimodal machine learning model involves a sequence of steps beginning with data acquisition and concluding with model training. The process starts by obtaining audio training data that represents the spatial distribution of audio frequencies within a scene. Concurrently, modality training data associated with the same scene is obtained. It should be mentioned that the relatedness between the audio and modality training data can vary depending on the specific task the data should address for the model. These two data types are then combined to form a training dataset / added to a training regiment. A machine learning model is subsequently trained using this dataset to produce the final multimodal machine learning model. This method utilizes specific data representations for the audio, such as audio-spatial (AS), frequency audio-spatial (FAS), and frequency audio-spatial temporal (FAST or FASTT). The choice of representation depends on the specific application; for instance, FAS data may be suited for classification tasks, while FAST data is often used for generative tasks. The training can be implemented with various model architectures, including transformers, convolutional neural networks, and diffusion models, and may employ training strategies like Joint Embedding Predictive Architecture (JEPA). The overall procedure is designed to enable the model to learn and internalize audio-spatial relationships from the provided data. It may alternatively be said that the overall procedure is to provide audio-spatial constraints between modalities to force the model to be comparably more audio-spatially coherent.

[0089] 1.1 Obtaining Audio Training Data Representing Spatial Distribution of Audio Frequencies

[0090] The acquisition of audio training data is one of many steps in the method. This data is structured to capture not just the audio content but also its spatial characteristics, representing how audio frequencies are distributed across multiple locations in a scene. The process can be executed by using a machine learning model configured to produce multichannel or spatial-frequency audio data. Alternatively, it can be generated directly from physical recordings. A practical implementation involves positioning multiple sound sensors, such as those in a microphone array, within a scene and recording audio signals from them simultaneously. The physical coordinates of each sensor are measured and recorded. Following the recording, a set of spatial locations within the scene is defined. For each of these locations, time delays between the signals arriving at different sensors are calculated. Beamforming processing is then applied, using these calculated delays to computationally focus on each spatial location and each frequency. This results in the generation of a plurality of frequency representations, one for each spatial location, which may also capture temporal variations. The collected raw data streams are saved to a storage system with metadata, including precise timestamps, for subsequent processing. The final data can be organized as a two-dimensional or three-dimensional frequency-spatial representation. It may alternatively be expressed that the final data can be organized as any one of AS, FAS, FAST, and FASTT as previously described herein.

[0091] 1.1.1 Positioning and Utilizing Sound Sensors within a Scene

[0092] The initial action for obtaining audio training data is the physical placement of a plurality of sound sensor(s) within the target scene. This arrangement is typically a microphone array, and its configuration is designed to effectively capture the spatial attributes of sound. The physical coordinates of each microphone in the array in a three-dimensional (x, y, z) system are recorded. The precise positioning of these sensors is a determining factor for the quality of the spatial information that can be extracted. The distances and relative orientations between the microphones influence the time-of-arrival differences of sound waves, which are subsequently used in beamforming calculations. An example of such a setup is illustrated in Figure 12. This step is a prerequisite for the simultaneous recording of audio signal(s), as the known geometry of the sensor array is required for the algorithms that calculate spatial audio features.

[0093] 1.1.2 Recording and Processing Audio Signals at Multiple Spatial Locations

[0094] Once the sound sensor(s) are positioned, audio signal(s) are recorded simultaneously from all sensors in the array. This simultaneous capture ensures that the time relationships between the signals at each microphone are preserved. The subsequent processing aims to derive spatial audio information from these multi-channel recordings. A grid of discrete spatial location(s) within the scene is first defined. For each point in this grid, the time delays of arrival for a sound originating from that point to each microphone are calculated. These calculations are based on the known microphone positions and the speed of sound, wherein the speed of sound varies based on various factors such as temperature, wind conditions, pressure, and air composition.

[0095] With the time delays established, a beamforming technique is applied. Beamforming is a signal processing method used to focus the listening direction of the microphone array to a specific point in space, effectively amplifying sound from that location while attenuating sound from others. Various beamforming algorithms can be used, including delay-and- sum, adaptive beamforming, neural beamforming, deconvolution, neural deconvolution, or the MUSIC algorithm. This process is repeated for all defined spatial location(s), resulting in a set of audio signal(s) or frequency representation(s), each corresponding to a specific point in the scene. This processed data forms the basis of the audio-spatial representations like FAS and FAST.

[0096] 1.1.3 Calculating Time Delays and Applying Beamforming Techniques

[0097] Following the simultaneous recording of audio from the microphone array, the time delays for sound arriving at each sensor are calculated for a predefined grid of spatial location(s). This calculation relies on the known physical coordinates of each microphone and a value for the speed of sound, which can be a constant or adjusted for environmental factors like air pressure and temperature. For each point in the spatial grid, the difference in distance to each microphone is computed, and this distance is converted into a time delay. Once the time delays are established for every spatial location, a beamforming algorithm is applied. This signal processing technique uses the calculated delays to computationally steer the focus of the microphone array. By applying the respective delay to each microphone's signal and then summing them, sound originating from the target location is constructively reinforced, while sound from other locations is attenuated. Alternative beamforming methods may be employed, such as adaptive beamforming, which dynamically adjusts to nullify interfering noise sources, or the MUSIC algorithm for high-resolution source direction finding. Neural beamforming represents another alternative that uses a trained model to perform this focusing task. Other beamforming methods include neural beamforming, neural deconvolution, deconvolution beamforming.

[0098] 1.1.4 Generating Frequency Representations for Spatial Locations

[0099] The beamforming process generates a respective frequency representation for each spatial location through either frequency-domain methods such as Steered-Response Power with Phase Transform (SRP-PHAT) or time-domain beamforming followed by frequency transformation using Fast Fourier Transform (FFT), optionally with windowing functions like Hanning windows to reduce spectral leakage.

[0100] The result for a single time chunk is a frequency spectrum for each spatial point. To capture how frequency content changes over time, this process is repeated on consecutive, often overlapping, audio segments, creating a spectrogram. Variations include generating Mel spectrograms, which use a frequency scale based on human perception, or other formats like MFCCs or chromagrams, depending on the features of interest for the specific application. The disclosure also describes a method where phase delays are multiplied with the Fourier-transformed buffer to directly generate a frequency-angle map, which constitutes the Frequency-Audio-Space (FAS) representation, and with the possibility to gather frames to create spectograms for FAST or FASTT as illustrated in Figure 13.

[0101] 1.2 Obtaining Modal Training Data Associated with the Scene

[0102] This step involves collecting data from other modalities that correspond to / is associated with the same scene from which the audio training data was gathered. This modality training data provides supplementary context, allowing the model to build a more complete understanding of the environment. The types of data that can be obtained are varied and may include visual data such as images and videos, text data, coordinate data, depth maps, and thermal data. Other potential data types include radar data, radio frequency data, vibration data, or even latent representations from other models. In some circumstances, the collection of this data is performed in conjunction with the audio recording to ensure that the information from all modalities is synchronized in time and space. For example, if video is used as a modal data type, it is captured at the same time and from a relevant perspective as the audio is being recorded. This temporal and spatial alignment may be a precursor to the subsequent step of combining the datasets for training.

[0103] It may in this context be noted that synchronized data collection is not strictly required, as the training data can have varied amounts of content-based, temporal, and spatial relatedness depending on the specific task and whether models are pre-trained. In some instances, it is sufficient to combine disparate data such as a cat's text description with unrelated acoustic scene data. This approach can train a model's ability to generate images of cats based solely on acoustic input. Such capability is achieved through training on separate text-acoustic and text-image pairs. The multimodal training process forces the model to learn shared representations that are valid across these cross-modal scenarios. This enables knowledge transfer even when direct temporal or spatial correspondence between modalities is absent.

[0104] 1.2.1 Collecting Visual, Depth, Radar, and Other Modal Data

[0105] In parallel with the audio data acquisition, data from other modalities is collected from the same scene. This process involves using a variety of sensors to capture supplementary information that provides context to the audio events. For instance, high-resolution cameras are used to capture visual data as still images or video streams. The precise three- dimensional coordinates of each sensor are measured and recorded to ensure proper spatial alignment with the audio data. Depth sensors, such as LiDAR or structured light cameras, may be used to generate depth maps that provide three-dimensional information about the scene geometry.

[0106] Depending on the application requirements, other sensors can be incorporated. Radar systems can be used to detect objects and their velocities, even in adverse weather conditions. Thermal cameras can capture heat signatures, and text data might be generated through transcription of speech or manual annotation of the scene content. All sensors begin recording simultaneously to capture a synchronized snapshot of the scene across all modalities. The selection of modalities is tailored to the intended use case of the final model, with the goal of gathering a rich, multi-faceted dataset of the scene. 1.2.2 Aligning Modal Training Data with Audio Training Data

[0107] After collecting data from all modalities, a preprocessing step is to ensure their alignment. This involves synchronizing the data both temporally and spatially. Temporal alignment is typically achieved by using a common clock for all recording devices or by using times- tamps embedded in the data files. This allows for the precise matching of a video frame, for example, with the corresponding segment of audio that was recorded at the exact same moment.

[0108] It may in this context be noted that the alignment process encompasses three distinct types of inter-modal correspondence that do not require strict spatial or temporal synchro- nization. Direct correspondence involves spatially and temporally aligned modality pairs such as synchronized audio-video or pixel-aligned image-depth data. Semantic correspondence includes conceptually related but spatial ly-temporally independent pairs, such as acoustic recordings with textual descriptions or audio-spatial data with object labels. Transitive correspondence enables indirect relationships learned through shared interme- diate modalities, wherein one modality relates to another through shared connections with a third modality, such as audio-spatial data connecting to radar data through shared text descriptions. There may be employed a correspondence-adaptive training approach that applies strong alignment constraints for directly corresponded pairs, uses weak semantic alignment for conceptually related pairs, and enables zero-shot transfer between modali- ties connected only through transitive relationships. This architecture permits training on heterogeneous datasets with varying alignment qualities without requiring spatial or temporal alignment as a prerequisite for learning meaningful cross-modal representations.

[0109] Spatial alignment establishes a geometric correspondence between the different sensor coordinate systems. This involves a calibration procedure to determine the mathematical transformation that maps a point in one modality's space to the corresponding point in another's. For example, it defines how a pixel coordinate in a camera image relates to a specific direction in the audio beamforming map. Successful alignment ensures that a sound source identified at a particular location in the audio data corresponds accurately to its visual representation in the image or video data. 1.3 Preparing and Processing Audio Training Data

[0110] Following the acquisition of raw audio recordings, a preparation and processing stage is undertaken to transform the data into a structured format suitable for the machine learning model. This stage involves steps such as data cleaning to remove noise or artifacts, formatting the data into a consistent structure, and aligning the audio training data with the corresponding modality training data to ensure temporal and spatial synchronization.

[0111] A specific procedure for preparing FAST (Frequency-Audio-Space-Time) data begins with loading recordings from a microphone array and gathering the physical coordinates of each microphone. An angle map is calculated to define the spatial grid. Subsequently, time delays and phase delays are calculated for each point on the grid relative to each microphone, accounting for the speed of sound. Beamforming is then applied by taking a buffer of the audio sample, applying a Fourier transform, and multiplying it with the calculated phase delays. This creates a frequency-angle map, or FAS data. To generate FAST data, this process can be extended over time, for instance by applying a Mel spectrogram to the beamformed audio wave for each spatial point over a time window. The resulting structured data is then ready to be saved and combined with other modalities.

[0112] 1.3.1 Cleaning and Formatting Audio Training Data

[0113] The initial stage of audio data preparation involves cleaning and formatting the raw recordings obtained from the microphone array. The cleaning process aims to improve the signal quality by removing unwanted artifacts. This can include applying digital filters to reduce steady background noise, such as hiss or hum, and using algorithms to detect and eliminate transient sounds like clicks or pops that may have occurred during recording.

[0114] Following the cleaning process, the audio data is standardized into a uniform format. All recordings are converted to the same file type, sample rate, and bit depth to ensure con- sistency across the entire dataset. The continuous recordings may also be segmented into smaller, fixed-length chunks or buffers. This standardization simplifies subsequent processing steps and ensures that the data is in a predictable and compatible structure before feature extraction begins.

[0115] 1.3.2 Aligning Audio Data with Modal Data for Synchronization To create a cohesive multimodal dataset, the processed audio (training) data should preferably be at least directly, semantically or transistively related with the data from all other modalities. This alignment process addresses both time and space. Temporal alignment uses the recorded timestamps to synchronize events across all data streams to the same moment in time, ensuring that events across different data streams correspond to the same point in time. This is achieved by pairing audio segments with corresponding video frames, text annotations, or other modal data points based on their recording timestamps.

[0116] Spatial synchronization establishes the geometric relationship between the audio scene and the other modal representations. Through a calibration process, a mapping is created between the coordinate system of the audio data, such as the beamformed angle map, and the coordinate systems of other sensors, like a camera's pixel grid. This alignment is what allows the system to associate a sound from a specific direction with the correct object in an image, a foundational requirement for the model to learn cross-modal spatial relationships.

[0117] 1.3.3 Extracting Relevant Spatial Features from Audio Data

[0118] The core of the audio processing pipeline is the extraction of spatial features from the multi-channel recordings. This step transforms the raw audio signal into a structured rep- resentation that explicitly encodes information about the location of sound sources. The primary method used is beamforming, which processes the signals from the microphone array to isolate the sound originating from numerous discrete points within the scene.

[0119] This procedure effectively generates a spatial map of the soundscape. For each point in the map, the system extracts an audio signal that represents the sound from that specific direction. The output is not a single audio stream but a collection of signals, each tagged with spatial coordinates. This process makes explicit the spatial attributes of the sound field, such as the direction, relative intensity, and position of different sources, as illustrated in Figure 6.

[0120] 1.3.4 Transforming Audio Data into Model-Compatible Formats The final step in the audio data preparation pipeline is to transform the extracted spatial features into a format suitable for ingestion by a machine learning model. This involves structuring the data into specific multi-dimensional representations, such as the Frequency-Audio-Space (FAS) or Frequency-Audio-Space-Time (FAST) formats. It may be noted that an Audio-Spatial (AS) representation can also be generated as part of the multi-dimensional data formats, wherein the AS representation comprises spatial coordinates (x, y) with time-domain samples, and wherein an inverse Fast Fourier Transform (IFFT) can be applied after the signal is phase delayed to retrieve the spatial signal in the time domain.

[0121] To create a FAS representation, a frequency spectrum is generated for each spatial point (x, y) at a single moment in time, resulting in a 3D data structure analogous to a multichannel image. The FAST representation extends this by adding a time axis; for each spatial point, a full spectrogram (frequency vs. time) is created, yielding a 4D data structure. These structured tensors are then normalized and organized into a dataset. This final format, which encapsulates spatial, frequency, and temporal information, is then ready to be combined with other modalities for training the model.

[0122] 1.4 Generating a Training Dataset by Combining Audio and Modal Training Data

[0123] This step involves the creation of a unified training dataset by merging the prepared au- dio training data with the modality training data. The objective is to construct a dataset where each entry contains synchronized information from multiple modalities, allowing the machine learning model to learn the relationships and correlations between them. For example, an audio spatial frequency image (FAS data) can be paired with a corresponding visual image of the scene. The combination process ensures that the data is aligned both temporally and spatially. This alignment is a key aspect, as it enables the model to associate specific sounds with their visual sources or other modal characteristics. The structure of the dataset can be adapted based on the training strategy. One approach is to use a joint embedding framework where pairs of data from different modalities, such as (Image, Audio) or (Image, Text), are composed of synchronized pairs, which allows the model to learn a shared latent space. An alternative approach involves concatenating the features from different modalities into a single, larger data tensor for each training example, which might be used for models like certain CNNs that are designed to process multi-channel inputs directly. This cross-modal training improves the model's ability to generalize and understand spa- tial features from multiple perspectives.

[0124] It should be noted that the above relates to direct correspondence. However, it is conceivable that semantic correspondence and transitive correspondence may also be employed in the combination process. In semantic correspondence, the model can be trained to produce enhanced audio outputss by using inputs such as text descriptions paired with conventional audio files, generating outputs that include both acoustic data and directional acoustic information. Through this semantic approach, the modalities learn to express shared semantic content such that the results converge without requiring explicit paired data. Transitive correspondence enables indirect learning relationships where modalities connect through intermediate representations, allowing the model to develop cross-modal understanding even when direct correspondence between specific modality pairs is not available in the training data. 1.5 Training a Machine Learning Model Using the Training Dataset

[0125] The training phase uses the generated multimodal dataset to train a machine learning model, with the goal of producing a model that has an enhanced understanding of audio- spatial relationships. The training can be performed using a variety of architectures, such as transformers, convolutional neural networks (CNNs), recurrent neural networks (RNNs), or diffusion models. The choice of architecture and training parameters, such as the optimizer (e.g., AdamW) and loss function (e.g., contrastive loss), will depend on the specific task, be it generation, classification, or another application. Hyperparameters like learning rate and batch size are adapted to the specific application. The training is executed by iteratively feeding batches of data from the dataset to the model. For each batch, the model produces an output, which is compared against a target or used in a comparative loss function to calculate an error signal. An optimizer then uses this error to update the model's weights via gradient descent. This procedure is repeated for many iterations or epochs until the model's performance on a validation set converges. A specific technique employed during training involves adding at least one encoder unit configured to process the audio training data. The encoder, which could be a convolutional neural network or a vision transformer, transforms the audio-spatial representation into a lower-dimensional latent vector that the core of the model can process. After the training is complete, this specialized encoder can optionally be removed from the final multimodal machine learning model. This is often done when the primary goal is to utilize the learned joint embedding space, allowing the model to perform tasks involving modalities other than the original audio-spatial data. Similarly, for generative tasks, a decoder unit can be added to translate the model's latent representation back into a specific output modality. This decoder can also be a temporary component added for a specific task and not part of the core saved model. The management of these units provides flexibility in how the trained model is ultimately deployed and used. The training process itself, typically driven by gradient descent, adjusts the model's parameters to minimize the loss function, thereby transferring the spatial constraints present in the training data into the model's internal representations. 1.6 Model Inference and Output Generation

[0126] Once training is complete, the method enables inference to generate new data. The inference process begins by providing the trained multimodal machine learning model with input data from one or more modalities. This input could be, for example, a text description, an image, or a silent video. The provided input data is first preprocessed using the same methods that were applied to the training data for that modality to ensure format consistency.

[0127] The preprocessed input is then fed into the trained model. If the input modality has a dedicated encoder unit that was retained after training, the data passes through this encoder first to be converted into the model's shared latent space. The model then processes this latent representation through its learned internal representations.

[0128] Following the processing of the input data, the model generates an output in a target modality. The model's internal latent representation is passed to a decoder unit, which translates the latent vector into a high-dimensional data format, such as an image, an audio waveform , or a text string.

[0129] The method produces outputs that maintain multimodal spatial coherence, where audio events align spatially with corresponding visual objects or other modal features. The quality of the generated output data is enhanced with respect to spatial characteristics such as sound source localization, directional accuracy, distance estimation, acoustic environment modeling, and movement tracking. These outputs can represent audio data, visual data, text data, coordinate data, radar data, radio frequency data, vibration data, depth maps, latent representations, or thermal data, with the specific output characteristics determined by the target modality and inference task.

[0130] The generated raw output from the model is then post-processed to convert it into a stan- dard, usable format. An audio waveform tensor is converted into a WAV file, while an image tensor is saved as a PNG or JPEG file. The final output is interpreted based on the task. For a generative task, it produces new media content. For a classification task, such as identifying a sound source, the output provides a class label or a set of coordinates derived from the model's spatial audio representation. This generation step may involve adding / having a specific decoder unit to translate the model's internal state into the desired output format, which can be optionally removed after inference if needed for deployment optimization.

[0131] 1.7 Performance Specifications and Requirements

[0132] The method enables the generation of multimodal content exhibiting spatial coherence, where generated audio events show a high degree of spatial alignment with corresponding visual objects. Perceived sound source locations accurately correlate with visual object positions in the output. For dynamic content such as video, the method produces models capable of rendering dynamic spatial changes at a minimum update rate of 25 frames per second, allowing for accurate capture of movements of sound sources or visual objects and evolving acoustic environments.

[0133] The resulting models reflect realistic spatial acoustics by incorporating characteristics of the acoustic environment, including reverberation and occlusion, that influence sound propagation. The method enables spatial differentiation and localization of at least five distinct sound sources concurrently within a given scene, ensuring clarity and separation of multiple audio events. Through the integration of diverse spatial features such as direction, distance, size, and motion from both audio and visual modalities, the method produces models with a unified and consistent spatial representation. The trained models exhibit improved generation quality on novel spatial arrangements and acoustic environments, demonstrating robust generalization of spatial understanding across a variety of unseen scenarios. This generalization capability results from the method's approach of training on spatially-explicit audio representations combined with other modalities, enabling the model to learn fundamental spatial relationships rather than memorizing specific scene configurations.

[0134] 2 Potential Applications

[0135] The disclosed technology has applications across several fields where a detailed understanding of audio-spatial relationships is beneficial. The method of training models with spatially explicit audio data enables the development of systems with an enhanced capac- ity to perceive, interpret, and generate spatially coherent multimodal content. Fields that can benefit from this include media creation and consumption, robotics and autonomous navigation, security and surveillance systems, and assistive technologies for accessibility.

[0136] 2.1 Audio-Visual Scene Understanding and Generation

[0137] In the domain of audio-visual media, a common challenge is the generation of audio that is spatially consistent with the visual content, and vice versa. Models often produce audio that is contextually appropriate but lacks correct spatial placement; for example, the sound of a car may not appear to originate from the car's location in a video frame. This disclosure addresses this by training models on audio-spatial data, such as the FAS representation, that is precisely aligned with visual data. A model trained with this method learns the geometric relationship between sound sources and their visual counterparts. Consequently, it can perform tasks like generating a spatially accurate audio track for a silent video, where sounds of objects emanate from their correct on-screen positions and move with them. This capability improves the realism and immersiveness of generated content, representing an advance over methods that only achieve contextual or temporal alignment without spatial correspondence.

[0138] 2.2 Robotics and Autonomous Systems For robotics and autonomous systems, situational awareness is a performance-limiting factor. Navigation and interaction often depend on line-of-sight sensors like cameras and LiDAR, which cannot detect objects or events that are occluded. Audio provides a complementary sensing modality that can perceive events from all directions, but conventional audio processing often lacks the precision to be useful for navigation. This disclosure allows a robotic system to build and maintain a detailed audio-spatial map of its surroundings. By processing multi-channel audio through a model trained with audio- spatial data, the system can localize sound sources with high precision. This enables the robot to detect and locate unseen hazards, such as a vehicle approaching from around a corner, or to identify the position of a person speaking to it. This enhanced perception complements existing sensors, leading to safer navigation in dynamic environments and more effective human-robot interaction.

[0139] 2.3 Surveillance, Security, and Monitoring

[0140] In surveillance applications, monitoring large areas with multiple audio and video feeds presents a challenge for human operators and automated systems. While audio can signal an event of interest, such as shouting or breaking glass, pinpointing the event's location in a noisy, complex environment is often difficult. This can delay response times and reduce the effectiveness of the surveillance system.

[0141] The disclosed method can be used to develop a system that not only detects an anomalous sound but also precisely identifies its point of origin. A model trained on audio-spatial data can analyze the incoming audio from a microphone array and overlay a visual indicator on the corresponding video feed at the exact location of the sound source. This provides immediate, actionable spatial intelligence to operators, allowing for a more rapid and targeted response than is possible with systems that provide only a non-localized audio alert. 2.4 Assistive Technologies and Accessibility

[0142] Assistive technologies for individuals with visual impairments aim to provide information about the surrounding environment to aid navigation and awareness. Existing solutions may use computer vision to identify objects or provide general directional sound cues, but they often lack the ability to convey a detailed spatial understanding of the scene.

[0143] This disclosure can be applied to create an assistive device that captures the ambient soundscape and translates it into a more descriptive form of feedback. A model trained with audio-spatial data can interpret the complex auditory scene and provide the user with rich spatial information, such as the location, distance, and movement of multiple sound sources. This information could be conveyed through spatialized audio headphones or a haptic interface. Such a system would offer a more detailed perception of the environment than existing technologies, potentially enhancing user mobility and safety. The terminology used herein is for the purpose of describing particular aspects only and is not intended to be limiting of the disclosure. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items. It will be further understood that the terms "comprises," "comprising," "includes," and / or "including" when used herein specify the presence of stated features, integers, actions, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, actions, steps, operations, elements, components, and / or groups thereof.

[0144] It will be understood that, although the terms first, second, etc., may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and, similarly, a second element could be termed a first element without departing from the scope of the present disclosure.

[0145] Relative terms such as "below" or "above" or "upper" or "lower" or "horizontal" or "vertical" may be used herein to describe a relationship of one element to another element as illustrated in the Figures. It will be understood that these terms and those discussed above are intended to encompass different orientations of the device in addition to the orientation depicted in the Figures. It will be understood that when an element is referred to as being "connected" or "coupled" to another element, it can be directly connected or coupled to the other element, or intervening elements may be present. In contrast, when an element is referred to as being "directly connected" or "directly coupled" to another element, there are no intervening elements present.

[0146] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. It will be further understood that terms used herein should be interpreted as having a meaning consistent with their meaning in the context of this specification and the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0147] It is to be understood that the present disclosure is not limited to the aspects described above and illustrated in the drawings; rather, the skilled person will recognize that many changes and modifications may be made within the scope of the present disclosure and appended claims. In the drawings and specification, there have been disclosed aspects for purposes of illustration only and not for purposes of limitation, the scope of the disclosure being set forth in the following claims.

Claims

Claims1. A computer-implemented method for generating a multimodal machine learning model, comprising:(a) obtaining audio training data representing a spatial distribution of audio fre- quencies at multiple locations within a scene;(b) obtaining modality training data associated with the scene;(c) generating a training dataset by combining the audio training data with the modality training data;(d) training a machine learning model using the training dataset to generate a multimodal machine learning model;2. The computer-implemented method of claim 1, wherein the modality training data comprises any of audio data, visual data, text data, coordinate data, radar data, radio frequency, vibration data, depth maps, latent representations, or thermal data.

3. The computer-implemented method of any of the preceding claims, wherein the audio training data comprises temporal variations of the audio frequencies at the multiple locations.

4. The computer-implemented method of any of the preceding claims, wherein the audio training data is generated using a machine learning model configured to produce multi-channel audio data or spatial-frequency audio data.

5. The computer-implemented method of any of the preceding claims, wherein the step of obtaining audio training data comprises:(a) positioning a plurality of sound sensors within the scene;(b) simultaneously recording audio signals using the plurality of sound sensors;(c) defining a plurality of spatial locations within the scene; (d) calculating time delays between the sound sensors for each spatial location;(e) applying beamforming processing to focus on each spatial location using the calculated time delays; and,(f) generating a plurality of frequency representations for the plurality of spatial locations.

6. The computer-implemented method of claim 5, wherein the step of applying the beamforming processing comprises any of: neural beamforming; delay-and-sum beamforming;adaptive beamforming; or, MUSIC algorithm processing.

7. The computer-implemented method of any of the preceding claims, wherein the step of training the machine learning model is based on any of: a transformer; a convolutional neural network; a recurrent neural network; a diffusion model; a generative adversarial network; a joint-embedding predictive architecture; a deep neural network; a generative neural network.

8. The computer-implemented method of any of the preceding claims, wherein the step of training the machine learning model comprises:(a) adding at least one encoder unit configured to process the audio training data; (b) training the machine learning model including the at least one encoder unit; and,(c) optionally removing the at least one encoder unit from the multimodal machine learning model.

9. The computer-implemented method of any of the preceding claims, wherein the audio training data is organized as a two-dimensional or a three-dimensional frequency- spatial representation.

10. A computer-implemented method for generating modality output data comprising:(a) providing a multimodal machine learning model according to any of claims 1 to 9;(b) providing, to the multimodal machine learning model, modality input data; (c) generating, using the multimodal machine learning model, modality output data.

11. The computer-implemented method of claim 10, wherein the modality output data comprises any of audio data, visual data, text data, coordinate data, radar data, radio frequency, vibration data, depth maps, latent representations, or thermal data.

12. The computer-implemented method of claim 11, wherein the modality input data comprises any of audio data, visual data, text data, coordinate data, radar data, radio frequency, vibration data, depth maps, latent representations, or thermal data.

13. The computer-implemented method of any of claims 10 to 12, wherein the step ofgenerating modality output data comprises:(a) adding at least one decoder unit configured to generate the modality output data;(b) optionally removing the at least one decoder unit from the multimodal machine learning model after generating the modality output data.

14. A computer-implemented method for classifying an audio source using spatial audio data, comprising:(a) providing a multimodal machine learning model according to any of claims 1 to 9;(b) providing, to the multimodal machine learning model, input spatial audio data associated with an audio source to generate a spatial audio representation; and,(c) classifying the audio source based on the spatial audio representation.