Real-time multi-object tracking, synchronization, and spatialization of media content using multimodal ai

WO2026039563A3PCT designated stage Publication Date: 2026-03-19FERRER JULIO
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Existing technologies face challenges in efficiently tracking, synchronizing, and spatializing audio and media content in real-time across multiple environments and devices, particularly in venues with varying locations and network conditions, often leading to desynchronization and suboptimal spatial audio experiences.

Method used

A Multimodal AI system integrated with wireless communication sensors uses methods like Time Difference of Arrival (TDoA), Single-Sided Two-Way Ranging (SS-TWR), and Phase Difference of Arrival (PDoA) signals to determine sound sources, combined with spatial transfer functions and AI models like Generative FDNs and Neural Acoustic Fields to generate spatialized audio, while leveraging Multi-Link Operation (MLO) for high-performance data transmission and synchronization across multiple venues.

Benefits of technology

Enables real-time, accurate, and synchronized multi-object tracking and spatialization of audio and media content across diverse environments, enhancing the immersive experience by maintaining precise timing and spatial coherence despite network variations and device positions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025041867_19032026_PF_FP_ABST
    Figure US2025041867_19032026_PF_FP_ABST
Patent Text Reader

Abstract

A system for synchronizing and spatializing media, comprising client device and server device. A method for operation includes: tracking the real-time spatial coordinates of both devices via wireless sensors; synchronizing media playback using a timestamp-based protocol that compensates for network latency and clock offset; and applying directional audio filters, such as Head-Related Transfer Functions (HRTF), on the client device to render audio appearing to originate from server device's physical location. The method is further characterized by compensating for acoustic propagation and signal processing delays to ensure accurate spatiotemporal alignment. The system may be enhanced by a multimodal Al configured to perform operations such as real-time voice-preserving translation and the generation of synchronized, spatialized haptic feedback based on semantic event markers.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEM AND METHOD FOR REAL-TIME MULTI-OBJECT TRACKING, SYNCHRONIZATION, AND SPATIALIZATION OF MEDIA CONTENT USING MULTIMODAL Al SYSTEM AND WIRELESS COMMUNICATION SENSOR

[0001] The present application claims the priority benefit of U.S. provisional patent applications No. 63 / 682,461 , No. 63 / 682,449, and No. 63 / 682,456 filed Aug. 13, 2024 and entitled “System and Method for Real-Time Multi-Object Tracking, Synchronization, and Spatialization of Media Content Using Multimodal Al System and Wireless Communication Sensor,” the disclosure of which is incorporated herein by scope of the appended claims.

[0002] This invention relates to medium, method, and system for the multiobject tracking, synchronization, and spatialization of one or more of media content using one or more of Multimodal Al (artificial intelligence) system and wireless communication sensor with one or more electronic device. Particularly, the invention relates to a medium, method, and system for the multi-object tracking, synchronization, and spatialization of one or more of audio and media content using one or more of Multimodal Al System and wireless communication sensor with one or more electronic device. More particularly, the invention relates to a medium, method, and system for real-time multi-object tracking, synchronization, and spatialization of one or more of audio and media content using one or more of Multimodal Al System and wireless communication sensor with one or more electronic device. Even more particularly, the invention relates to a medium, method, and system for real-time multi-object tracking, synchronization, and spatialization of one or more of audio and media content using one or more of Multimodal Al System and wireless communication sensor with one or more electronic device within and around one or more of venue, public, private, office, open space, educational institution, store, storefront, home, vehicle, chat, gaming, virtual, and an other environment. Specifically, the invention relates to a medium, method, and system for real-time conversational, transactional, and ambient multiobject tracking, synchronization, and spatialization of one or more of audio and media content using one or more of Multimodal Al System and wireless communication sensor with one or more electronic device within and around one or more of venue, public, private, office, open space, educational institution, store, storefront, home, vehicle, chat, gaming, virtual, and an other environment. More specifically, the invention relates to a medium, method, and system for real-time conversational, transactional, and ambient multi-object tracking, synchronization, and spatialization of one or more of audio and media content using one or more of Multimodal Al System and wireless communication sensor with one or more electronic device within and around one or more of venue, public, private, office, open space, educational institution, store, storefront, home, vehicle, chat, gaming, virtual, and an other environment over a communications network. The invention provides alternative methods for real-time conversational, transactional, and ambient multi-object tracking, synchronization, and spatialization of static and non-static of one or more client-side device, server-side device, and wireless communication sensor.

[0003] According to embodiments of the invention, platform is provided for real-time multi-object tracking, synchronization, and spatialization of one or more of one or more audio content, media content, mixed audio content, and other data input during user’s interaction with one or more of one or more Multimodal Al System and wireless communication sensor. For example, multi-object tracking may determine sound source by leveraging the time delay between signals received by one or more wireless communication sensor. For example, multiobject tracking may determine sound source by leveraging one or more of one or more Time Difference of Arrival (TDoA), Single-Sided Two-Way Ranging (SS- TWR), Double-Sided Two-Way Ranging (DS-TWR), and Phase Difference of Arrival (PDoA) signals received by one or more wireless communication sensor. For example, multi-object tracking may determine the sound pressure level (SPL) difference between user’s head and static and non-static object. For example, multi-object tracking may determine the SPL difference between static and nonstatic user’s head and one or more object. For example, one or more PDoA signal, time delay, and SPL difference may determine the orientation and position between user’s head and object.

[0004] For example, data input may comprise text, audio, video, images, physics-informed, metadata tags, language models, modalities, sensory inputs, and other data input.

[0005] For example, audio content may comprise of one or more object-based audio, channel-based audio, scene-based audio, metadata, and an other audio signal input. For example, the metadata may include the position and gain of one or more of one or more object-based, scene-based, and channel-based audio signals, and other rendering metadata. For example, audio content may comprise audio generated by Multimodal Al System. For example, audio content may comprise synchronized audio generated by Multimodal Al System. For example, audio content may comprise spatialized audio generated by Multimodal Al System. For example, audio content may comprise one or more of one or more synchronized and spatialized audio generated by Multimodal Al System. For example, audio content may comprise one or more of one or more audio content, mono audio content, stereo audio content, multiple audio content, binaural audio content, ambisonic audio content, and alternative audio content.

[0006] For example, binaural or spatial audio content may be produced by processing an audio signal with a spatial transfer function. Said spatial transfer function may comprise, without limitation, a head-related impulse response (HRIR), a head-related transfer function (HRTF), a binaural room impulse response (BRIR), or a Spatial Room Impulse Response (SRIR). A key objective of this processing is often to improve the sense of externalization for the listener. To this end, a filter may be used to modify the interaural coherence (IC), a critical perceptual cue which can be calculated from said spatial transfer function. Furthermore, said HRIR, HRTF, or SRIR may be personalized, for example, measured for or adapted to a specific user, or may be non-personalized, for example, based on a generic acoustical model or mannequin.

[0007] In a particular embodiment, a target coherence value is established to serve as a benchmark for a realistic, externalized sound field. This target, known as the diffuse-field interaural coherence, is calculated by averaging the interaural coherence values derived from a multiplicity of HRTFs, wherein each HRTF corresponds to a different direction of sound arrival. Said BRIR, SRIR, or a functionally similar representation of a reverberant environment, may be derived from one or more methods, including: a) direct physical measurement in a real environment; b) physics-based numerical simulation, such as methods based on virtual 3D models of an environment; c) microphone array processing and rendering techniques, such as the Spatial Decomposition Method (SDM), Spatial Impulse Response Rendering (SIRR), or Rendering from Arbitrary Microphone Array (REPAIR); or d) Learning-based modeling, wherein a model is trained to generate or predict an acoustic response.

[0008] For example, said learning-based modeling may include, but is not limited to, models such as Differentiable Feedback Delay Networks (FDNs) which are optimized to a target response, Generative FDNs (GFDNs) which generate a response based on a conditioning input, Neural Acoustic Fields (NAFs) which learn a continuous representation of a sound field, or Spatialized Scattering Delay Networks. Notably, certain parametric or generative models, such as Directional Feedback Delay Networks (DFDNs), can directly synthesize a spatially reverberant audio signal in response to an input, thereby obviating the need to explicitly generate, store, and convolve with a complete BRIR or SRIR data set. Once derived or otherwise obtained, the application of said one or more spatial transfer functions or the execution of said generative models to produce one or more binaural or spatial audio content may be performed by one or more real-time operations. Such operations include not only the primary convolution with an HRTF, BRIR, or SRIR, but may also include the direct synthesis from a parametric model or the application of said filter designed to steer the signal’s interaural coherence towards the target diffuse-field value to boost externalization. These operations may include, but are not limited to, convolution or one or more matrix operations.

[0009] For example, a position of client-side device relative to server-side device may be calculated using multilateration operation, angle of arrival, and angle of departure methods, and an other geometric positioning technique. For example, audio content may utilize one or more of one or more bandpass transfer function, headphone transfer function (HpTF), compensation filters, speed- adaptive equalization filters, loudness normalization, noise-cancellation, speech enhancement, voice isolator, Al-powered volume management, Al-based filters and an other DSP algorithm, so as to regularize audio signals, and adjust one or more timber, spectral cue, localization, and an other attribute of sound. For example, audio content may comprise one or more of one or more speech, ambient, music, sound effects, silence, noise, and an other sound.

[0010] For example, media content may comprise one or more of one or more audio and media content. For example, media content may comprise generated by Multimodal Al System. For example, media content may be digital content. For example, media content may be an other type of content other than digital content.For example, Multimodal Al Media Application may produce audio and media content. For example, Multimodal Al Media Application may produce compressed and uncompressed audio and media content. For example, Multimodal Al Media Application may translate media content from foreign and non-foreign languages. For example, media content may comprise facial animation generated by Multimodal Al Media Application. For example, media content may comprise of embed digital signatures. For example, alternative media content may comprise of media content with associated metadata describing one or more sound source positioning, sound pressure, multiplicity of audio objects and channel sound signals, and an other spatial property.

[0011] For example, audio and media content may be generated by one or more server-side device. For example, audio and media content may be generated by one or more client-side device. For example, audio and media content may be accessed by user using one or more client-side device and server-side device.

[0012] For example, audio and media content may be accessed by user using wire earbuds, wireless earbuds, wire headphones, wireless headphones, brain on a chip device, ear transducer, bone conductor transducer, and an other apparatus with one or more speakers. For example, audio and media content may be accessed by user using wire virtual reality (VR) headset, wireless VR headset, wire spatial computing headset, wireless spatial computing headset, wire computing glasses, wireless computing glasses, and an other apparatus with one or more speakers. For example, audio content may be accessed by user using speaker system. For example, audio content may be accessed by user using one or more of one or more server-side device and speaker system. For example, audio content may be accessed by user using one or more of one or more client-side device and speaker system.

[0013] For example, one or more wireless communication sensor may access and localize audio and media content using SPL. For example, two or more wireless communication sensor may be located on user’s head, on user’s head and in user’s ear, or on user’s head and in both user’s ears. For example, one or more wireless communication sensor may comprise one or more of an Ultra- Wideband (UWB) anchor, a Wi-Fi access point, a Bluetooth beacon, and an other transmitting and receiving device. For example, one or more wireless communication sensor may be at a location other than on user’s head, on user’s head and in user’s ear, or on user’s head and in both user’s ears. For example, three or more wireless communication sensor may improve the accuracy and resolution of direction in 3D space. For example, the resolution and sampling rate of Multimodal Al System recording may increase to improve accuracy. For example, one or more wireless communication sensor may be located on user’s person, headset, client-side device, and server-side device.

[0014] For example, audio and media content may be processed, analyzed, and modified using one or more of one or more application sensor. For example, an application sensor may comprise microphone, visual sensor, thermal sensor, vibratory sensor, camera, LIDAR, radar, position sensor, biomarker sensor, ambient light sensor (ALS), electrocardiogram sensor (ECG), and an other sensor. For example, position sensor may comprise one or more of one or more headmotion tracking sensors, infrared camera to recognize hand gestures, understand depth or map a room, eye motion tracking sensors, haptic sensors, magnetic sensors, and an other position sensor. For example, ALS may comprise time-of- day data. For example, ECG may comprise electrical activity of the heart. For example, application sensor may apply data to Multimodal Al System as data input for pattern recognition, predictive analysis, automated response, real-time monitoring, alerts and notifications, performance optimization, safety monitoring, user personalization, multi-sensor integration, behavior analysis, decision-making, deep learning, predictive health, and an other automated task.

[0015] For example, mixed audio content may expand, combine, edit, delete, fuse, transcribe, and the like one or more of one or more audio content track, ambient audio track, and data input to create one or more mixed audio content. For example, the mixed audio content may adjust the SPL, panning, speed- adaptive equalization (EQ), Al-powered volume management, effects, and dynamics of individual audio elements to create balance and polish.

[0016] For example, audio and media content may be accessed by user over a network. For example, audio and media content may be accessed by user over one or more of Internet, private virtual network, extranet, fiber optic network, wide area network (WAN), local area network (LAN), Bluetooth network, ultra-wideband network (UWB), RF, wired network, wireless network, near field magnetic inductance, sonic waves, ultrasound waves, infrared waves, wireless mesh protocol, satellite internet constellation system, and an other type of network. For example, audio and media content may be accessed by user via TCP, UDP, WebSocket, WebRTC connection, and an other real-time communication. For example, content may be accessed by user within virtual boundaries around a geographical area (geo-fencing). For example, the virtual boundaries may be coordinates of a real-world geographical area.

[0017] Embodiments are described for medium, method, and system for realtime multi-object tracking, synchronization, and spatialization using Multimodal Al System to process and integrate in real-time audio and media content in, to, and from one or more wireless communication sensor and electronic device. For example, Multimodal Al System comprises multiple sensor data, applications, models, and data input to generate rich contextual outputs that enhance complex, comprehensive, unifying, and nuanced understanding and accuracy. For example, Multimodal Al System may generate more adaptive, natural, context-aware, and human-like interactions using Multimodal Al System Application. For example, Multimodal Al System Application may comprise of one or more of client-side voice-activated Al application, voice-to-text Al application, audio-to-client-device application, authentication Al application, on-device Al foundation model, and an other application. For example, Multimodal Al System Application may comprise of one or more of server-side text-to-speech Al application, text-to-audio Al application, audio-to-dialog Al application, location Al application, dialog supervisor Al application, privacy guardrail Al application, marketing Al application, server-based Al model, and an other application.

[0018] As used herein, the term "natively Multimodal Al model" refers to a single, unified artificial intelligence architecture that processes input streams ofmultiple data modalities (such as audio and sensor data) and directly generates an output stream without conversion to an intermediate, human-readable format like transcribed text. This end-to-end processing is distinct from prior art "cascaded" systems that use a pipeline of separate unimodal models (e.g., speech-to-text followed by text-to-speech). The native model achieves this efficiency by converting all input modalities into a shared mathematical latent space, where cross-modal reasoning occurs before the output is directly decoded into the desired modality, thereby preserving nuances like emotional tone that are lost in text conversion.

[0019] For example, Multimodal Al System Application is architected for low- latency, bidirectional data streaming, and may comprise one or more of streaming text-to-speech (TTS) engine for dynamically generating audio from text inputs as they are received, and real-time speech-to-text (STT) transcription engine for converting live audio inputs into a corresponding text stream. For example, in another real-time embodiment, the Multimodal Al System Application may use a Multimodal transformer model that processes one or more of one or more physics-informed text, verbal, gesture, touch, vital, facial, visual, audio, and an other data inputs natively.

[0020] For example, Multimodal Al System Application may use generative Al, natural language processing (NLP), Transformers, Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), Diffusion Models, Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), Transformers, Physics Informed Neural Networks (PINN), Large Behavior Models (LBM), Large Language Models(LLM), mixture-of-experts (MoE), quantized model, machine learning Al application and the like, to generate informed, integrated, cohesive, insightful, and leveraged data across various modalities.

[0021] In particular, according to further embodiments of the invention, one or more of one or more audio content, media content, mixed audio content, and an other data input synchronize in real-time with static and non-static on-screen media produced by Multimodal Al Media Application. For example, on-screen media may be projected using one or more digital projector, 3D digital display, 3D projection mapping, multi-view stream display, hologram projector, led screen, billboard, robotics, vehicle, computer, television, console, headset, and an other projection device. For example, on-screen media produced may be using one or more lipsyncing general-purpose robotic humanoid, an other robot with a video display, and an other reception device. For example, one or more of one or more speaker and microphone may be integrated into various locations on a humanoid robot, said locations comprising the head and other non-head portions of the robot's chassis. For example, synchronization is performed by one or more of one or more computer algorithm, clock, and sensor. Alternatively, or additionally, Multimodal Al System Application receives synchronization input from one or more of one or more user and designer.

[0022] According to further embodiments of the invention, Multimodal Al Media Application may monitor, expand, combine, edit, delete, fuse, transcribe, and the like, audio content, media content, mixed audio content, and an other data input. For example, Multimodal Al System Application generates communicationbetween server-side and client-side device to produce real-time on-going dialog within and around one or more of venue, public, private, office, open space, educational institution, store, storefront, home, vehicle, chat, gaming, virtual, and an other environment. For example, Multimodal Al System Application may comprise one or more location Al application. For example, location Al application tracks multitude of client-device to determine the near-precise location of each client device controlled by a user in public environment. For example, location Al application may comprise one or more of one or more cellular network, Wi-Fi signals, GPS, satellite internet constellation system, UWB, gyroscope, barometer, accelerometer, raw sensor data, and an other position locator to assist in determining location of server-side and client-side device. According to yet other embodiments of the invention, the communication environment may include multiple venues within different locations and / or time zones viewing the same media content simultaneously, maybe even synchronously.

[0023] Furthermore, the input signals may comprise data from a wireless local area network (WLAN), including signal characteristics received from a plurality of radio links across different frequency bands, such as 2.4 GHz, 5 GHz, and 6 GHz, managed through a protocol that enables Multi-Link Operation (MLO). The unique signature of signal strengths, latencies, and interference levels across these multiple concurrently managed links can be used as a distinct input to enhance the accuracy and reliability of the context determination.

[0024] According to other embodiments of the invention, the communication environment is configured to leverage said MLO to provide high-performance data transmission between device. MLO enables client device and a network access point to establish and concurrently utilize multiple radio links. This capability is harnessed by the system to facilitate demanding use cases, such as the synchronized delivery and presentation of media content to multiple venues, which may be situated in different physical locations and / or time zones. The system may be configured to dynamically manage data flows across these multiple links to optimize for specific performance objectives, such as throughput, latency, or reliability. In further embodiments, the system is configured to select and operate in one of a plurality of MLO modes based on application requirements or real-time network analysis. These modes may include: a) a Simultaneous Multi-Link (SML) mode, wherein the system is configured to aggregate the bandwidth of at least two of the established radio links by transmitting and / or receiving data packets simultaneously across them. This mode is preferentially employed for applications requiring maximum data throughput and minimal latency, such as the synchronized streaming of ultra-high-definition, for example, 8K or higher, media content; b) an Alternating Multi-Link (AML) mode, wherein the system is configured to monitor the performance of the established radio links and dynamically select a single optimal link for data transmission at any given moment. The system can perform seamless, near-instantaneous switching between links to avoid transient interference or congestion. This mode is preferentially employed to ensure robust and uninterrupted connectivity, which is critical for maintaining near-precise synchronization of media playback across all participating venues, therebypreventing desynchronization events caused by network degradation on any single frequency band.

[0025] For example, Multimodal Al System Application may comprise privacy guardrail Al application monitored by one or more of application, employees, contractors, and the like within and around one or more of venue, public, private, office, open space, educational institution, hospital, museum, stadium, arena, concert, sport event, convention center, hotel, store, storefront, home, vehicle, chat, gaming, virtual, and an other environment. For example, Multimodal Al System Application may comprise one or more authentication Al application. For example, authentication Al application may comprise one or more of one or more physical biometric, behavior biometric, face identification, finger identification, eye identification, and an other identifier.

[0026] For example, marketing Al application may generate real-time multiobject tracking, synchronization, and spatialization of advertisement and promotion to integrate with audio and media content. For example, marketing Al application may comprise one or more of one or more audience segmentation monitoring, behavior targeting, personalization, and customization Al application, programmatic Al application, analytics Al application, fraud prevention Al application, and an other function of marketing. For example, marketing Al application determines user demographics, interests, sentiments, behaviors and other criteria to predict the effectiveness of advertisement and promotion for a desired outcome. For example, marketing Al application automates ad inventory management, strategy optimization, churn prevention, A / B testing, product review, social media posting, marketing accounting, and an other task automation.

[0027] According to further embodiments of the invention, audio content, media content, mixed audio content, and an other data input may be generated by user’s text commands, verbal commands, gesture-based commands, touch-based commands, vital signs, facial expressions, eye position, and an other signal. According to yet other embodiments of the invention, audio content, media content, mixed audio content, and an other data input may be processed, saved, stored, and used to train and finetune Multimodal Al System, Multimodal Al System Application, Multimodal Al Media Application, Al Audio Mixing System and Multimodal Deconstruction Al System. These features are available in one or more in-venue settings, on-edge settings, remote settings, and others. These features are available on multiple platforms and device. For example, the on-edge setting may receive and process the commands on the same electronic device. For example, the remote settings may receive and process the commands on different electronic device. For example, client-side device and server-side device may receive and process commands on the same electronic device. For example, communication between client-side device and server-side device may comprise one or more of one or more unicast transmission, multicast transmission, and broadcast transmission.

[0028] According to yet further embodiments of the invention, Al Audio Mixing System is configured to receive and demultiplex a data stream originating from a remote server-side device, said stream comprising a plurality of discrete audio subtracks and a corresponding stream of semantic event markers associated with aparticipant in a live event. The engine utilizes said processor to execute instructions that continuously process the stream of semantic event markers to interpret real-time narrative context of the participant's actions and physiological state. In response to the interpreted narrative context, the engine is further configured to dynamically and automatically adjust a set of mixing and rendering parameters for each of the discrete audio sub-tracks, said parameters comprising one or more of gain, equalization, audio compression, and the selective application of spatialization filters. The engine thereby generates a final, coherent binaural audio output that is not a static representation of the captured audio, but is instead a context-aware auditory experience that is cinematically adapted in real-time to heighten the dramatic effect of the actions and states described by the semantic event markers. According to yet further embodiments of the invention, a Multimodal Deconstruction Al System may be one or more of a specialized deep learning architecture, for example, a transformer-based or a hybrid convolutional- recurrent neural network (CRNN) architecture, specifically designed for real-time, on-device processing. The model is configured to receive and process a plurality of heterogeneous, time-synchronized data streams simultaneously, said streams comprising: the continuous audio stream of a user's verbal command received from client device, the raw audio stream captured by server-side device's own microphones, a time-series data stream from an integrated Inertial Measurement Unit (IMU), and a time-series data stream from one or more integrated biometric sensors, such as an Electrocardiogram (ECG) sensor. A key aspect of the model is its ability to directly map the user's vocal command waveform to an internal intent vector without an intermediate text transcription step, thereby reducing latency and preserving vocal nuance. Concurrently, the model performs real-time deconstruction of the participant’s reality by utilizing two integrated sub-networks: a source separation sub-network that deconstructs the participant's raw microphone audio into a plurality of discrete audio content sub-tracks, such as ‘voice,’ ‘footsteps,’ and ‘impacts’; and an event detection sub-network that processes the IMU and biometric sensor data to identify and classify significant occurrences, generating a corresponding stream of semantic event markers. The model then utilizes the derived user intent vector to selectively package or prioritize the generated sub-tracks and event markers into the final, high-information-density data package that constitutes server-side response.

[0029] According to yet further embodiments of the invention, memory data store information may comprise one or more of one or more audio content, media content, mixed audio content, and other data input before, during, and after user’s interaction with one or more Multimodal Al System. For example, memory data store may comprise server-side master application, client-side master application, server-side streaming application, client-side Multimodal Al System Application, client-side Multimodal Al Media Application, server-side Multimodal Al System Application, server-side Multimodal Al Media Application, client-side Al Audio Mixing System, server-side Multimodal Deconstruction Al System and an other application. For example, memory data store may be located in one or more clientside device. For example, memory data store may be located in one or moreserver-side device. For example, memory data store may be located in one or more of one or more client-side device and server-side device.

[0030] For example, server-side master application and client-side master application are designed for synchronizing and managing data network packages and acting as the central coordinator in a distributed system. Server-side master application and client-side master application ensure data consistency, reliability, and efficient communication across client device, servers, and other device. For example, server-side master application and client-side master application may manage the cache, encoding, decoding, compression, and encryption of data packages to optimize network usage and secure data in transit.

[0031] For example, server-side streaming application may be responsible for ingesting audio content, media content, mixed audio content, and other data inputs, processing, and streaming to client device while ensuring synchronization. For example, server-side streaming application may encode, buffer, transcode, distribute content, and generate real-time synchronization. For example, serverside streaming application may stream multiple copies of the data for redundancy.

[0032] According to further embodiments of the invention, the memory data store may track and store the state and action of the user. For example, the state represents a user’s current situation or configuration using client device at a given time. For example, a state may comprise past purchase activity data, present purchase activity data, stock market data, medical data, foreign language command, or an other state. For example, present purchase activity data may comprise strokes and clicks client device controlled by a user records as the user selects, stores, and manages items in an e-commerce cart intended for purchase. For example, a state may comprise environmental conditions data, location tracking data, health bio signs data, and text sentiment data of the user. For example, an action may comprise data regarding the decision and the steps taken by the user using client device as a result of a prompt presented by the Multimodal Al System application. For example, a dialog supervisor Al application manages the context, strategy, and dialog flow with the user. For example, a dialog supervisor Al application may comprise intent recognition, context management, response generation, decision making, flow control, error handling, policy management, customer support, natural language understanding (NLU), natural language generation (NLG), personalization, e-commerce, and other dialog functions. For example, a dialog supervisor Al application may recommend that the user use client device to relocate to an alternative location, close a purchase order, encourage the user’s mood, or take another action to achieve a specific objective. For example, a dialog supervisor Al application may offer a reward to a user who achieves a desired outcome influenced by a dialog supervisor Al application prompt. For example, a dialog supervisor Al may process one or more user preferences to automatically select a language for the media content, wherein said selection comprises performing real-time translation of the media content or selecting a version of the media content in a preferred language.

[0033] According to further embodiments of the invention, dialog supervisor Al application may comprise generative Al, natural language processing (NLP), Transformers, Generative Adversarial Networks (GANs), VariationalAutoencoders (VAEs), Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), Transformers, Physics Informed Neural Networks (PINN), Large Behavior Models (LBM), Large Language Models(LLM), mixture-of-experts (MoE), quantized model, machine learning Al application, and the like that manages conversation of state and actions, and an other algorithm that manages conversation. For example, a dialog supervisor Al application may ask to determine user intent, query the database, provide recommendations based on profile and memory data, analyze sensor data, identify data required to act, track, execute transactions, provide answers, and perform other tasks. For example, a dialog supervisor Al application may comprise multiple agents that make multiple calls to the same dialog supervisor Al application.

[0034] In the context of the present disclosure, the "first delay" represents the acoustic propagation time, which is a calculated simulation of the real-world time required for a sound wave to travel through a physical medium, such as air, from the location of a sound source, such as server-side device, to the location of a listener, such as client-side device. The primary purpose of calculating and applying this first delay is to provide a powerful psychoacoustic cue for distance perception, thereby creating a realistic and immersive spatial audio experience. The determination of this first delay may be accomplished by one or more of the following methods.

[0035] In one embodiment, the first delay is determined indirectly via a modelbased calculation. This method leverages the location coordinates of the device. First, client-side device calculates the straight-line Euclidean distance (D) between its own real-time coordinates (Xc,Yc,Zc) and the coordinates of server-side device (Xs,Ys,Zs). The first delay is then calculated by dividing the distance (D) by the speed of sound (vsound). In some implementations, a standard value for the speed of sound, such as 343 meters per second, is used. In other, more advanced implementations, client-side device may be configured to use onboard sensors, such as a thermometer, to determine local environmental conditions and select a more accurate value for the speed of sound. The accuracy of the first delay calculated by this method is directly dependent on the accuracy of the underlying positioning system providing the coordinates.

[0036] In another embodiment, the first delay is determined by a direct active measurement using a reference signal. In this method, server-side device is configured to emit a dedicated, known reference sound signal, which may be an audible tone or an inaudible ultrasonic pulse. At the moment of emission, serverside device records an emission timestamp (Temit) and transmits this timestamp to client-side device over a low-latency data channel. Client-side device's microphone continuously listens for said reference signal. Upon detection, clientside device records an arrival timestamp (Tarrive) on its own synchronized clock. The first delay is then calculated directly as the difference between the arrival and emission timestamps (FirstDelay=Tarrive-Temit). This method provides a direct measurement of the actual acoustic travel time within the specific physical environment.

[0037] In a further embodiment, the first delay is determined by a direct passive measurement using cross-correlation. In this method, client-side device comparesthe "clean" audio track received from server-side device (the source signal) with the audio that is contemporaneously recorded by its own microphone from the physical environment (the received signal). Client-side device's processor is configured to perform a mathematical cross-correlation between the source signal and the received signal. The output of the cross-correlation function reveals the time offset at which the two signals are most similar. This time offset corresponds directly to the acoustic propagation time, or the first delay.

[0038] Regardless of the method used for its calculation, the resulting first delay value is subsequently used in the audio processing pipeline. It is applied to the final binaural stereo output, often in conjunction with loudness compensation signals, to render a believable sense of distance. Furthermore, the first delay is a critical component of a final compensating delay, wherein it is combined with a calculated second delay (corresponding to processing latency) to ensure the final audio playback is aligned with the system's synchronized event timeline, thereby creating an audio experience that is both physically realistic and technically synchronized.

[0039] According to further embodiments of the invention, server-side device for real-time multi-object tracking, synchronizing, and spatializing of one or more audio and media content using Multimodal Al System and wireless communication sensor includes: processor; data storage operably connected with processor; memory operably connected with processor, memory comprising one or more of server-side master application, server-side streaming application, server-side Multimodal Al Media Application, server-side Multimodal Al Media System and server-side Multimodal Deconstruction Al System; server-side playback device and server-side sensor system operably connected with processor; server-side computing device operably connected with processor and configured to communicate over one or more networks and input / output (I / O) device with one or more sensor systems, speaker systems, wireless communication sensor systems, and a an other audio transducer; server-side local interface operably connected with processor and configured to communicate over network with client-side device under user's control; server-side master application configured to receive message or packet comprising media content selected by the user over the network from client-side device; server-side master application configured to obtain user’s preferences; server-side master application configured to generate personalized audio and media content using user preferences; server-side master application configured to receive and process real-time data packets; server-side local interface configured to transmit to client-side device via the network both generated audio and media content and server-side timing information, wherein said timing information and content enable client-side device to substantially synchronize and spatialize its playback with corresponding playback by serverside playback device.

[0040] It is to be understood that the embodiments described herein are illustrative and not restrictive. The scope of the invention is defined by the appended claims rather than the foregoing descriptions, and all changes that come within the meaning and range of equivalency of the claims are therefore intended to be embraced therein. The various functions and processes described hereinmay be implemented in hardware, software, or a combination thereof. When implemented in software, the functions may be stored on or transmitted over as one or more instructions or code on a non-transitory computer-readable medium.DESCRIPTION OF THE DRAWINGS

[0041] FIGURE 1 is a schematic block diagram of a networked environment for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor.

[0042] FIGURE 2 is a schematic block diagram of server-side computing device in an alternative configuration of a networked environment for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor.

[0043] FIGURE 3 illustrates a system comprising: one or more client-side device, wherein at least one of the one or more client-side device is embodied as earbuds configured to be worn by a human user; and server-side device embodied as a humanoid robot, wherein the one or more client-side device are configured to synchronize with server-side device.

[0044] FIGURE 4 is a flowchart for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 4 applies to a method using Multimodal Al, as viewed from client-side.

[0045] FIGURE 5 is a flowchart for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 5 applies to a method using Multimodal Al, as viewed from server-side.

[0046] FIGURE 6 is a flowchart for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 6 applies to a method using a single, natively Multimodal Al, as viewed from client-side.

[0047] FIGURE 7 is a flowchart for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 7 applies to a method using a single, natively Multimodal Al, as viewed from server-side.

[0048] FIGURE 8 is a flowchart for client-side device to perform tracking, synchronization, and spatialization of media content of client-side playback with server-side playback using a Multimodal Al System and wireless communication sensors. Server-side device is configured to generate audio content track and ambient audio track, and one or more mixed audio content. FIGURE 8 applies to a method using Multimodal Al, as viewed from server-side.

[0049] FIGURE 9 is a flowchart for client-side device to perform tracking, synchronization, and spatialization of media content of client-side playback with server-side playback using a Multimodal Al System and wireless communication sensors. This method enables client-side device to synchronize its media playback with server-side device in both time and 3D space. It calculates network latency and the physical positions of both device, then uses this data to apply directionalaudio filters. The result is a near-perfectly timed, spatialized audio experience that sounds as if it is coming from the server's precise physical location after compensating for all acoustic and processing delays. FIGURE 9 applies to a method using Multimodal Al, as viewed from client-side.

[0050] FIGURE 10 is a flowchart for client-side device to perform tracking, synchronization, and spatialization of media content of client-side playback with server-side playback using a Multimodal Al System and wireless communication sensors. This method enables server-side device to synchronize its media playback with client-side device in both time and 3D space. It calculates network latency and the physical positions of both device, then uses this data to apply directional audio filters. The result is a near-perfectly timed, spatialized audio experience that sounds as if it is coming from the server's precise physical location after compensating for acoustic and processing delays. FIGURE 10 applies to a method using Multimodal Al, as viewed from server-side.

[0051] FIGURE 11 is a flowchart for client-side device performs spatializing media content in a two-dimensional (2D) plane. FIGURE 11 applies to a method using Multimodal Al, as viewed from client-side.

[0052] FIGURE 12 is a flowchart for dynamically updating spatialization of media content on client-side device. FIGURE 12 applies to a method using Multimodal Al, as viewed from client-side.

[0053] FIGURE 13 is a flowchart for dynamically updating the spatialization of media content on client-side device to track a moving server-side device. FIGURE 13 applies to a method using Multimodal Al, as viewed from client-side.

[0054] FIGURE 14 is a flowchart for predictive spatialization of media content on client-side device. FIGURE 14 applies to a method using Multimodal Al, as viewed from client-side.

[0055] FIGURE 15 is a flowchart for predictive spatialization of multi-object track media content on client-side device using a Multimodal Al System and wireless communication sensors. FIGURE 15 applies to a method using Multimodal Al, as viewed from client-side.

[0056] FIGURE 16 is a flowchart for improving location accuracy via fusion for real-time spatialization. FIGURE 16 applies to a method using Multimodal Al, as viewed from client-side.

[0057] FIGURE 17 is a flowchart for adaptively selecting a positioning protocol for real-time spatialization. FIGURE 17 applies to a method using Multimodal Al, as viewed from client-side.

[0058] FIGURE 18 is a flowchart for a method for improving the accuracy of acoustic measurements for real-time spatialization. FIGURE 18 applies to a method using Multimodal Al, as viewed from client-side.

[0059] FIGURE 19 is a flowchart for a method for applying a spatialization filter to a plurality of audio system formats includes. FIGURE 19 applies to a method using Multimodal Al, as viewed from client-side.

[0060] FIGURE 20 is a flowchart for a method for processing streamed media content on client-side device. FIGURE 20 applies to a method using Multimodal Al, as viewed from client-side.

[0061] FIGURE 21 is a flowchart for a method for processing streamed media content on server-side device. FIGURE 21 applies to a method using Multimodal Al, as viewed from server-side.

[0062] FIGURE 22 is a flowchart for a method for processing streamed media content on client-side device. FIGURE 22 applies to a method for synchronizing with a server and rendering, position-aware audio on client-side device.

[0063] FIGURE 23 is a flowchart for a method for processing streamed media content on client-side device. FIGURE 23 applies to a method for dynamically reconfiguring a multi-speaker audio system based on real-time position data from client device.

[0064] FIGURE 24 is a flowchart for a method for processing and demultiplexing a data stream. FIGURE 24 applies to a method for rendering dynamic, context-aware binaural audio using client-side Al Audio Mixing System from client device.

[0065] FIGURE 25 is a flowchart for a method for processing and performing semantic event detection to generate corresponding stream of event markers describing a participant's actions. FIGURE 25 applies to a method for generating a high-information-density data stream for remote auditory rendering using serverside Multimodal Deconstruction Al System from server device.

[0066] FIGURE 26 is a flowchart for a method for rendering synchronized, multi-sensory haptic and auditory experience. FIGURE 26 applies to a method for rendering synchronized, multi-sensory haptic and auditory experience using client-side Al Sensory Experience Application, from client device.

[0067] FIGURE 27 is a flowchart for a method for generating high-information- density data stream for enabling a remote multi-sensory experience. FIGURE 27 applies to a method for generating high-information-density data stream for enabling a remote multi-sensory experience using server-side Al Sports Analysis Application from server device.

[0068] FIGURE 28 is a flowchart for a method for rendering real-time, spatialized, and voice-preserving translated audio experience. FIGURE 28 applies to a method for using client-side Multimodal Al Media Application to receive audio stream in a source language, perform live translation into a target language while preserving the vocal identity of the original speaker, and render the resulting translated audio as a fully synchronized and spatialized experience on client device.

[0069] FIGURE 29 is a flowchart for a method for enabling remote, real-time, spatialized audio translation. FIGURE 29 applies to server-side method for generating and transmitting data stream to client device, wherein said data stream comprises a primary audio track in a source language, corresponding real-time spatialization metadata, and the necessary timing data to allow the remote client device to perform synchronized, spatialized, and translated rendering.DETAILED DESCRIPTION

[0070] While the present invention is susceptible of embodiment in many different forms, there is shown in the drawings and will herein be described in detail one or more specific embodiments, with the understanding that the present disclosure is to be considered as exemplary of the principles of the invention and not intended to limit the invention to the specific embodiments shown and described. In the following description and in the several figures of the drawings, like reference numerals are used to describe the same, similar or corresponding parts in the several views of the drawings.

[0071] The system and method for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor includes a plurality of components such as one or more of electronic components, hardware components, and computer software components. A number of such components can be combined or divided in the system. An example component of the system includes a set and / or series of computer instructions written in or implemented with any of a number of programming languages, as will be appreciated by those skilled in the art.

[0072] The system and method for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor includes a plurality of components such as one or more input and output (I / O) device, analogue to digital and digital to analogue device (AD / DA), amplifier device, active and passive speaker systems, and an other electronic device.

[0073] The system in one example employs one or more computer-readable signal-bearing media. The computer-readable signal bearing media store software, firmware and / or assembly language for performing one or more portions of one or more implementations of the invention. The computer-readable signalbearing medium for the system in one example comprises one or more of a magnetic, electrical, optical, biological, and atomic data storage medium. For example, the computer-readable signal-bearing medium comprises floppy disks, magnetic tapes, CD-ROMs, DVD-ROMs, hard disk drives, downloadable files, files executable “in the cloud,” and electronic memory.

[0074] FIGURE 1 is a schematic block diagram of a networked environment for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor via one or more of one or more electronic device, that comprises client-side networked environment 105, server-side networked environment 110, and network 115. Network 115 comprises one or more of one or more speaker system 135, microphone system 137, wireless communication sensor 165, input and output (I / O) device 198, Internet, private virtual network, extranet, fiber optic network, wide area network (WAN), local area network (LAN), wired network, wireless network, a satellite internet constellation system and an other type of network.

[0075] Client-side networked environment 105 comprises client-side device 120 and client-side playback device 125 that is operably connected with client-side device 120. Client-side device 120 comprises, for example, one or more of one ormore tablet 120, phone 120, smart device 120, virtual reality headset 120, wireless virtual reality headset 120, augmented reality headset 120, wireless augmented reality headset 120, wire computing glasses 120, wireless computing glasses, computer program 120, computer browser 120, media player 120, game console 120, virtual device 120, neuro link device 120, wearable device 120 and an other computing device 120.

[0076] Client-side device 120 runs one or more applications. Client-side device 120 deploys over network 115.

[0077] Client-side playback device 125 is configured to play media content. For example, client-side playback device 125 plays media content received from client-side device 120. Alternatively, or additionally, client-side playback device 125 plays media content received directly over network 115. For example, clientside playback device 125 comprises one or more of one or more headphone 125, earphone 125, earbud 125, earworn wearable 125, screen 125, television 125, monitor 125, in-venue projector 125, home theater 125, three-dimensional digital projector 125, and an other client-side playback device 125. For example, clientside playback device 125 comprises one or more of one or more open headphone 125, semi-open headphone 125, closed headphone 125, and an other type of headphone 125.

[0078] Client-side playback device 125 operates in an environment with clientside sensor system 130. For example, sensor system 130 comprises one or more of one or more visual sensor 130, microphone 130, thermal sensor 130, vibratory sensor 130, position sensor 130, camera 130, LIDAR 130, radar 130, biomarker sensor 130, ambient light sensor (ALS), electrocardiogram sensor (ECG), and an other sensor 130. For example, the sensor system 130 comprises one or more of one or more head motion tracking sensor 130, eye motion tracking sensor 130, infrared camera to recognize hand gestures, understand depth or map a room 130, haptic sensor 130, magnetic sensor 130 and an other position sensor 130. For example, the sensor system 130 may comprise a spatial resolution of three degrees of freedom in all directions for head tracking and six degrees or more for motion tracking so as to provide one or more smooth audio transitions and stable sound imaging.

[0079] Client-side device 120 is configured to generate a dynamically updated spatial audio experience based on the relative positions of the client and a server. A key aspect of this process is the calculation of an operative direction vector, which dictates the perceived direction of the audio source. This operative vector is intelligently derived to accommodate scenarios both with and without the use of head tracking. The system first establishes a baseline direction vector representing the direct line-of-sight between the client and server device. This baseline vector may be used directly as the operative direction vector for spatialization. Alternatively, in embodiments where the user's head orientation is monitored by a sensor, the system applies a rotation transformation to the baseline vector based on the real-time head orientation data. The resulting rotated vector then becomes the operative direction vector. This dual-path approach ensures a robust and seamless spatial audio experience whether the head tracking feature is active, disabled, or unavailable.

[0080] Client-side playback device 125 is configured to operate in an environment having a speaker system 135. Said speaker system 135 may be any of a plurality of types, including, but not limited to, a channel-based system, an object-based system, a scene-based system, or another three-dimensional (3D) audio system. For instance, a channel-based system may be a single-channel or multi-channel configuration comprising one or more speakers placed in designed positions within a physical space. An object-based system may utilize audio objects, such as sound sources, that contain metadata describing the intended position and other spatial properties of the object. A scene-based system may utilize a sound field representation, such as one or more spherical harmonic basis functions, that describes how sound pressure changes as a function of time and direction. The physical speaker system 135 may comprise one or more audio transducers, such as woofers, to reproduce a low-frequency range, and tweeters, to reproduce a high-frequency range.

[0081] Client-side playback device 125 may operates in an environment with microphone system 137. For example, microphone system 137 may comprise one or more of one or more microphone array, spot microphone, and an other microphone system.

[0082] Client-side device 120 comprises one or more client-side memory 140 and client-side data storage 145.

[0083] Client-side memory 140 is defined herein as including both volatile and nonvolatile memory and data storage components. For example, client-side memory 140 may comprise one or more sequential access. For example, clientside memory 140 comprises one or more client-side buffers. Volatile components are those that do not retain data values upon loss of power. Nonvolatile components are those that retain data upon loss of power. For example, clientside memory 140 may comprise one or more random access memory (RAM), read-only memory (ROM), hard disk drive, solid-state drive, USB flash drive, memory card, floppy disk, optical disc such as compact disc (CD) or digital versatile disc (DVD), magnetic tape, and other memory components. For example, RAM may comprise one or more static random access memory (SRAM), dynamic random access memory (DRAM), magnetic random access memory (MRAM), and other forms of RAM. For example, ROM may comprise one or more programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and other forms of ROM.

[0084] Client-side memory 140 comprises one or more client-side master application 150, client-side Multimodal Al Media Application 152, client-side Multimodal Al System Application 160, client-side Al Audio Mixing System 162, and client-side Al Sensory Experience Application 163.

[0085] Client-side memory 140 further comprises client-side device unique identifier. Client-side device unique identifier is a number unique to this particular device. In other words, each device in the world will have its own number that no other such device will have. A copy of client-side device’s unique identifier, known as client-side unique identifier, will be transmitted by client-side device 120 in a message or packet to server-side computing device 170. Then, a copy of client-side transmitted unique identifier, known as server-side unique identifier, will be transmitted back from server-side computing device 170 to client-side device 120. Server-side unique identifier received by client-side device 120 will then be compared with client-side device’s unique identifier to help determine the integrity of the messages and as a security check.

[0086] Optionally, client-side memory 140 further comprises an other clientside application (not pictured). The other client-side application comprises one or more additional client-side application, additional client-side service, additional client-side process, and additional client-side functionality. For example, an other client-side application runs background services. For example, an other client-side application runs boot processes. For example, an other client-side application runs other client-side applications.

[0087] Client-side data storage 145 comprises one or more single database, multiple database, cloud application platform, relational database, no-sequel database, flash memory, solid state memory, and an other client-side data storage device. Client-side data storage 145 may be located in a single installation that may be local to server-side computing device 170. Alternatively, client-side data storage 145 may be located in a single installation that may be local to client-side device 120. Alternatively, client-side data storage 145 may be distributed in a plurality of locations. Client-side data storage 145 may be distributed in a plurality of geographical locations. Client-side data storage 145 may be distributed in a plurality of geographical locations located in the same time zone. Client-side data storage 145 may be distributed in a plurality of geographical locations, wherein not the geographical locations are located in the same time zone.

[0088] Client-side data storage 145 comprises one or more of item prices, order information, media content, and other information. For example, the media content comprises one or more of a live broadcast, simultaneously recorded broadcast, audio track, a video track, an other media track, a motion picture, a commercial, a motion picture trailer, a demonstration (“demo”), a commentary, extra content, and an other form of additional content. The media content comprises one or more of media data, media content files, and other media content. The motion picture comprises one or more of a feature-length theatrical production, short-film production, an animated production, a broadcast television production, live events, a pay television production, a documentary, a commercial, a trailer, and an other motion picture. The media data comprises one or more audio track, multi-channel track, commentary, and other media data. The audio track comprises one or more English language audio track, audio track in a language other than English, and customized audio track. The commentary comprises one or more commentary by one or more directors of a motion picture, a commentary by one or more actors in a motion picture, a commentary by contributors to a motion picture other than the directors and actors, and commentary by persons other than contributors to a motion picture.

[0089] Client-side master application 150 is configured to store playable media content, such as segmented or non-segmented media tracks, in client-side data storage 145. Optionally, the application 150 also performs media processing on the playable content. This processing may involve passing the content through oneor more digital signal processing (DSP) algorithms — such as bandpass transfer functions, headphone transfer functions (HpTFs), compensation filters, speed- adaptive equalization filters, or Al-based filters — to regularize audio signals or to adjust sound attributes including timbre, spectral cues, and localization. Specifically, an HpTF may be applied to smooth, accentuate, or cancel frequency response fluctuations for a particular pair of headphones. The application 150 may automatically upload a predetermined HpTF to a connected pair of headphones or, upon receiving headphone data such as a model, serial number, or other identifying information, may select and upload a customized HpTF from client-side data storage 145. Furthermore, the application 150 may determine a processing delay associated with these functions and adjust client-side playback timeline to compensate.

[0090] For example, client-side master application 150 parses the playable media content into a chronological sequence that substantially matches the sequence of the motion picture. For example, client-side master application 150 writes the playable media content to one or more of client-side data storage 145 and client-side memory 140. For example, client-side master application 150 writes the playable media content to a media content file located in one or more clientside data storage 145 and client-side memory 140.

[0091] For example, client-side Multimodal Al Media Application 152 may produce audio and media content.

[0092] For example, client-side Multimodal Al System Application 160 may comprise one or more client-side voice-activated Al application, voice-to-text Al application, audio-to-client-device application, authentication Al application, on- device Al foundation model, real-time Al application, and another application.

[0093] Client-side master application 150 is configured to connect with serverside networked environment 110 so as to substantially synchronize between server-side networked environment 110 and client-side device 120 media content played on client-side playback device 125. Client-side playback device 125 comprises one or more of one or more screen 125, television 125, monitor 125, cellular phone 125, laptop computer 125, desktop computer 125, notebook 125, tablet 125, headset 125, channel-based playback system 125, object-based playback system 125, scene-based playback system 125, 3d audio playback system 125, and an other client-side playback device 125. Client-side playback device 125 plays for the user one or more audio media content, video media content, and an other form of media content. For example, networked environment 115 may be synchronized with other sensory experiences such as, for example, one or more smoke effects, fire effects, lasers, fireworks, drones and water droplets, moving chairs, and the like. For example, more than one client-side playback device 125 may be used simultaneously.

[0094] As explained below in greater detail, particularly in Figure 9, client-side master application 150 is configured to perform one or more of sampling and recording client-side running media play time (CRMPT) at which client-side media player plays the media on the client. The CRMPT is defined as the elapsed running time for customized media content that is being played by client-side media player on the client. If no customized media content is being played by client-side mediaplayer, the CRMPT is defined as zero. The CRMPT recorded by client-side master application 150 represents a real world time value based on the host system clock of the client. Then, client-side master application 150 creates client-side message or packet that it transmits to server-side networked environment 110.

[0095] As explained below in greater detail, particularly in Figure 12, client-side Multimodal Al System Application 160 is configured to determine by client-side device, its own position (xc, yc, zc) via multilateration using three or more wireless communication sensors 165 or alternatively, may be accessed as a predetermined value corresponding to a known, fixed location.

[0096] Server-side networked environment 110 comprises server-side data storage 155, server-side computing device 170 that is operably connected with server-side data storage 155, and server-side playback device 196 that is operably connected with server-side data storage 155. Server-side data storage 155 is a second location where, as mentioned above in relation to client-side data storage 145, client-side master application 150 may store the playable media content.

[0097] Server-side data storage 155 comprises one or more of item prices, order information, media content, and other information. For example, the media content files comprise one or more of one or more audio track, video track, an other media track, motion picture, commercial, motion picture trailer, demonstration (“demo”), commentary, extra content, and an other form of additional content. The media content comprises one or more media data, media content files, and other media content. The motion picture comprises one or more of one or more feature-length theatrical production, short-film production, animated production, broadcast television production, live events, pay television production, documentary, commercial, trailer, and an other motion picture. The media data comprises one or more of one or more audio track, multi-channel track, commentary, and other media data. The audio track comprises one or more of one or more English language audio track, audio track in a language other than English, and customized audio track. The commentary comprises one or more commentary by one or more director of a motion picture, commentary by one or more actor in a motion picture, commentary by contributors to a motion picture other than the directors and actors, and commentary by persons other than contributors to a motion picture.

[0098] Server-side data storage 155 comprises one or more single database, multiple database, cloud application platform, relational database, no-sequel database, flash memory, solid state memory, and an other server-side data storage device. Server-side data storage 155 may be located in a single installation that may be local to client-side device 120. Alternatively, server-side data storage 155 may be located in a single installation that may be local to server-side computing device 170. Alternatively, server-side data storage 155 may be distributed in a plurality of locations. Server-side data storage 155 may be distributed in a plurality of geographical locations. Server-side data storage 155 may be distributed in a plurality of geographical locations located in the same time zone. Server-side data storage 155 may be distributed in a plurality of geographical locations, wherein not the geographical locations are located in the same time zone.

[0099] Server-side computing device 170 comprises one or more server, computer, cloud-computing device, and distributed computing system.

[0100] Server-side computing device 170 may be located in a single installation. Alternatively, server-side computing device 170 may be distributed in a plurality of geographical locations. For example, server-side computing device 170 may be distributed in a plurality of geographical locations located in the same time zone. For example, server-side computing device 170 may be distributed in a plurality of geographical locations wherein not the geographical locations are located in the same time zone.

[0101] Server-side playback device 196 is configured to play media content. For example, server-side playback device 196 plays media content received from server-side computing device 170. Alternatively, or additionally, server-side playback device 196 plays media content received directly over network 115. For example, server-side playback device 196 comprises one or more of one or more in-venue projector 196, home theater 196, television 196, monitor 196, three- dimensional digital projector 196, and an other device 196.

[0102] Server-side playback device 196 is configured to operate in an environment having a speaker system 135. Said speaker system 135 may be any of a plurality of types, including, but not limited to, a channel-based system, an object-based system, a scene-based system, or another three-dimensional (3D) audio system. For instance, a channel-based system may be a single-channel or multi-channel configuration comprising one or more speakers placed in designed positions within a physical space. An object-based system may utilize audio objects, such as sound sources, that contain metadata describing the intended position and other spatial properties of the object. A scene-based system may utilize a sound field representation, such as one or more spherical harmonic basis functions, that describes how sound pressure changes as a function of time and direction. The physical speaker system 135 may comprise one or more audio transducers, such as woofers, to reproduce a low-frequency range, and tweeters, to reproduce a high-frequency range.

[0103] Server-side playback device 196 operates in an environment with server-side sensor system 197. For example, sensor system 197 comprises one or more of one or more visual sensor 197, microphone 197, thermal sensor 197, vibratory sensor 197, position sensor 197, camera 197, LIDAR 197, radar 197, biomarker sensor 197, ambient light sensor (ALS), electrocardiogram sensor (ECG), and an other sensor 197. For example, the sensor system 197 comprises one or more of one or more head motion tracking sensor 197, eye motion tracking sensor 197, infrared camera to recognize hand gestures, understand depth or map a room 197, haptic sensor 197, magnetic sensor 197 and an other position sensor 197. For example, the sensor system 197 may comprise a spatial resolution of three degrees of freedom in all directions for head tracking and six degrees or more for motion tracking so as to provide one or more smooth audio transitions and stable sound imaging.

[0104] Server-side playback device 196 is configured to communicate with server-side computing device 170. For example, server-side playback device 196 communicates with server-side computing device 170 using one or more of one ormore satellite, antenna, cable, network 115, and an other communication method. Server-side playback device 196 comprises one or more of one or more digital projector 196, hologram projector 196, led screen 196, screen 196, television 196, monitor 196, cellular phone 196, laptop computer 196, notebook 196, tablet 196, headset 196, multi-channel playback system 196, and an other server-side playback device 196.

[0105] Server-side computing device 170 comprises server-side memory 175. Server-side memory 175 is defined herein as including both volatile and nonvolatile memory and data storage components. For example, server-side memory 175 may comprise one or more sequential access. For example, server-side memory 175 comprises one or more server-side buffers. Volatile components are those that do not retain data values upon loss of power. Nonvolatile components are those that retain data upon loss of power. For example, server-side memory 175 may comprise one or more random access memory (RAM), read-only memory (ROM), hard disk drive, solid-state drive, USB flash drive, memory card, floppy disk, optical disc such as compact disc (CD) or digital versatile disc (DVD), magnetic tape, and other memory components. For example, RAM may comprise one or more static random access memory (SRAM), dynamic random access memory (DRAM), magnetic random access memory (MRAM), and other forms of RAM. For example, ROM may comprise one or more programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and other forms of ROM.

[0106] Server-side computing device 170 comprises one or more of serverside master application 180, server-side streaming application 185, server-side Multimodal Al media application 190, server-side Al Sports Analysis Application 194, server-side Multimodal Al System Application 195 and server-side Multimodal Deconstruction Al System 199. Server-side master application 185 is configured to provide synchronization timing information to one or more of client-side master application 150, and server-side streaming application 185.

[0107] Optionally, server-side computing device 170 further comprises an other server-side application (not pictured). The other server-side application comprises one or more of additional server-side application, additional server-side service, additional server-side process, and additional server-side functionality.

[0108] For example, the other server-side application runs background services. For example, the other server-side application runs boot processes. For example, the other server-side application runs other server-side application.

[0109] As explained below in greater detail, particularly in Figure 10, serverside master application 180 is configured to perform one or more of sampling and recording server-side running media play time (SRMPT) at which server-side media player plays the media on the server. The SRMPT is defined as an elapsed running time for customized media content that is being played by server-side media player on the server. If no customized media content is being played by server-side media player, the SRMPT is defined as zero. For example, the motion picture’s SRMPT time might clock in at 6 minutes, 10 seconds, and 10 frames. The SRMPT recorded by server-side master application 180 represents a real world time value based on the host system clock of the server. Then, server-side masterapplication 180 creates server-side message or packet that it transmits to clientside master application 150.

[0110] As explained below in greater detail, particularly in Figure 5, server-side master application 180 is configured to store media content. For example, serverside master application 180 stores playable media content in server-side data storage 155. The playable media content comprises one or more segmented media content track, non-segmented media content track, audio stimulus, and an other playable media content. Optionally, server-side master application 180 performs media processing of the playable media content.

[0111] Server-side master application 180 is configured to store playable media content to be played by server-side playback device 196. For example, server-side master application 180 stores the playable media content in serverside data storage 155. Optionally, server-side master application 180 performs media processing of the playable media content. For example, server-side master application 180 passes the playable media content through one or more bandpass transfer function, compensation filters, speed-adaptive equalization filters, Al- based filters and an other DSP algorithm, to regularize audio signals, and adjust one or more timber, spectral cue, localization, and an other attribute of sound. For example, server-side master application 180 parses the playable media content into a chronological sequence that substantially matches the sequence of the motion picture. For example, server-side master application 180 writes the playable media content to one or more server-side data storage 155 and serverside memory 175. For example, server-side master application 180 writes the playable media content to a media content file located in one or more server-side data storage 155 and server-side memory 175.

[0112] For example, server-side Multimodal Al Media Application 190 may produce audio and media content.

[0113] For example, server-side Multimodal Al System Application 195 may comprise of one or more of server-side text-to-speech Al application, text-to-audio Al application, audio-to-dialog Al application, location Al application, dialog supervisor Al application, privacy guardrail Al application, marketing Al application, server-based Al model, real-time Al application and an other application.

[0114] Server-side streaming application 185 segments media content for deployment via network 115 to client-side device 120. Server-side streaming application 185 supports multiple alternate data streams, two or more of which can have different bit rates from each other. For example, the multiple alternate data streams might have same bit rates from each other. Server-side streaming application 185 also allows for client-side device 120 to switch streams intelligently as network bandwidth changes. Server-side streaming application 185 also provides for media encryption and user authentication over encrypted connections.

[0115] Speaker system 135 may receive audio via one or more server-side computing device 170, network 115, an analogue connection, and a digital connection. For example, speaker systems 135 may be one or more actively and passively used with amplifiers.

[0116] Microphone system 137, may receive audio via speaker system 137. Microphone system 137 may be placed on the wall, on the speaker, or within thevenue environment. Microphone system 137 may used to calibrate speaker system 135 and verify the speaker conditions.

[0117] Input and output (I / O) device 198 may receive audio via one or more directly from server-side computing device 170, over the network 115, an analogue connection, and a digital connection. For example, input and output (I / O) device 198 may perform one or more digital to analogue and analogue to digital conversions. For example, input and output (I / O) device 198 may route audio to one or more of one or more amplifier, speaker and an other electronic device.

[0118] As explained below in greater detail, particularly in Figure 5, server-side Multimodal Al System Application 195 is configured to determine by client-side device, its own position (xs, ys, zs) via multilateration using three or more wireless communication sensors 165 or alternatively, may be accessed as a predetermined value corresponding to a known, fixed location.

[0119] FIGURE 2 is a schematic block diagram of server-side computing device 170 in an alternative configuration of a networked environment for real-time synchronization of media content via one or more of one or more electronic device and speaker system.

[0120] Server-side computing device 170 comprises one or more server-side data storage 155, server-side playback device 196 (not pictured), server-side memory 175, server-side processor 210, and server-side local interface 220. Server-side local interface 220 is operationally connected with one or more serverside data storage 155, server-side memory 180, and server-side processor 210. \Server-side memory comprises one or more server-side master application 180, server-side streaming application 185, server-side Multimodal Al Media Application 190, and server-side Multimodal Al System Application 195. For example, server-side processor 210 comprises server-side computer. For example, server-side local interface 220 comprises a bus. For example, serverside local interface 220 comprises a bus and further comprises one or more of an accompanying address / control bus or other bus structure.

[0121] Software components stored in one or more server-side memory 175 and server-side data storage 155 are executable by server-side processor 210. In this respect, the term executable means a program file that is in a form that can ultimately be run by server-side processor 210. For example, a compiled program is executable if it may be translated into machine code in a format that can be loaded into a random access portion of server-side memory 175 and run by serverside processor 210. For example, source code is executable if it may be expressed in a proper format, such as object code, that may be loaded into a random access portion of server-side memory 175 and run by server-side processor 210. For example, source code is executable if it may be interpreted by another executable program to generate instructions in a random access portion of server-side memory 175 and run by server-side processor 210. An executable program may be stored in one or more portions or components of server-side memory 175. For example, server-side memory 175 comprises one or more random access memory (RAM), read-only memory (ROM), hard disk drive, solid-state drive, USB flash drive, memory card, floppy disk, optical disc such as compact disc (CD) or digital versatile disc (DVD), magnetic tape, and other memory components.

[0122] One or more data and components stored in one or more server-side memory 175 and server-side data storage 155 are executable by server-side processor 210. For example, server-side processor 210 can execute one or more server-side master application 180, server-side streaming application 185, serverside Multimodal Al Media Application 190, and server-side Multimodal Al System Application 195.

[0123] For example, as an alternative to the setup in FIGURE 1 with serverside data storage 155 separate from server-side computing device 170, serverside data storage 155 may be located in server-side computing device 170. For example, server-side data storage 155 may be located in server-side memory 175.

[0124] Server-side processor 210 comprises one or more processors. Serverside memory 175 comprises one or more memories. For example, server-side memory 175 comprises at least one memory configured to operate in a parallel processing circuit. In such a case, server-side local interface 220 may serve as network 115. For example, server-side local interface 220 may facilitate communication between two processors. For example, server-side local interface 220 may facilitate communication between processor and a memory. For example, server-side local interface 220 may facilitate communication between two memories. Server-side local interface 220 may comprise additional systems designed to coordinate this communication. For example, server-side local interface 220 may comprise a system to perform load balancing. Server-side processor 210 may comprise an electrical processor. Alternatively, or additionally, server-side processor 210 may comprise a non-electrical processor.

[0125] Any logic or application described herein, including but not limited to server-side master application 180, server-side streaming application 185, serverside Multimodal Al Media Application 190, and server-side Multimodal Al System Application 195 that comprises software or code can be embodied in any non- transitory computer-readable medium for use by or in connection with an instruction execution system such as, for example, server-side processor 210 in a computer system or other system. In this sense, the logic may comprise, for example, statements including instructions and declarations that can be fetched from the computer-readable medium and can be executed by the instruction execution system. In the context of the present disclosure, a computer-readable medium can be any medium that can contain, store, or maintain the logic or application described herein for use by or in connection with the instruction execution system. For example, the computer-readable medium may comprise one or more RAM, ROM, hard disk drive, solid-state drive, USB flash drive, memory card, floppy disk, optical disc such as a CD or a DVD, magnetic tape, and other memory components. For example, RAM may comprise one or more SRAM, DRAM, MRAM, and other forms of RAM. For example, ROM may comprise one or more PROM, EPROM, EEPROM, and other forms of ROM.

[0126] FIGURE 3 illustrates a system wherein a user 305 wears client-side device 310 that synchronizes with server-side device 315, which is shown in this embodiment as a humanoid robot. Server-side device 315 may feature a digital screen 320 for displaying synchronized media content, and its head may comprisemechanized or biological components, such as moving eyes, eyelids, a nose, and lips, configured to synchronize with client-side device 310. While illustrated as a humanoid robot, in other embodiments server-side device 315 is not so limited. For example, server-side device 315 may be a vehicle, an autonomous vehicle, a wearable, or another mechanical or biological device. In another embodiment, server-side device 315 may be an intelligent environment that acts as a digital counterpart to a physical location, such as a hospital, museum, or stadium. Said intelligent environment may be configured to process information and communicate intelligently with client-side device 310. In any embodiment, clientside device 310 may synchronize with any combination of the digital, mechanical, or biological components of server-side device 315. Furthermore, the system accommodates multiple states of relative motion, wherein either client-side device 310 or server-side device 315 may be in motion while the other is static, or both may be in motion concurrently in a two-dimensional (2D) or three-dimensional (3D) plane.

[0127] FIGURE 4 is a flowchart of method 400 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 4 applies to a method using Multimodal Al, as viewed from client-side.

[0128] The order of the steps in the method 400 is not constrained to that shown in Figure 4 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0129] According to this method, client-side device performs tracking, synchronization, and spatialization of media content of client-side playback with server-side playback using a Multimodal Al System and wireless communication sensors.

[0130] In block 405, receiving, by client-side device controlled by a user, one or more user commands, said commands comprising text commands, verbal commands, gesture-based commands, or touch-based commands, and receiving application sensor data and location Al application data to be processed and played on client-side device in coordination with server-side playback of media content by server-side computing device. Block 405 transfers control to block 410.

[0131] Next, in block 410, obtaining, by client-side device, attributes of the user using an authentication Al application. Block 410 transfers control to block 415.

[0132] Next, in block 415, obtaining, by client-side device, stored personal profile information of the user. Block 415 transfers control to block 420.

[0133] Next, in block 420, processing, by client-side device, verbal commands using a voice-activated Al application and a voice-to-text Al application to generate a user text prompt. Block 420 transfers control to block 425.

[0134] Next, in block 425, sending, by client-side device, client-side unique identifier, the application sensor data, the location Al application data, and the user text prompt to server-side device. Block 425 transfers control to block 430.

[0135] Next, in block 430, receiving, by client-side device from server-side device or a cloud database, an audio content track, an ambient audio track, serverside device's location coordinates (xs, ys, zs), media-related data, and server-side device unique identifier. Block 430 transfers control to block 435.

[0136] Next, in block 435, creating, by client-side device, client-side packet comprising client-side unique identifier and client-side start host time (CSHT). Block 435 transfers control to block 440.

[0137] Next, in block 440, sending, by client-side device, said client-side packet to server-side device. Block 440 transfers control to block 445.

[0138] Next, in block 445, receiving and processing, by client-side device, server-side packet comprising server-side unique identifier, the CSHT, server-side start host time (SSHT), server-side end host time (SEHT), and server-side running media play time (SRMPT). Block 445 transfers control to block 450.

[0139] Next, in block 450, synchronizing in real-time, by client-side device, client-side playback with server-side playback using said CSHT, SSHT, SEHT, and SRMPT. Concurrently determining, by client-side device, its own position (xc, yc, zc) via multilateration using three or more wireless communication sensors. Block 450 transfers control to block 455.

[0140] Next, in block 455, calculating, by client-side device, an azimuth angle 0 using 6 = arctan2((ys - yc) / (xs - xc)). Calculating, by client-side device, an elevation angle cp using cp = arctan((zs - zc) I sqrt((xs - xc)2+ (ys - yc)2)).

[0141] Block 455 transfers control to block 460.

[0142] Next, in block 460, calculating, by client-side device, a first delay corresponding to the acoustic propagation time from server-side device location to client-side device location. Block 460 transfers control to block 465.

[0143] Next, in block 465, generating, by client-side device, using the audio content track and the ambient audio track, one or more mixed audio content. Applying one or more audio filters to the mixed audio content, wherein said application comprises: selecting, from a database, a pair of Head-Related Impulse Responses (HRIR), said pair comprising a left-ear HRIR and a right-ear HRIR corresponding to the calculated azimuth and elevation angles (0, <p). Block 465 transfers control to block 470.

[0144] Next, in block 470, generating a left channel of the binaural stereo output by convolving the mixed audio content with the left-ear HRIR. Generating a right channel of the binaural stereo output by convolving the mixed audio content with the right-ear HRIR. Block 470 transfers control to block 475.

[0145] Next, in block 475, applying the calculated first delay and one or more loudness compensation signals to the resulting binaural stereo output to render distance. Further processing the binaural stereo output using one or more of a Binaural Room Impulse Response (BRIR), a Headphone Transfer Function (HpTF), and a voice isolator, wherein said BRIR may be a pre-measured response or may be generated by rendering a more general spatial representation, such as a Spatial Room Impulse Response (SRIR), for the current user orientation. Block 475 transfers control to block 480.

[0146] Next, in block 480, calculating, by processor on client-side device, a second delay corresponding to the total signal processing latency introduced by the application of said spatialization filters. Block 480 transfers control to block 485.

[0147] Next, in block 485, applying, by client-side device, a final compensating delay to client-side playback, said final compensating delay accounting for the firstdelay and the second delay to ensure the processed audio is aligned with the synchronized playback timeline. Block 485 transfers control to block 490.

[0148] Next, in block 490, playing back, by client-side device, the fully synchronized and spatialized mixed audio and media content. Block 490 then terminates the process.

[0149] It is to be understood that while the embodiment described in blocks 405 through 490 specifies the selection and application of Head-Related Impulse Responses (HRIR) and Binaural Room Impulse Responses (BRIR), the scope of the invention is not limited to these specific filter formats. The spatialization filters and responses applied by client-side device may be derived, synthesized, or rendered from any suitable spatial audio representation. Such representations include, but are not limited to, higher-order Ambisonics, spherical harmonic coefficients, or Spatial Room Impulse Responses (SRIRs). The process of calculating orientation angles (0, <p) and using them to generate a specific binaural output from these broader spatial representations is contemplated as being within the scope of the present invention.

[0150] FIGURE 5 is a flowchart of method 500 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 5 applies to a method using Multimodal Al, as viewed from server-side.

[0151] The order of the steps in the method 500 is not constrained to that shown in Figure 5 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0152] According to this method, client-side device performs tracking, synchronization, and spatialization of media content of client-side playback with server-side playback using a Multimodal Al System and wireless communication sensors.

[0153] In block 505, receiving, by server-side computing device, client-side unique identifier, application sensor data, location Al application data, and a user text prompt from client-side device controlled by a user. Block 505 transfers control to block 510.

[0154] Next, in block 510, processing, by server-side device, the user text prompt, potentially in conjunction with dialog supervisor Al data, via a text-to- speech Al application to produce a primary audio content track. Block 510 transfers control to block 515.

[0155] Next, in block 515, generating, by server-side device, an ambient audio track using the received location Al application data to reflect the user's environment. Block 515 transfers control to block 520.

[0156] Next, in block 520, concurrently determining, by server-side device, its own real-time location coordinates (xs, ys, zs), for example, via multilateration using three or more wireless communication sensors with known positions. Block 520 transfers control to block 525.

[0157] Next, in block 525, transmitting, by server-side device to client-side device, the audio content track, the ambient audio track, server-side device's location coordinates (xs, ys, zs), media-related data and server-side device uniqueidentifier to be used for synchronization, and spatialization by the client. Block 525 transfers control to block 530.

[0158] Next, in block 530, waiting for and subsequently receiving, by serverside device from client-side device, client-side packet, said packet comprising client-side unique identifier and client-side start host time (CSHT). Block 530 transfers control to block 535.

[0159] Next, in block 535, immediately upon receipt of client-side packet, creating, by server-side device, server-side packet comprising server-side unique identifier, the received CSHT, server-side start host time (SSHT), server-side end host time (SEHT), and server-side running media play time (SRMPT). Block 535 transfers control to block 540.

[0160] Next, in block 540, transmitting, by server-side device, said server-side packet to client-side device, thereby providing the necessary timing data for clientside device to perform playback synchronization. Block 540 transfers control to block 545.

[0161] Next, in block 545, initiating, by server-side device, playback of the primary audio content track, said playback being internally synchronized with the timing information dispatched to client-side device to maintain a coherent shared experience. Block 545 then terminates the process.

[0162] FIGURE 6 is a flowchart of method 600 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 6 applies to a method using a single, natively Multimodal Al, as viewed from client-side.

[0163] The order of the steps in the method 600 is not constrained to that shown in Figure 6 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0164] According to this method, client-side device performs tracking, synchronization, and spatialization of media content of client-side playback with server-side playback using a single, natively Multimodal Al System and wireless communication sensors. Client-side device is configured to communicate with server-side device via real-time audio and data stream.

[0165] In block 605, receiving, by client-side device controlled by a user, one or more user commands, said commands comprising text commands, verbal commands, gesture-based commands, or touch-based commands, and receiving application sensor data and location Al application data. Block 605 transfers control to block 610.

[0166] Next, in block 610, obtaining, by client-side device, attributes of the user using an authentication Al application. Block 610 transfers control to block 615.

[0167] Next, in block 615, capturing, by client-side device, a verbal command from the user as a continuous audio stream. Block 615 transfers control to block 620.

[0168] Next, in block 620, transmitting, by client-side device, the continuous audio stream, client-side unique identifier, the application sensor data, and the location Al application data to server-side device stream using a single, natively Multimodal Al model to directly generate client-side response, wherein saidprocessing occurs without an intermediate text transcription step to reduce response latency. Block 620 transfers control to block 625.

[0169] Next, in block 625, receiving, by client-side device from server-side device, a response packet. Said packet comprises an audio content track, an ambient audio track, server-side device's location coordinates (xs, ys, zs), media-related data, and server-side device unique identifier. Block 625 transfers control to block 630.

[0170] Next, in block 630, creating, by client-side device, client-side packet comprising client-side unique identifier and client-side start host time (CSHT). Block 630 transfers control to block 635.

[0171] Next, in block 635, sending, by client-side device, said client-side packet to server-side device. Block 635 transfers control to block 640.

[0172] Next, in block 640, receiving and processing, by client-side device, server-side packet comprising server-side unique identifier, the CSHT, server-side start host time (SSHT), server-side end host time (SEHT), and server-side running media play time (SRMPT). Block 640 transfers control to block 645.

[0173] Next, in block 645, synchronizing in real-time, by client-side device, client-side playback with server-side playback using said CSHT, SSHT, SEHT, and SRMPT. Concurrently determining, by client-side device, its own position (xc, yc, zc) via multilateration using three or more wireless communication sensors. Block 645 transfers control to block 650.

[0174] Next, in block 650, calculating, by client-side device, an azimuth angle 6. Calculating, by client-side device, an elevation angle q>. Said calculation comprises: a) first, determining a baseline direction vector between server-side device's location coordinates (xs,ys,zs) and client-side device's own location coordinates (xc,yc,zc); b) second, applying a rotation transformation to the baseline direction vector based on real-time orientation data received from the user's head tracking sensor; c) finally, calculating the azimuth and elevation angles from the resulting rotated vector. Block 650 transfers control to block 655.

[0175] Next, in block 655, calculating, by client-side device, a first delay corresponding to the acoustic propagation time from server-side device location to client-side device location. Block 655 transfers control to block 660.

[0176] Next, in block 660, generating, by client-side device, using the audio content track and the ambient audio track, one or more mixed audio content. Applying one or more audio filters to the mixed audio content, wherein said application comprises: Selecting, from a database, a pair of Head-Related Impulse Responses (HRIR), said pair comprising a left-ear HRIR and a right-ear HRIR corresponding to the calculated azimuth and elevation angles (0, q>). Block 660 transfers control to block 665.

[0177] Next, in block 665, generating a left channel of the binaural stereo output by convolving the mixed audio content with the left-ear HRIR. Generating a right channel of the binaural stereo output by convolving the mixed audio content with the right-ear HRIR. Block 565 transfers control to block 670.

[0178] Next, in block 670, applying the calculated first delay and one or more loudness compensation signals to the resulting binaural stereo output to render distance. Further processing the binaural stereo output using one or more of aBinaural Room Impulse Response (BRIR), a Headphone Transfer Function (HpTF), speed-adaptive equalization filters, noise-cancellation, and a voice isolator, wherein said BRIR may be a pre-measured response or may be generated by rendering a more general spatial representation, such as a Spatial Room Impulse Response (SRIR), for the current user orientation. Block 670 transfers control to block 675.

[0179] Next, in block 675, calculating, by processor on client-side device, a second delay corresponding to the total signal processing latency introduced by the application of said spatialization filters. Block 675 transfers control to block 680.

[0180] Next, in block 680, applying, by client-side device, a final compensating delay to client-side playback, said final compensating delay accounting for the first delay and the second delay to ensure the processed audio is aligned with the synchronized playback timeline. Block 680 transfers control to block 685.

[0181] Next, in block 685, playing back, by client-side device, the fully synchronized and spatialized mixed audio and media content. Block 685 then terminates the process.

[0182] It is to be understood that while the embodiment described in blocks 605 through 685 specifies the selection and application of Head-Related Impulse Responses (HRIR) and Binaural Room Impulse Responses (BRIR), the scope of the invention is not limited to these specific filter formats. The spatialization filters and responses applied by client-side device may be derived, synthesized, or rendered from any suitable spatial audio representation. Such representations include, but are not limited to, higher-order Ambisonics, spherical harmonic coefficients, or Spatial Room Impulse Responses (SRIRs). The process of calculating orientation angles (0, <p) and using them to generate a specific binaural output from these broader spatial representations is contemplated as being within the scope of the present invention.

[0183] FIGURE 7 is a flowchart of method 700 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 7 applies to a method using a single, natively Multimodal Al, as viewed from server-side.

[0184] The order of the steps in the method 700 is not constrained to that shown in Figure 7 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0185] According to this method, client-side device performs tracking, synchronization, and spatialization of media content of client-side playback with server-side playback using a single, natively Multimodal Al System and wireless communication sensors. Server-side device is configured to communicate with client-side device via real-time audio and data stream.

[0186] In block 705, receiving, by server-side device, a continuous audio stream, client-side unique identifier, application sensor data, and location Al application data from client-side device in real-time. Block 705 transfers control to block 710.

[0187] Next, in block 710, processing, by server-side device, the received continuous audio stream using a single, natively multimodal Al model, wherein said model directly interprets the audio input to generate a contextual response withoutan intermediate text transcription step, thereby minimizing processing latency. Block 710 transfers control to block 715.

[0188] Next, in block 715, generating, by server-side device, one or more audio content tracks based on the response from the multimodal Al model, and concurrently generating an ambient audio track using the received location Al application data. Block 715 transfers control to block 720.

[0189] Next, in block 720, determining, by server-side device, its own real-time location coordinates (xs, ys, zs). Block 720 transfers control to block 725.

[0190] Next, in block 725, transmitting, by server-side device to client-side device, a response packet. Said packet comprises an audio content track and an ambient audio track, server-side device's location coordinates (xs, ys, zs), media-related data, and server-side device unique identifier. Block 725 transfers control to block 730.

[0001] Next, in block 730, waiting for and subsequently receiving, by serverside device from client-side device, client-side packet, said packet comprising client-side unique identifier and client-side start host time (CSHT). Block 730 transfers control to block 735.

[0002] Next, in block 735, immediately upon receipt of client-side packet, creating, by server-side device, server-side packet comprising server-side unique identifier, the received CSHT, server-side start host time (SSHT), server-side end host time (SEHT), and server-side running media play time (SRMPT). Block 735 transfers control to block 740.

[0003] Next, in block 740, transmitting, by server-side device, said server-side packet to client-side device, thereby providing the necessary timing data for clientside device to perform playback synchronization. Block 740 transfers control to block 745.

[0191] Next, in block 745, initiating, by server-side device, playback of the primary audio content track, said playback being internally synchronized with the timing information dispatched to client-side device to maintain a coherent shared experience. Block 745 then terminates the process.

[0192] FIGURE 8 is a flowchart of method 800 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 8 applies to a method using a Multimodal Al, as viewed from server-side.

[0193] The order of the steps in the method 800 is not constrained to that shown in Figure 8 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0194] According to this method, client-side device performs tracking, synchronization, and spatialization of media content of client-side playback with server-side playback using a Multimodal Al System and wireless communication sensors. Server-side device is configured to generate audio content track and ambient audio track, and one or more mixed audio content.

[0195] In block 805, processing, by server-side device, audio content track, ambient audio track, and data input; generating, by server-side device, using audio content track and ambient audio track, one or more mixed audio content;processing, by server-side device, mixed audio content and data input. Block 805 transfers control to block 810.

[0196] Next, in block 810, uploading, by server-side device, mixed audio content to server-side device or cloud database. Block 810 transfers control to block 815.

[0197] Next, in block 815, receiving, by client-side device, from server-side device or cloud database, mixed audio content; and processing, by client-side device, mixed audio content. Block 815 then terminates the process.

[0198] FIGURE 9 is a flowchart of method 900 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 9 applies to a method using a Multimodal Al, as viewed from client-side.

[0199] The order of the steps in the method 900 is not constrained to that shown in Figure 9 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0200] According to this method, client-side device performs tracking, synchronization, and spatialization of media content of client-side playback with server-side playback using a Multimodal Al System and wireless communication sensors. This method enables client-side device to synchronize its media playback with server-side device in both time and 3D space. It calculates network latency and the physical positions of both device, then uses this data to apply directional audio filters. The result is a near-perfectly timed, spatialized audio experience that sounds as if it is coming from the server's precise physical location after compensating for acoustic and processing delays.

[0201] In block 905, receiving, by client-side device, server-side message or packet from server-side master application, server-side message or packet comprising one or more of server-side unique identifier, client-side start host time (CSHT), server-side start host time (SSHT), server-side end host time (SEHT), and server-side running media play time (SRMPT). Block 905 transfers control to block 910.

[0202] Next, in block 910, reading, by client-side device, server-side message or packet into one or more client-side buffers. Block 910 transfers control to block 915.

[0203] Next, in block 915, recording, by client-side device, client-side end host time (CEHT) upon reception of the packet. Block 915 transfers control to block 920.

[0204] Next, in block 920, evaluating and verifying, by client-side device, the logical consistency of the timing data within server-side message or packet. Calculating a half-round-trip time (HRT), by client-side device, using the formula: HRT = ((CEHT - CSHT) - (SEHT - SSHT)) / 2; calculating a playback time offset (TPO), representing the clock offset between the client and server, by clientside device, using the formula: TPO = ((SSHT - CSHT) + (SEHT - CEHT)) I 2. Block 920 transfers control to block 925.

[0205] Next, in block 925, computing, by client-side device, a synchronized client-side running media play time (CRMPT) to align with the server's mediatimeline, using one or more of the SRMPT and the TPO. Block 925 transfers control to block 930.

[0206] Next, in block 930, concurrently sending, by client-side device to serverside device, a request to obtain server-side device location coordinates. Block 930 transfers control to block 935.

[0207] Next, in block 935, receiving, by client-side device, said server-side device location coordinates (xs, ys, zs); calculating, by client-side device, its own position (xc, yc, zc) via multilateration using signals received from three or more wireless communication sensors with known location coordinates. Block 935 transfers control to block 940.

[0208] Next, in block 940, calculating, by client-side device, an azimuth angle 0 using 9 = arctan2((ys - yc) I (xs - xc)). Calculating, by client-side device, an elevation angle <p using cp = arctan((zs - zc) / sqrt((xs - xc)2+ (ys - yc)2)). Block 940 transfers control to block 945.

[0209] Next, in block 945, calculating, by client-side device, a first delay corresponding to the acoustic propagation time from server-side device location to client-side device location. Block 945 transfers control to block 950.

[0210] Next, in block 950, applying, by the client-side, one or more audio filters to an audio content, wherein said application comprises: Selecting a Head-Related Transfer Function (HRTF) or Head-Related Impulse Response (HRIR) signal corresponding to the calculated azimuth and elevation angles. Applying the selected HRTF / HRIR signal to render directionality. Block 950 transfers control to block 955.

[0211] Next, in block 955, applying, by the client-side, loudness compensation signals to adjust for the relative distance between the device. Further processing the audio using one or more of a Binaural Room Impulse Response (BRIR), a Headphone Transfer Function (HpTF), speed-adaptive equalization filters, noisecancellation, and a voice isolator, wherein said BRIR may be a pre-measured response or may be generated by rendering a more general spatial representation, such as a Spatial Room Impulse Response (SRIR), for the current user orientation. Block 955 transfers control to block 960.

[0212] Next, in block 960, calculating, by processor on client-side device, a second delay corresponding to the signal processing latency introduced by the application of said audio filters. Block 960 transfers control to block 965.

[0213] Next, in block 965, applying, by client-side device, a final compensating delay to client-side playback, said final compensating delay accounting for the first delay and the second delay to ensure the processed audio is aligned with the synchronized timeline. Block 965 transfers control to block 970.

[0214] Next, in block 970, synchronizing, by client-side device, using the CRMPT, client-side playback of synchronized and spatialized mixed audio and media content to server-side playback of audio and media content. Block 970 then terminates the process.

[0215] It is to be understood that while the embodiment described in blocks 905 through 970 specifies the selection and application of Head-Related Impulse Responses (HRIR) and Binaural Room Impulse Responses (BRIR), the scope of the invention is not limited to these specific filter formats. The spatialization filtersand responses applied by client-side device may be derived, synthesized, or rendered from any suitable spatial audio representation. Such representations include, but are not limited to, higher-order Ambisonics, spherical harmonic coefficients, or Spatial Room Impulse Responses (SRIRs). The process of calculating orientation angles (0, >) and using them to generate a specific binaural output from these broader spatial representations is contemplated as being within the scope of the present invention.

[0216] FIGURE 10 is a flowchart of method 1000 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 10 applies to a method using a Multimodal Al, as viewed from server-side.

[0217] The order of the steps in the method 1000 is not constrained to that shown in Figure 10 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0218] According to this method, client-side device performs tracking, synchronization, and spatialization of media content of client-side playback with server-side playback using a Multimodal Al System and wireless communication sensors. This method enables server-side device to synchronize its media playback with client-side device in both time and 3D space. It calculates network latency and the physical positions of both device, then uses this data to apply directional audio filters. The result is a near-perfectly timed, spatialized audio experience that sounds as if it is coming from the server's precise physical location after compensating for acoustic and processing delays.

[0219] In block 1005, receiving, by server-side computing device, an initial client-side packet from client-side master application, the packet comprising one or more of client-side unique identifier and client-side start host time (CSHT). Block 1005 transfers control to block 1010.

[0220] Next, in block 1010, recording, by server-side device, server-side start host time (SSHT) immediately upon reception of client-side packet. Block 1010 transfers control to block 1015.

[0221] Next, in block 1015, generating or retrieving, by server-side device, a non-spatialized audio content track and any related media content based on an application state or a user request. Block 1015 transfers control to block 1020.

[0222] Next, in block 1020, recording, by server-side device, server-side end host time (SEHT) and determining an authoritative server-side running media play time (SRMPT) corresponding to the state of the media content to be played. Block 1020 transfers control to block 1025.

[0223] Next, in block 1025, creating, by server-side device, server-side response packet, said packet comprising server-side unique identifier, the received CSHT, the recorded SSHT, the recorded SEHT, the SRMPT, and the non-spatialized audio content track. Block 1025 transfers control to block 1030.

[0224] Next, in block 1030, transmitting, by server-side device, said serverside response packet to client-side device, thereby providing client-side device with the necessary data to calculate synchronization and to perform spatialization. Block 1030 transfers control to block 1035.

[0225] Next, in block 1035, concurrently determining, by server-side device, its own real-time location coordinates (xs, ys, zs) via multilateration or other means. Block 1035 transfers control to block 1040.

[0226] Next, in block 1040, waiting for and subsequently receiving, by serverside device, a secondary request from client-side device, said request being for server-side device's location coordinates. Block 1040 transfers control to block 1045.

[0227] Next, in block 1045, transmitting, by server-side device in response to said secondary request, its determined location coordinates (xs, ys, zs) to clientside device. Block 1045 transfers control to block 1050.

[0228] Next, in block 1050, initiating playback, by server-side device, of its own media content through one or more speaker systems, said playback being internally aligned with the SRMPT to ensure it is synchronized with the corresponding playback occurring on client-side device. Block 1050 then terminates the process.

[0229] FIGURE 11 is a flowchart of method 1100 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 11 applies to a method using a Multimodal Al, as viewed from client-side.

[0230] The order of the steps in the method 1100 is not constrained to that shown in Figure 11 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0231] According to this method, client-side device performs spatializing media content in a two-dimensional (2D) plane.

[0232] In block 1105, receiving, by client-side device, signals from two wireless communication sensors with known location coordinates in a 2D plane. Block 1105 transfers control to block 1110.

[0233] Next, in block 1110, calculating, by client-side device, the respective distances to each of the two wireless communication sensors. Calculating, by client-side device, its 2D position (xc, yc) by determining the intersection point of two circles, wherein the center of each circle is the known location of one of the wireless communication sensors and the radius of each circle is the calculated distance to said sensor. Block 1110 transfers control to block 1115.

[0234] Next, in block 1115, receiving, by client-side device, the current 2D location coordinates (xs, ys) of server-side device or other tracked object. Block 1115 transfers control to block 1120.

[0235] Next, in block 1120, calculating, by client-side device, an azimuth angle, 9, between client-side device and server-side device in the horizontal plane, using the formula 0 = arctan2((ys - yc) / (xs - xc)). Block 1120 transfers control to block 1125.

[0236] Next, in block 1125, calculating, by client-side device, one or more acoustic propagation delays based on the 2D distance between the calculated client position (xc, yc) and the received server position (xs, ys). Block 1125 transfers control to block 1130.

[0237] Next, in block 1130, generating, by client-side device, a filter tracking response based on the calculated azimuth angle and acoustic propagation delay.Applying the filter tracking response to one or more audio content, thereby rendering the audio to be perceived as originating from the direction and distance of server-side device within the 2D plane. Block 1130 then terminates the process.

[0238] FIGURE 12 is a flowchart of method 1200 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 12 applies to a method using a Multimodal Al, as viewed from client-side.

[0239] The order of the steps in the method 1200 is not constrained to that shown in Figure 12 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0240] According to further embodiments of the invention, a method for dynamically updating spatialization of media content on client-side device.

[0241] In block 1205, storing, by client-side device, a set of previous serverside device location coordinates (xs, ys, zs) and a set of previous orientation angles (0, cp) in a memory buffer. Block 1205 transfers control to block 1210.

[0242] Next, in block 1210, receiving, at a pre-determined interval, a current set of server-side device location coordinates (xs, ys, zs). Block 1210 transfers control to block 1215.

[0243] Next, in block 1215, calculating, by client-side device, a current set of orientation angles (0, cp) based on the current relative positions of client-side device and server-side device. Block 1215 transfers control to block 1220.

[0244] Next, in block 1220, calculating a displacement magnitude between the previous and current server-side device location coordinates; comparing, by clientside device, the calculated displacement magnitude to a pre-determined threshold. If the displacement magnitude exceeds the pre-determined threshold, recalculating, by client-side device, one or more first delays corresponding to the new acoustic propagation time and generating a new filter tracking response based on the current server-side device coordinates (xs, ys, zs) and current orientation angles (0, cp). Block 1220 transfers control to block 1225.

[0245] Next, in block 1225, applying, by client-side device, a smooth transition algorithm, such as a cross-fade, between the previous filter tracking response and the newly calculated filter tracking response to prevent abrupt audible changes in the spatialized audio output. Block 1225 transfers control to block 1230.

[0246] Next, in block 1230, replacing, by client-side device, the previous server-side device location coordinates and previous orientation angles in the memory buffer with the current coordinates and current angles. Block 1230 transfers control to block 1235.

[0247] Next, in block 1235, continuing playback, by client-side device, of the media content using the newly updated and smoothly transitioned spatialization parameters. Block 1235 then terminates the process.

[0248] FIGURE 13 is a flowchart of method 1300 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 13 applies to a method using a Multimodal Al, as viewed from client-side.

[0249] The order of the steps in the method 1300 is not constrained to that shown in Figure 13 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0250] In yet further embodiments of the invention, a method for dynamically updating spatialization of media content on client-side device to track a moving server-side device includes.

[0251] In block 1305, storing, by client-side device, a set of previous serverside device location coordinates (x_s_prev, y_s_prev, z_s_prev) and a set of previous spatialization filter parameters in a memory buffer. Block 1305 transfers control to block 1310.

[0252] Next, in block 1310, receiving, by client-side device, a current set of server-side device location coordinates (x_s_curr, y_s_curr, z_s_curr). Block 1310 transfers control to block 1315.

[0253] Next, in block 1315, calculating a displacement magnitude, D, between the previous and current server-side device location coordinates, wherein D = sqrt((x_s_curr - x_s_prev)2+ (y_s_curr - y_s_prev)2+ (z_s_curr - z_s_prev)2). Comparing, by client-side device, the calculated displacement magnitude D to a pre-determined displacement threshold, wherein said threshold represents a minimum perceptually significant change in position. If the displacement magnitude D exceeds the pre-determined displacement threshold. Block 1315 transfers control to block 1320.

[0254] Next, in block 1320, re-calculating, by client-side device, a current set of orientation angles (6, cp) based on the current relative positions of client-side device and server-side device. Block 1320 transfers control to block 1325.

[0255] Next, in block 1325, re-calculating one or more acoustic propagation delays and generating a new set of spatialization filter parameters based on the current server-side device coordinates and current orientation angles. Block 1325 transfers control to block 1330.

[0256] Next, in block 1330, applying, by client-side device, a smooth transition algorithm between the previous spatialization filter parameters and the new spatialization filter parameters. For example, said smooth transition algorithm may comprise a cross-fade implemented via a linear interpolation over a defined time interval to prevent abrupt, audible artifacts in the spatialized audio output. Block 1330 transfers control to block 1335.

[0257] Next, in block 1335, replacing, by client-side device, the previous server-side device location coordinates and previous spatialization filter parameters in the memory buffer with the current coordinates and current spatialization filter parameters. Block 1335 transfers control to block 1340.

[0258] Next, in block 1340, continuing playback, by client-side device, of the media content using the newly updated and smoothly transitioned spatialization filter parameters. Block 1340 then terminates the process.

[0259] FIGURE 14 is a flowchart of method 1400 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 14 applies to a method using a Multimodal Al, as viewed from client-side.

[0260] The order of the steps in the method 1400 is not constrained to that shown in Figure 14 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0261] Yet, in other embodiments of the invention, a method for predictive spatialization of media content on client-side device.

[0262] In block 1405, receiving, by client-side device, a current set of serverside device location coordinates (xs, ys, zs) and server-side velocity vector (vxs, vys, vzs), said velocity vector being derived from one or more of server-side device accelerometer and changes in wireless communication sensor data. Block 1405 transfers control to block 1410.

[0263] Next, in block 1410, concurrently receiving, by client-side device, a current set of client-side device location coordinates (xc, yc, zc) and client-side velocity vector (vxc, vyc, vzc), said velocity vector being derived from one or more of client-side device accelerometer and changes in wireless communication sensor data. Block 1410 transfers control to block 1415.

[0264] Next, in block 1415, defining a future time interval, t; calculating, by client-side device, a future client-side device location coordinate (xc_future, yc_future, zc_future), using the formulas: xc_future = xc + vxc * t; yc_future = yc + vyc * t; zc_future = zc + vzc * t. Block 1415 transfers control to block 1420.

[0265] Next, in block 1420, calculating, by client-side device, a future serverside device location coordinate (xs_future, ys_future, zs_future), using the formulas: xs_future = xs + vxs * t; ys_future = ys + vys * t; zs_future = zs + vzs * t. Block 1420 transfers control to block 1425.

[0266] Next, in block 1425, calculating, by client-side device, a future azimuth angle, 0_future, using the predicted future coordinates: 9_future = arctan2((ys_future - yc_future) I (xs_future - xc_future)). Calculating, by client-side device, a future elevation angle, <p_future, using the predicted future coordinates: cp_future = arctan((zs_future - zc_future) / sqrt((xs_future - xc_future)2+ (ys_future - yc_future)2)). Block 1425 transfers control to block 1430.

[0267] Next, in block 1430, generating a set of predictive spatialization filter parameters based on said future azimuth angle, future elevation angle, and a predicted future acoustic propagation delay. Block 1430 transfers control to block 1435.

[0268] Next, in block 1435, storing, by client-side device, the predictive spatialization filter parameters. Block 1435 transfers control to block 1440.

[0269] Next, in block 1440, applying, by client-side device, the stored predictive spatialization filter parameters to the audio content at the future time t, thereby compensating for system latency and aligning the audio spatialization with the predicted real-world positions of the device at the moment of playback. Block 1440 then terminates the process.

[0270] FIGURE 15 is a flowchart of method 1500 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 15 applies to a method using a Multimodal Al, as viewed from client-side^

[0271] The order of the steps in the method 1500 is not constrained to that shown in Figure 15 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0272] Still, in other embodiments of the invention, a method for predictive spatialization of multi-object track media content on client-side device using a Multimodal Al System and wireless communication sensors includes.

[0273] In block 1505, receiving, by client-side device, a current set of serverside device location coordinates (xs, ys, zs) from one or more of server-side master application and one or more wireless communication sensors. Block 1505 transfers control to block 1510.

[0274] Next, in block 1510, receiving, by client-side device, server-side velocity vector (vxs, vys, vzs), said velocity vector being derived from one or more of server-side device accelerometer and changes in wireless communication sensor data. Block 1510 transfers control to block 1515.

[0275] Next, in block 1515, concurrently receiving, by client-side device, a current set of client-side device location coordinates (xc, yc, zc) from one or more wireless communication sensors. Receiving, by client-side device, client-side velocity vector (vxc, vyc, vzc), said velocity vector being derived from one or more of client-side device accelerometer and changes in wireless communication sensor data. Block 1515 transfers control to block 1520.

[0276] Next, in block 1520, defining a future time interval, t, said time interval corresponding to a predictive look-ahead duration. Storing, by client-side device, the current client-side device location coordinates (xc, yc, zc) and current serverside device location coordinates (xs, ys, zs). Block 1520 transfers control to block 1525.

[0277] Next, in block 1525, calculating, by client-side device, a future clientside device location coordinate (xc_future, yc_future, zc_future), using the formulas: xc_future = xc + vxc * t; yc_future = yc + vyc * t; zc_future - zc + vzc * t. Block 1525 transfers control to block 1530.

[0278] Next, in block 1530, calculating, by client-side device, a future serverside device location coordinate (xs_future, ys_future, zs_future), using the formulas: xs_future = xs + vxs * t; ys_future = ys + vys * t; zs_future = zs + vzs * t; Calculating, by client-side device, a future azimuth angle, 0_future, using the predicted future coordinates: 0_future = arctan2((ys_future - yc_future) / (xs_future - xc_future)). Calculating, by client-side device, a future elevation angle, cp_future, using the predicted future coordinates: cp_future = arctan((zs_future - zc_future) / sqrt((xs_future - xc_future)2+ (ys_future - yc_future)2)). Block 1530 transfers control to block 1535.

[0279] Next, in block 1535, storing, by client-side device, the calculated future azimuth and elevation angles. Block 1535 transfers control to block 1540.

[0280] Next, in block 1540, generating a set of predictive spatialization filter parameters based on said future azimuth angle, future elevation angle, and a predicted future acoustic propagation delay derived from the future coordinates. Block 1540 transfers control to block 1545.

[0281] Next, in block 1545, applying, by client-side device, the stored predictive spatialization filter parameters to the audio content, therebycompensating for system latency and aligning the audio spatialization with the predicted real-world positions of the device at the moment of playback. Block 1545 then terminates the process.

[0282] It is to be understood that while the method described in blocks

[1505] through

[1545] details the predictive spatialization process for a single server-side object, the present invention is not so limited. In embodiments involving multiobject track media content, client-side device is configured to receive a set of location coordinates and velocity vectors for a plurality of objects. The computational steps for predicting future positions, calculating future angles, and generating predictive spatialization filters are then performed for each object in the set, either serially or in parallel. The final audio output is a composite mix of the individually spatialized audio tracks corresponding to each object, thereby creating a cohesive and dynamic audio environment.

[0283] FIGURE 16 is a flowchart of method 1600 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 16 applies to a method using a Multimodal Al, as viewed from client-side.

[0284] The order of the steps in the method 1600 is not constrained to that shown in Figure 16 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0285] In additional embodiments of the invention, a method for improving location accuracy via fusion for real-time spatialization.

[0286] In block 1605, enabling, by client-side master application, a high- accuracy positioning mode on client-side device and concurrently sending a request to server-side device to activate a corresponding high-accuracy positioning mode. Block 1605 transfers control to block 1610.

[0287] Next, in block 1610, receiving, by server-side device, said request from client-side device and, in response, enabling its own high-accuracy positioning mode. Block 1610 transfers control to block 1615.

[0288] Next, in block 1615, concurrently performing, by both client-side device and server-side device, a location fusion process, wherein said process comprises: receiving, by each device, location-related data from a plurality of its respective available wireless communication protocols, said protocols comprising one or more of a Global Navigation Satellite System (GNSS), a satellite internet constellation system, Wi-Fi network data, Bluetooth beacon signals, Ultra- Wideband (UWB) signals, and cellular network data; and applying, by a location fusion Al application on each device, a sensor fusion algorithm to the aggregated data, wherein said algorithm weighs the data from each protocol based on factors such as signal strength and known accuracy to compute a single, high-fidelity location coordinate for each respective device. Block 1615 transfers control to block 1620.

[0289] Next, in block 1620, transmitting, by server-side device, its computed high-fidelity location coordinate (xs,ys,zs) to client-side device. Block 1620 transfers control to block 1625.

[0290] Next, in block 1625, receiving, by client-side device, server-side device's high-fidelity location coordinate. Block 1625 transfers control to block 1630.

[0291] Next, in block 1630, utilizing, by client-side device, both its own internally computed high-fidelity location coordinate (xc,yc,zc) and the received high-fidelity server-side location coordinate (xs,ys,zs) for subsequent spatialization calculations. For example, said high-fidelity coordinates are used to calculate one or more of an azimuth angle, an elevation angle, and one or more acoustic propagation delays, thereby improving the accuracy and stability of the resulting spatialized audio content. Block 1630 then terminates the process.

[0292] FIGURE 17 is a flowchart of method 1700 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 17 applies to a method using a Multimodal Al, as viewed from client-side.

[0293] The order of the steps in the method 1700 is not constrained to that shown in Figure 17 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0294] In yet additional embodiments of the invention, a method for adaptively selecting a positioning protocol for real-time spatialization.

[0295] In block 1705, enabling, by client-side master application, an adaptive positioning mode on one or more of client-side device and server-side device. Block 1705 transfers control to block 1710.

[0296] Next, block 1710, continuously monitoring, by the device, a set of conditions, said conditions comprising one or more of the signal strength of available positioning protocols, the power consumption requirements of the device, and the location accuracy requirements of a current application state. Said available positioning protocols may include a Global Navigation Satellite System (GNSS) such as GPS, a satellite internet constellation system, Wi-Fi, Bluetooth, Ultra-Wideband (UWB), cellular network data, and another protocol. In response to a change in the monitored conditions, selecting, by a location management Al application, an optimal positioning protocol from the available positioning protocols. For example, selecting UWB when high-precision, low-latency tracking is required and the signal is available, or selecting Wi-Fi or cellular data when the device is indoors or power conservation is prioritized over high precision. Block 1710 transfers control to block 1715.

[0297] Next, block 1715, switching, by the device, to the selected optimal positioning protocol as the primary source for location data. Block 1715 transfers control to block 1720.

[0298] Next, block 1720, receiving, by the device, location coordinates from the currently selected optimal positioning protocol. Block 1720 transfers control to block 1725.

[0299] Next, block 1725, utilizing said location coordinates to perform subsequent spatialization calculations, including the calculation of one or more of an azimuth angle, an elevation angle, and one or more acoustic propagation delays, thereby ensuring the spatialized audio content is rendered using the mostappropriate data source for the current networked environment. Block 1725 then terminates the process.

[0300] FIGURE 18 is a flowchart of method 1800 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 18 applies to a method using a Multimodal Al, as viewed from client-side.

[0301] The order of the steps in the method 1800 is not constrained to that shown in Figure 18 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0302] In alternative embodiments of the invention, a method for improving the accuracy of acoustic measurements for real-time spatialization.

[0303] In block 1805, continuously monitoring, by client-side device, an ambient air temperature and a relative humidity using one or more integrated environmental sensors. Block 1805 transfers control to block 1810.

[0304] Next, in block 1810, calculating, by client-side device, an adjusted speed of sound, c, in real-time based on the monitored ambient air temperature and relative humidity. For example, using a comprehensive formula that accounts for the effect of water vapor on the molar mass and specific heat ratio of air. Utilizing, by client-side device, the adjusted speed of sound c in subsequent acoustic distance calculations. For example, using the adjusted speed of sound c when calculating a distance d from an acoustic time-of-flight measurement t, wherein d=c-t, thereby reducing errors caused by variations in both temperature and humidity. Block 1810 then terminates the process.

[0305] FIGURE 19 is a flowchart of method 1900 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 19 applies to a method using a Multimodal Al, as viewed from client-side.

[0306] The order of the steps in the method 1900 is not constrained to that shown in Figure 19 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0307] In more embodiments of the invention, a method for applying a spatialization filter to a plurality of audio system formats includes.

[0308] Next, in block 1905, calculating, by client-side device, a filter tracking response, wherein said response comprises a set of parameters representing the real-time spatial relationship between client-side device and one or more sound sources, said parameters including one or more of an azimuth angle, an elevation angle, and a propagation delay. Block 1905 transfers control to block 1910.

[0309] Next, in block 1910, applying, by client-side device, the filter tracking response to a selected audio content format, wherein the method of application is adapted based on the architecture of the selected format. Block 1910 transfers control to block 1915.

[0310] Next, in block 1915, for an object-based audio system: applying the filter tracking response to one or more individual audio objects, thereby rendering each object at the specific spatial position defined by the filter parameters. Block 1915 transfers control to block 1920.

[0311] Next, in block 1920, for embodiments featuring a channel-based audio system, such as a stereo, 5.1, or 7.1 system, the process comprises utilizing the filter tracking response to generate a binaural Tenderer or a virtualizer. Said virtualizer is configured to map the fixed channels of the system to a three- dimensional soundfield to simulate an immersive audio experience. Block 1920 transfers control to block 1925.

[0312] Next, in block 1925, for embodiments featuring a scene-based audio system, such as an Ambisonics system, the process comprises utilizing the filter tracking response to rotate the entire captured soundfield in real-time, thereby aligning the orientation of the audio with the user's head orientation. Block 1925 transfers control to block 1930.

[0313] Next, in block 1930, for a parametric binaural signal: utilizing the filter tracking response to modify the spatial parameters of the parametric binaural signal, thereby adjusting the perceived location of the audio content. Block 1930 transfers control to block 1935.

[0314] Next, in block 1935, for a pre-rendered binaural audio track: applying the filter tracking response to rotate the entire pre-rendered binaural soundfield, thereby aligning the orientation of the audio with the user's head orientation in realtime. Block 1935 then terminates the process.

[0315] FIGURE 20 is a flowchart of method 2000 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 20 applies to a method using a Multimodal Al, as viewed from client-side.

[0316] The order of the steps in the method 2000 is not constrained to that shown in Figure 20 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0317] According to a streaming embodiment of the invention, a method for processing streamed media content on client-side device.

[0318] In block 2005, establishing a streaming session between client-side device and server-side streaming application, wherein real-time transport protocol is selected based on the low-latency requirements of the interactive Al response loop. Block 2005 transfers control to block 2010.

[0319] Next, in block 2010, continuously receiving, by client-side device, a stream of data packets, wherein said packets comprise one or more audio content tracks and spatialization metadata, such as the real-time position of the audio source (xs,ys,zs) generated by server-side device. Block 2010 transfers control to block 2015.

[0320] Next, in block 2015, placing, by client-side device, the received data packets into an adaptive jitter buffer. Said buffer is configured to dynamically adjust its size based on real-time network conditions, such as latency and packet arrival variance, and the latency tolerance of the current application state; for example, by minimizing the buffer size during rapid interactive conversation to prioritize responsiveness, while allowing the buffer to grow during passive content playback to maximize stability. Block 2015 transfers control to block 2020.

[0321] Next, in block 2020, periodically creating and sending, by client-side device, client-side packet containing client-side unique identifier and client-sidestart host time (CSHT) to server-side device to initiate a synchronization event. Block 2020 transfers control to block 2025.

[0322] Next, in block 2025, receiving and processing, by client-side device, server-side packet comprising server-side unique identifier, the original CSHT, server-side start host time (SSHT), server-side end host time (SEHT), and serverside running media play time (SRMPT). Synchronizing, in real-time, client-side playback timeline with server-side playback timeline using said received timestamps. Block 2025 transfers control to block 2030.

[0323] Next, in block 2030, after extracting data packets from the buffer according to the synchronized timeline, detecting any lost packets and applying a spatially-aware packet loss concealment (PLC) algorithm. Said algorithm intelligently prioritizes the reconstruction of audio data most critical to human perception of spatial location. For example, the algorithm may prioritize the interpolation of high-frequency transients essential for calculating interaural time differences (ITD) and levels (I LD), thereby minimizing perceived audio source drift or instability that would otherwise result from data loss. Block 2030 transfers control to block 2035.

[0324] Next, in block 2035, decoding, by client-side device, the synchronized and repaired data to generate one or more clean audio channels formatted for direct input into a spatialization engine. Block 2035 transfers control to block 2038.

[0325] Next, in block 2038, calculating, by client-side device, a final set of spatialization angles (0,4>) to be used for rendering, wherein said calculation comprises: a) determining a baseline direction vector between server-side device's location coordinates (xs,ys,zs) and client-side device's own determined location coordinates (xc,yc,zc); b) applying a rotation transformation to the baseline direction vector based on real-time orientation data received from a user's head tracking sensor; and c) calculating the azimuth and elevation angles from the resulting rotated vector. Block 2038 transfers control to block 2040.

[0326] Next, in block 2040, applying, by client-side device, one or more audio filters to the decoded audio based on the calculated spatialization angles (6,<|)) to generate a spatialized audio output, wherein the rendering is performed by convolving the audio in real-time with binaural data, such as a Head-Related Transfer Function (HRTF) or a Binaural Room Impulse Response (BRIR), and wherein said binaural data is continuously updated using the client's high- frequency position sensor data to keep the sound image spatially stable, and wherein said BRIR is selectable from a database of pre-measured responses or is generated by rendering a more general spatial representation, such as a Spatial Room Impulse Response (SRIR). Block 2040 transfers control to block 2045.

[0327] Next, in block 2045, calculating, by client-side device, a first delay corresponding to the acoustic propagation time from server-side device's location to client-side device's location. Block 2045 transfers control to block 2050.

[0328] Next, in block 2050, calculating, by client-side device, a second delay corresponding to the total signal processing latency introduced by the application of the spatialization filters in the preceding step. Block 2050 transfers control to block 2055.

[0329] Next, in block 2055, applying, by client-side device, a final compensating delay to client-side playback timeline, said final delay accounting for the calculated first delay (acoustic propagation) and the second delay (processing latency). Block 2055 transfers control to block 2060.

[0330] Next, in block 2060, playing back, by client-side device, the fully processed mixed audio content, which is now aligned with server-side playback timeline as a result of the multi-stage synchronization and delay compensation, and is rendered as a fully spatialized audio experience. Block 2060 then terminates the process.

[0331] FIGURE 21 is a flowchart of method 2100 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 21 applies to a method viewed from server-side.

[0332] The order of the steps in the method 2100 is not constrained to that shown in Figure 21 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0333] According to a streaming embodiment of the invention, a method for processing streamed media content on server-side device.

[0001] In block 2105, accepting, by server-side streaming application, a request from client-side device to establish a streaming session, wherein real-time transport protocol is selected based on the low-latency requirements of an interactive Al response loop. Block 2105 transfers control to block 2110.

[0002] Next, in block 2110, preparing, by server-side application, one or more content streams for transmission, wherein said content streams comprise one or more of an audio content track, an ambient audio track, mixed audio content, media content, and content generated by a Multimodal Al Media Application. Block 2110 transfers control to block 2115.

[0003] Next, in block 2115, continuously packaging, by server-side device, the generated audio content tracks into a series of data packets. Encapsulating, within said data packets, spatialization metadata, said metadata comprising the real-time position of the audio source (xs,ys,zs) corresponding to server-side device's location or a virtual source location. Block 2115 transfers control to block 2120.

[0004] Next, in block 2120, transmitting, by server-side streaming application to client-side device, the continuous stream of spatially-enriched data packets. Block 2120 transfers control to block 2125.

[0005] Next, in block 2125, concurrently with the streaming process of Block 2120, monitoring for and receiving, from client-side device, client-side packet sent to initiate a synchronization event, said packet containing client-side unique identifier and client-side start host time (CSHT). Block 2125 transfers control to block 2130.

[0006] Next, in block 2130, upon receiving a client-side packet, server-side device records the server-side start host time (SSHT) corresponding to the time of receipt; extracts the original CSHT from client-side packet; queries a master server clock to obtain server-side running media play time (SRMPT); constructs serverside packet comprising server-side unique identifier, the extracted CSHT, the recorded SSHT, the SRMPT, and server-side end host time (SEHT) recorded uponcompletion of the packet's construction; and transmits said server-side packet to client-side device. Block 2130 then transfers control to block 2135.

[0007] Next, in block 2135, continuing the process of streaming data packets (Block 2120) and monitoring for subsequent synchronization events (Block 2125) until a termination signal is received from client-side device. Block 2135 transfers control to block 2140.

[0008] Next, in block 2140, upon receiving a disconnect signal from client-side device or detecting a session timeout, terminating the transmission of data packets and closing the streaming session. Block 2140 then terminates the process.

[0334] FIGURE 22 is a flowchart of method 2200 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 22 applies to a method viewed from client-side.

[0335] The order of the steps in the method 2200 is not constrained to that shown in Figure 22 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0336] According to an embodiment of the invention, a method for synchronizing with a server and rendering, position-aware audio on client-side device.

[0337] In block 2205, determining, by client-side device, its own real-time spatial coordinates (xc, yc, zc) and orientation using one or more integrated position sensors, such as head motion tracking sensors. Block 2205 transfers control to block 2210.

[0338] Next, in block 2210, creating, by client-side device, a data packet containing its unique identifier along with its recently captured position and orientation metadata, said packet timestamped with client-side Start Host Time (CSHT), and transmitting said packet over a network to server-side computing device. Block 2210 transfers control to block 2215.

[0339] Next, in block 2215, receiving, by client-side device, a synchronization packet in response from the server, said packet comprising the original CSHT it sent, server-side Start Host Time (SSHT), server-side End Host Time (SEHT), and server-side Running Media Play Time (SRMPT). Block 2215 transfers control to block 2220.

[0340] Next, in block 2220, processing, by client-side device, the synchronization packet to calculate round-trip network latency and its clock's offset relative to the server's clock, and adjusting its internal playback timeline to be synchronized with the server's SRMPT. Block 2220 transfers control to block 2225.

[0341] Next, in block 2225, receiving, by client-side device, a dedicated audio content stream from the server, said stream comprising customized content, such as dialogue in a specific language. Block 2225 transfers control to block 2230.

[0342] Next, in block 2230, rendering, by client-side device, its received audio stream for playback. The rendering is performed by convolving the audio in realtime with binaural data, such as a Head-Related Transfer Function (HRTF) or a Binaural Room Impulse Response (BRIR), which is continuously updated using the client's high-frequency position sensor data to keep the sound image spatially stable. Furthermore, said BRIR may be selected from a database of pre-measuredresponses or may be generated by rendering a more general spatial representation, such as a Spatial Room Impulse Response (SRIR), according to the current user orientation. The process then repeats from block 2205 to provide continuous updates.

[0343] FIGURE 23 is a flowchart of method 2300 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 23 applies to a method viewed from server-side.

[0344] The order of the steps in the method 2300 is not constrained to that shown in Figure 23 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0345] According to an embodiment of the invention, a method for dynamically reconfiguring a multi-speaker audio system based on real-time position data from client device.

[0346] In block 2305, receiving, by server-side device, a data packet over a network from client-side device, said packet containing the client's unique identifier and its real-time position and orientation metadata. Block 2305 transfers control to block 2310.

[0347] Next, in block 2310, performing, by server-side device, an immediate synchronization handshake upon receipt of client-side packet by extracting clientside Start Host Time (CSHT) and transmitting a response packet containing the echoed CSHT, a newly captured Server-Side Start Host Time (SSHT), and the current Server-Side Running Media Play Time (SRMPT). Block 2310 transfers control to block 2315.

[0348] Next, in block 2315, processing, by server-side device, the client's location data by comparing the user's current position to a predefined ideal listening position as described in one or more stored audio layout files. Block 2315 transfers control to block 2320.

[0349] Next, in block 2320, calculating, by server-side device, a new audio signal configuration to shift the perceived sound image of the main speaker system in response to a significant deviation between the user's current and ideal positions, said calculation leveraging the precedence effect (Haas Effect) to determine near-precise micro-delays for each speaker signal. Block 2320 transfers control to block 2325.

[0350] Next, in block 2325, transmitting, by server-side device, the reconfigured audio signals with calculated delays to the corresponding channels of the physical multi-speaker system, and concurrently transmitting a separate, customized audio stream to client-side device for its own local rendering. Block 2325 transfers control to block 2330.

[0351] Next, in block 2330, initiating or continuing, by server-side device, its own playback of the primary audio content, with said playback being internally aligned with the SRMPT timing information it dispatched to the client to ensure the public sound field remains coherent with the private sound field. The process then repeats from block 2305 as new location data is received.

[0352] FIGURE 24 is a flowchart of method 2400 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal AlSystem and wireless communication sensor. FIGURE 24 applies to a method viewed from client-side.

[0353] The order of the steps in the method 2400 is not constrained to that shown in Figure 24 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0354] According to an embodiment of the invention, a method for rendering a dynamic, context-aware binaural audio experience comprises receiving a data packet containing a plurality of separated audio sub-tracks and a corresponding stream of semantic event markers from a remote device, and processing said event markers with client-side Al Audio Mixing System to automatically adjust mixing and spatialization parameters of the sub-tracks in real-time to reflect the narrative context of the live event.

[0355] In block 2405, capturing, by client-side device controlled by a user, a verbal command, said command comprising a request to initiate a focused- listening session, and transmitting said verbal command as a continuous audio stream along with client-side unique identifier to server-side device for real-time processing. Block 2405 transfers control to block 2410.

[0356] Next, in block 2410, receiving, from server-side device, a response comprising a high-information-density data package and server-side location coordinates (xs,ys,zs). Said package is demultiplexed into its constituent parts: a plurality of discrete audio sub-tracks and a stream of semantic event markers. Block 2410 transfers control to block 2415.

[0357] Next, in block 2415, processing, by client-side Al Audio Mixing System, the stream of semantic event markers to understand the real-time narrative context and to dynamically and automatically adjust a set of mixing and rendering parameters for each discrete audio sub-track to heighten dramatic effect. Block 2415 transfers control to block 2420.

[0358] Next, in block 2420, concurrently determining, by client-side device, its own position (xc,yc,zc) and transmitting client-side packet comprising client-side start host time (CSHT) to server-side device to initiate a synchronization handshake. Block 2420 transfers control to block 2425.

[0359] Next, in block 2425, receiving server-side packet comprising serverside running media play time (SRMPT) and using said timing data to synchronize client-side playback timeline. Block 2425 transfers control to block 2430.

[0360] Next, in block 2430, calculating, by client-side device, an azimuth angle 0, an elevation angle 4>, and a first delay based on the relative positions of the user (xc,yc,zc) and the participant (xs,ys,zs). Block 2430 transfers control to block 2435.

[0361] Next, in block 2435, applying one or more audio filters to each of the one or more audio content sub-tracks, wherein said application is dictated by both the spatialization parameters (0,<t») and the dynamic mixing parameters determined by the Al mixing engine, to generate a final spatialized audio content. For example, the ‘impacts’ sub-track is rendered from the calculated azimuth and elevation, and its gain is momentarily increased in response to a [EVENT: TACKLE] marker. Block 2435 transfers control to block 2440.

[0362] Next, in block 2440, playing back, by client-side device, the fully synchronized and spatialized audio content, thereby providing a dynamically directed, narrative-driven auditory experience that adapts moment-to-moment based on the Al's interpretation of the live event. The process then continues from block 2410.

[0363] FIGURE 25 is a flowchart of method 2500 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 25 applies to a method viewed from server-side.

[0364] The order of the steps in the method 2500 is not constrained to that shown in Figure 25 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0365] According to an embodiment of the invention, a method for generating a high-information-density data stream for remote auditory rendering comprises applying server-side Multimodal Deconstruction Al System to captured sensor data to perform real-time source separation on raw audio to create distinct audio sub-tracks, concurrently performing semantic event detection to generate a corresponding stream of event markers describing a participant's actions, and transmitting said sub-tracks and event markers as a single data packet.

[0366] In block 2505, receiving, by server-side computing device worn by a participant, client-side unique identifier and a continuous audio stream comprising a verbal command from client-side device controlled by a user. Block 2505 transfers control to block 2510.

[0367] Next, in block 2510, processing, by server-side device, the continuous audio stream using a single, server-side Multimodal Deconstruction Al System, wherein said model performs real-time deconstruction of the participant’s auditory reality to generate server-side response without an intermediate text transcription step. Said deconstruction comprises: applying a source separation algorithm to raw microphone input to generate a plurality of discrete audio content subtracks, such as a ‘voice’ sub-track, a ‘footsteps’ sub-track, and an ‘impacts’ subtrack; and concurrently performing semantic event detection by analyzing integrated sensor data, such as data from an Inertial Measurement Unit (IMU) or an Electrocardiogram (ECG) sensor, to generate a stream of machine-readable event markers, such as [EVENT: TACKLE, G-FORCE: 9.2] or [EVENT: SPRINT, HEART_RATE: 180BPM], Block 2510 transfers control to block 2515.

[0368] Next, in block 2515, generating, by server-side device, a high- information-density data package based on said server-side response, said package comprising the plurality of separated audio content sub-tracks and the stream of semantic event markers. Block 2515 transfers control to block 2520.

[0369] Next, in block 2520, concurrently determining, by server-side device, its own real-time location coordinates (xs,ys,zs) via multilateration or another high- precision positioning method. Block 2520 transfers control to block 2525.

[0370] Next, in block 2525, transmitting, by server-side device to client-side device, the high-information-density data package and server-side device's location coordinates (xs,ys,zs). Block 2525 transfers control to block 2530.

[0371] Next, in block 2530, subsequently receiving and processing, by serverside device, client-side packet comprising client-side start host time (CSHT) for initiating a synchronization handshake. Block 2530 transfers control to block 2535.

[0372] Next, in block 2535, transmitting, by server-side device to client-side device, server-side packet comprising server-side running media play time (SRMPT) and other timing data to enable playback synchronization. Block 2535 then continues the process from block 2505 pending new data.

[0373] FIGURE 26 is a flowchart of method 2600 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 26 applies to a method viewed from client-side.

[0374] The order of the steps in the method 2600 is not constrained to that shown in Figure 26 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0375] According to an embodiment of the invention, a method for rendering a synchronized, multi-sensory haptic and auditory experience comprises receiving, by client-side device, a high-information-density data package comprising a plurality of discrete audio sub-tracks and a corresponding stream of semantic event markers; processing, by client-side Al Sensory Experience Application, said stream of semantic event markers to understand the real-time context of an event; dynamically generating, based on said context, a first set of audio mixing parameters and a second, corresponding set of haptic command signals; and rendering a final output by applying said audio mixing parameters to the audio subtracks while concurrently actuating said haptic command signals on a peripheral haptic device in synchronization with the audio playback.

[0376] In block 2605, capturing, by client-side device controlled by a user, a verbal command to join a live sports broadcast, such as a soccer or American football match, and concurrently establishing a communicative coupling with one or more peripheral haptic device. Block 2605 transfers control to block 2610.

[0377] Next, in block 2610, transmitting said command and a unique identifier to server-side device. Block 2610 transfers control to block 2615.

[0378] Next, in block 2615, receiving, from server-side device, a high- information-density data package derived from the live sports broadcast. Said package is generated by server-side Al that processes real-time game feeds, and the package is demultiplexed into its constituent parts: a plurality of discrete audio sub-tracks, such as Crowd-Noise, On-Field-Dialogue, Ball Impacts, and Commentary; and a stream of sports-specific semantic event markers, such as [EVENT:TACKLE], [EVENT: GOAL], and [EVENT: PLAYER_SPRINT]. Block 2615 transfers control to block 2620.

[0379] Next, in block 2620, processing, by client-side Al Sensory Experience Application, the stream of semantic event markers to understand the real-time game action. Based on this understanding, the engine dynamically generates mixing parameters for the audio sub-tracks and corresponding haptic command signals. For example, upon receiving an [EVENT: TACKLE, INTENSITY: HIGH] marker, the engine increases the gain on the Ball Impacts sub-track andgenerates a HAPTIC_JOLT_SHARP command. Block 2620 transfers control to block 2625.

[0380] Next, in block 2625, establishing a virtual position for the user within the sports arena, such as a specific seat in the stands or a "player-follow" mode, and performing a time synchronization handshake with the server to align with the live broadcast timeline. Block 2625 transfers control to block 2630.

[0381] Next, in block 2630, continuously receiving location coordinates for key game elements, such as the ball or specific players, from the server and calculating the real-time azimuth (0), elevation (4>), and first delay (acoustic propagation) for each sound source relative to the user's virtual position. Block 2630 transfers control to block 2635.

[0382] Next, in block 2635, generating a final spatialized audio soundscape by applying audio filters to each audio sub-track. The rendering is dictated by both spatialization parameters and dynamic mixing parameters determined by the Al engine. For example, the spatialization parameters ensure that the sound of a ball impact is rendered from its corresponding on-field location. Block 2635 transfers control to block 2640.

[0383] Next, in block 2640, concurrently generating a final haptic output command. This comprises selecting a haptic effect from a library based on the game event marker and spatializing said effect by directing it to a location on the haptic device corresponding to the on-field event's direction. For example, a tackle occurring on the user's left would trigger a haptic jolt on the left side of a haptic vest. Block 2640 transfers control to block 2645.

[0384] Next, in block 2645, calculating a plurality of compensating delays, including a second delay for audio processing latency and a third delay for haptic system latency, to ensure sensory outputs are synchronized with the live broadcast. Block 2645 transfers control to block 2650.

[0385] Next, in block 2650, rendering the final multi-sensory experience by playing back the dynamic, spatialized audio soundscape while simultaneously actuating the synchronized and spatialized haptic commands on the peripheral haptic device, thereby allowing the user to hear and feel the action of the live sports match with unprecedented immersion. The process may then continue from block 2615.

[0386] In a further embodiment, the method may be applied to the "player- follow" mode established in block 2625, wherein the user selects a particular player to follow via the initial verbal command or a subsequent user input. In this mode, client-side Al Sensory Experience Engine is configured to filter the stream of semantic event markers, prioritizing those that are associated with the selected player's unique identifier. The audio mixing parameters are adjusted to amplify the audio sub-tracks associated with the selected player, such as their on-field dialogue or breathing, while relevant sounds are spatialized from the player's realtime on-field coordinates. Furthermore, the haptic command signals are generated primarily in response to the selected player's direct physical interactions, allowing the user to feel synchronized, spatialized haptic feedback corresponding to events such as the player sprinting, being tackled, or striking a ball. This creates a deeplyimmersive, first-person sensory experience, effectively placing the user in the virtual position of their chosen player on the field.

[0387] FIGURE 27 is a flowchart of method 2700 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 27 applies to a method viewed from server-side.

[0388] The order of the steps in the method 2700 is not constrained to that shown in Figure 27 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0389] According to an embodiment of the invention, a method for generating a high-information-density data stream for enabling a remote multi-sensory experience comprises applying server-side Multimodal Analysis Al System Application to captured, real-time sensor data; concurrently performing, by said Al Sports Analysis Application, real-time source separation on a raw audio feed to create a plurality of distinct audio sub-tracks and performing semantic event detection on the sensor data to generate a corresponding stream of event markers that describe the real-time context; and packaging and transmitting said audio subtracks and said stream of event markers as a single, synchronized data stream to a remote client-side device for processing and rendering.

[0390] In block 2705, receiving, by server-side device, a request from clientside device to join a live broadcast of a sporting event, such as a soccer or American football match. Block 2705 transfers control to block 2710.

[0391] Next, in block 2710, concurrently ingesting, by server-side device, a plurality of real-time data feeds associated with the live sporting event. Said data feeds may comprise the main broadcast video and audio, isolated on-field microphone feeds, player and ball telemetry data (for example, from GPS or RFID tags), and statistical game data. Block 2710 transfers control to block 2715.

[0392] Next, in block 2715, processing, by server-side Al Sports Analysis Application, the plurality of ingested data feeds in real-time to understand the game's narrative context and key actions. Block 2715 transfers control to block 2720.

[0393] Next, in block 2720, generating, by server-side device based on the Al's analysis, a high-information-density data package. Said package comprises two primary components: a plurality of discrete audio sub-tracks, such as Crowd Noise, On-Field Dialogue, Ball Impacts, and Commentary; and a corresponding stream of sports-specific semantic event markers, such as [EVENT: TACKLE, PLAYERJD: 54], [EVENT: GOAL_SCORED], and [INTENSITY: HIGH], Block 2720 transfers control to block 2725.

[0394] Next, in block 2725, appending real-time location coordinates for key game elements, such as the ball or specific players, to the data package, said coordinates being extracted from the player and ball telemetry data feed. Block 2725 transfers control to block 2730.

[0395] Next, in block 2730, continuously transmitting the complete high- information-density data package to client-side device via a low-latency stream. This provides the client with necessary content and metadata to generate a multi-sensory experience that reflects the live game action. Block 2730 transfers control to block 2735.

[0396] Next, in block 2735, subsequently receiving, by server-side device from client-side device, a time synchronization packet, and in response, transmitting a packet containing server-side running media play time (SRMPT) and other timing data to enable client-side device to align its playback with the live broadcast timeline. Block 2735 transfers control to block 2740.

[0397] Next, in block 2740, continuing to process the live game feeds and stream the resulting data package for the duration of the sporting event, thereby providing a continuous, real-time feed for the immersive client-side experience. The process may then await a termination command from client-side device.

[0398] FIGURE 28 is a flowchart of method 2800 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 28 applies to a method viewed from client-side.

[0399] The order of the steps in the method 2800 is not constrained to that shown in Figure 28 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0400] According to an embodiment of the invention, a method for rendering real-time, spatialized, and voice-preserving translated audio experience comprises: receiving, by client-side device, a data stream comprising a primary audio track in a source language and corresponding spatialization metadata; processing, by client-side Multimodal Al Media Application, the source language audio to perform real-time translation into a target language, wherein said application concurrently synthesizes the translated audio to preserve the vocal identity and emotional tone of the original speaker; applying one or more spatialization filters to the resulting translated audio track based on the spatialization metadata to generate a binaural output; and applying one or more compensating delays to ensure the final output is aligned with a synchronized playback timeline.

[0401] In block 2805, establishing a streaming session between client-side device and server-side streaming application, wherein real-time transport protocol is selected based on the low-latency requirements of the interactive Al response loop. Block 2805 transfers control to block 2810.

[0402] Next, in block 2810, continuously receiving, by client-side device, a stream of data packets, wherein said packets comprise a primary audio track in a source language and spatialization metadata, such as the real-time position of the audio source (Xs, Ys, Zs). Block 2810 transfers control to block 2815.

[0403] Next, in block 2815, placing the received data packets into an adaptive jitter buffer to ensure a smooth and stable stream by dynamically adjusting its size based on real-time network conditions. Block 2815 transfers control to block 2820.

[0404] Next, in block 2820, performing a time synchronization handshake with server-side device by exchanging timestamp packets, such as CSHT, SSHT, and SRMPT, to align client-side playback timeline with server-side timeline. Block 2820 transfers control to block 2825.

[0405] Next, in block 2825, after extracting data from the buffer, detecting any lost packets and applying a spatially-aware Packet Loss Concealment (PLC) algorithm to reconstruct the missing audio data from the source language track. Block 2825 transfers control to block 2830.

[0406] Next, in block 2830, decoding the synchronized and repaired data to generate a clean, source language audio channel. Block 2830 transfers control to block 2835.

[0407] Next, in block 2835, processing, by client-side Multimodal Al Media Application, the clean source language audio channel to perform real-time translation. Said application translates the speech from the source language to a target language selected by the user. In a preferred implementation, the multimodal Al concurrently analyzes the vocal characteristics, such as pitch, cadence, and emotional tone, of the original speaker to generate a final translated audio track that preserves the vocal identity of the original speaker. Block 2835 transfers control to block 2840.

[0408] Next, in block 2840, applying one or more audio filters to the final translated audio track, such as convolving the translated audio with a selected pair of Head-Related Impulse Responses (HRIR) corresponding to a calculated azimuth (0) and elevation (cp), to generate a binaural stereo output. Block 2840 transfers control to block 2845.

[0409] Next, in block 2845, calculating, by client-side device, a first delay corresponding to the acoustic propagation time from the source's location to the client's location. Block 2845 transfers control to block 2850.

[0410] Next, in block 2850, calculating, by client-side device, a second delay corresponding to the total signal processing latency introduced by both the translation Al and the spatialization filters. Block 2850 transfers control to block 2855.

[0411] Next, in block 2855, applying a final compensating delay to client-side playback timeline, said final delay accounting for both the acoustic propagation delay (first delay) and the total processing latency (second delay). Block 2855 transfers control to block 2860.

[0412] Next, in block 2860, playing back, by client-side device, the fully processed audio content, which is now rendered as a synchronized, spatialized, and live-translated audio experience in the user's preferred language. Block 2860 then terminates the process.

[0413] FIGURE 29 is a flowchart of method 2900 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 29 applies to a method viewed from server-side.

[0414] The order of the steps in the method 2900 is not constrained to that shown in Figure 29 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0415] According to an embodiment of the invention, a method for enabling a remote, real-time, spatialized audio translation comprises: retrieving or generating, by server-side device, a primary audio track in a source language; concurrently determining the real-time location coordinates of said server-side device to serveas spatialization metadata; transmitting said primary audio track and said location coordinates to a remote client-side device via a continuous data stream; and subsequently responding to a synchronization request from client-side device by transmitting a set of server-side timing data, thereby providing the client with all necessary information to perform a remote translation and synchronized, spatialized rendering.

[0416] In block 2905, establishing, by server-side streaming application, realtime streaming session with client-side device in response to a content request. Block 2905 transfers control to block 2910.

[0417] Next, in block 2910, retrieving or generating, by server-side device, a primary audio track in its original source language, for example, a foreign language. Block 2910 transfers control to block 2915.

[0418] Next, in block 2915, concurrently determining, by server-side device, its own real-time location coordinates (Xs, Ys, Zs), which represent the position of the audio source. Block 2915 transfers control to block 2920.

[0419] Next, in block 2920, commencing the continuous transmission of data packets to client-side device via a low-latency stream. Said stream comprises both the primary audio track in the source language and the corresponding real-time spatialization metadata (Xs, Ys, Zs). Block 2920 transfers control to block 2925.

[0420] Next, in block 2925, subsequently receiving, by server-side device from client-side device, a time synchronization packet, said packet comprising clientside start host time (CSHT). Block 2925 transfers control to block 2930.

[0421] Next, in block 2930, in response to the synchronization request, creating and transmitting, by server-side device, server-side synchronization packet to client-side device. Said packet comprises the original CSHT, server-side start host time (SSHT), server-side end host time (SEHT), and server-side running media play time (SRMPT), thereby providing client-side device with all necessary information to align its playback timeline. Block 2930 transfers control to block 2935.

[0422] Next, in block 2935, continuing, by server-side device, its own local playback of the source language audio track, said playback being aligned with server-side running media play time (SRMPT) to ensure synchronization with client-side device. The process may then continue to stream content and respond to subsequent synchronization requests.

[0423] For example, one or more of audio, video, and another entertainment format can be playing on client-side. For example, one of more of audio, video, and another entertainment format can be played on server-side.

[0424] For example, the steps of the flowchart depicted in Figure 8 may be implemented by one or more of client-side networked environment 105 and clientside device 120.

[0425] For example, the steps of the flowchart depicted in Figure 11 , Figure 12, Figure 13, Figure 14, Figure 15, Figure 16, Figure 17, Figure 18, and Figure 19 may be implemented by one or more of server-side networked environment 110 and server-side device 170.

[0426] For example, client-side master application and server-side master application may be implemented in client-side Multimodal Al System Application and server-side Multimodal Al System Application.

[0427] For example, client-side Multimodal Al System Application and serverside Multimodal Al System Application may communicate directly to other Multimodal Al Systems located in other networks.

[0428] For example, client-side Al Audio Mixing System and server-side Multimodal Deconstruction Al System may communicate directly to other Multimodal Al Systems located in other networks.

[0429] For example, instead of being located in client-side memory 140, one or more of client-side master application 150, client-side Multimodal Al Media Application 152, client-side Multimodal Al System Application 160, client-side Al Audio Mixing System 162 and client-side Al Sensory Experience Application 163 may be located in a section of client-side device 120 other than client-side memory 140.

[0430] For example, instead of being located in server-side memory 175, one or more of server-side master application 180, server-side streaming application 185, server-side Multimodal Al Media Application 190, server-side Al Sports Analysis Application 194, server-side Multimodal Al System Application 195, and server-side Multimodal Deconstruction Al System 199 may be located in one or more of server-side data storage 155 and a section of server-side computing device 170 other than server-side memory 175. For example, instead of being a freestanding component of server-side networked environment 110, server-side memory 175 may be located in server-side computing device 170.

[0431] For example, client-side data storage 145 may be separate from clientside device 120 rather than being comprised in client-side device 120. For example, server-side data storage 155 may be comprised in server-side computing device 170 rather than being separate from server-side computing device 170.

[0432] For example, instead of being two separate entities, client-side device 120 and client-side playback device 125 may be combined in client-side device 120. For example, server-side device 170 and server-side playback device 196 may be included in server-side device 170.

[0433] While the above representative embodiments have been described with certain components in exemplary configurations, it will be understood those skilled in the art that other representative embodiments can be implemented using one or more of different configurations and different components. For example, it will be understood by those skilled in the art that the order of certain fabrication steps and certain components can be altered without substantially impairing the functioning of the invention. The representative embodiments and disclosed subject matter, which have been described in detail herein, have been presented by way of example and illustration and not by way of limitation. It will be understood by those skilled in the art that various changes may be made in the form and details of the described embodiments, resulting in equivalent embodiments that remain within the scope of the invention. It is intended, therefore, that the subject matter in the above description shall be interpreted as illustrative and shall not be interpreted in a limiting sense.

Claims

CLAIMSWhat is claimed is:1 . A method performed by client-side computing device, said method, comprising: receiving, from server-side device, plurality of data packets comprising audio content and associated spatialization metadata, wherein said spatialization metadata defines a location of virtual sound source; determining, by processor of client-side device, a location of client-side device; calculating one or more spatialization parameter based on said spatialization metadata and the determined location of client-side device; and applying one or more audio filters to said audio content based on said one or more spatialization parameter to generate spatialized audio output.

2. The method of claim 1 , further comprising receiving one or more commands from a user, wherein said commands are received via one or more of text input, verbal input, gesture-based input, or touch-based input.

3. The method of claim 1 , further comprising obtaining one or more attributes of a user via authentication Al application.

4. The method of claim 1 , wherein client-side device receives signals from plurality of wireless sensors, and wherein determining the location of client-side device comprises calculating said location via multilateration using said signals.

5. The method of claim 1 , wherein determining the location of client-side device comprises applying sensor fusion algorithm to location-related data received from plurality of wireless communication protocols to compute a single, high-fidelity location coordinate.

6. The method of claim 1 , wherein determining the location of client-side device comprises dynamically selecting optimal positioning protocol from plurality of available positioning protocols based on a set of monitored conditions.

7. The method of claim 1 , further comprising: receiving, from head tracking sensor, real-time head orientation data of a user; and calculating, said one or more spatialization parameter by applying rotation transformation to baseline direction vector based on said real-time head orientation data.

8. The method of claim 7, wherein applying said one or more audio filters comprises generating left channel of spatialized audio output by convolving said audio content with a left-ear Head-Related Impulse Response (HRIR), and generating right channel of the spatialized audio output by convolving said audio content with a right-ear HRIR.

9. The method of claim 7, wherein said one or more audio filters comprise a Binaural Room Impulse Response (BRIR) generated by rendering a Spatial Room Impulse Response (SRIR).

10. The method of claim 7, wherein generating said spatialized audio output is derived from spatial audio representation comprising one or more of higher-order Ambisonics or spherical harmonic coefficients.11 . The method of claim 1 , further comprising: storing a set of previous spatialization filter parameter; calculating a displacement magnitude between a previous and a current location of virtual sound source; andin response to said displacement magnitude exceeding a pre-determined threshold, re-calculating said one or more spatialization parameter and applying smooth transition algorithm between the previous and re-calculated spatialization filter parameter.

12. The method of claim 1 , further comprising calculating a Half-Round-Trip Time (HRT) and a Playback Time Offset (TPO) using a plurality of client-side and server-side timestamps to synchronize client-side playback timeline.

13. The method of claim 1, wherein calculating said one or more spatialization parameter is predictive and further comprises: receiving client-side velocity vector and server-side velocity vector; and generating said one or more spatialization parameter based on calculated future location of client-side device and calculated future location of virtual sound source.

14. The method of claim 13, wherein said predictive calculation is performed for each of plurality of individual audio objects.

15. The method of claim 1 , further comprising: calculating acoustic propagation time delay, and signal processing latency delay; and applying, a final compensating delay to playback of said spatialized audio output based on acoustic propagation time delay and signal processing latency delay.

16. The method of claim 15, further comprising applying one or more loudness compensation signals to said spatialized audio output to render a perceived distance.

17. The method of claim 15, wherein client-side device monitors ambient air temperature and relative humidity using one or more environmental sensors, and wherein said acoustic propagation time delay is calculated using adjusted speed of sound based on monitored ambient air temperature and monitored relative humidity.

18. The method of claim 1 , further comprising processing said spatialized audio output with one or more additional filters selected from the group consisting of: Headphone Transfer Function (HpTF), speed-adaptive equalization filter, noise-cancellation filter, and voice isolator.

19. The method of claim 1 , further comprising placing said received data packets into adaptive jitter buffer, wherein size of said adaptive jitter buffer is dynamically adjusted based on a current application state.

20. Client-side computing system, comprising: one or more processor; and memory storing instruction that, when executed by said one or more processor, cause the system to perform the method of claim 1 .

21. The system of claim 20, wherein the method further comprises applying sensor fusion algorithm to compute a single, high-fidelity location coordinate.

22. The system of claim 20, wherein the method further comprises calculating said one or more spatialization parameter predictively based on one or more velocity vector.

23. The system of claim 20, wherein the method further comprises applying smooth transition algorithm between filter sets.

24. Non-transitory computer-readable storage medium having instruction stored thereon that, when executed by one or more processor of client-side device, cause said client-side device to perform the method of claim 1 .

25. Non-transitory computer-readable storage medium of claim 24, wherein the method further comprises applying sensor fusion algorithm to compute a high-fidelity location coordinate.

26. A method performed by server-side computing device, the method comprising: establishing streaming session with client-side device; providing primary audio content track for transmission; determining real-time location coordinate data for an audio source; encapsulating, within plurality of data packets, said primary audio content track and associated spatialization metadata comprising said real-time location coordinate data; transmitting, to said client-side device, continuous stream of said data packets; and participating in bi-directional time synchronization process with client-side device to enable synchronized playback.

27. The method of claim 26, wherein providing said primary audio content track comprises: receiving, from client-side device, continuous audio stream corresponding to verbal command and receiving application sensor data; and processing said audio stream and said sensor data using a single, natively Multimodal Al model to directly generate said primary audio content track as a response.

28. The method of claim 26, wherein determining said real-time location coordinate data comprises applying sensor fusion algorithm to compute a single, high-fidelity location coordinate.

29. The method of claim 26, further comprising transmitting server-side velocity vector to client-side device to enable predictive spatialization by client-side device.

30. The method of claim 26, wherein participating in said bi-directional time synchronization process comprises: receiving client-side packet comprising client-side start host time (CSHT); querying master clock to obtain server-side running media play time (SRMPT); and transmitting server-side packet to client-side device comprising the CSHT and the SRMPT.31 . The method of claim 26, further comprising generating ambient audio track based on location data received from client-side device.

32. The method of claim 26, further comprising initiating playback of said primary audio content track, said playback being internally synchronized with timing data transmitted to client-side device.

33. Server-side computing system, comprising: one or more processor; and memory storing instruction that, when executed by said one or more processor, cause the system to perform the method of claim 26.

34. The system of claim 33, wherein providing said primary audio content track comprises processing data from client-side device using a single, natively Multimodal Al model.

35. The system of claim 33, wherein the method further comprises applying sensor fusion algorithm to compute a high-fidelity location coordinate.

36. The system of claim 33, wherein the method further comprises transmitting serverside velocity vector to client-side device.

37. Non-transitory computer-readable storage medium having instruction stored thereon that, when executed by one or more processor of server-side device, cause server-side device to perform the method of claim 26.

38. Non-transitory computer-readable storage medium of claim 37, wherein providing said primary audio content track comprises processing data using a single, natively Multimodal Al model.

39. Non-transitory computer-readable storage medium of claim 37, wherein the method further comprises applying sensor fusion algorithm.

40. Non-transitory computer-readable storage medium of claim 37, wherein the method further comprises transmitting server-side velocity vector.

41. Non-transitory computer-readable storage medium of claim 37, wherein participating in said bi-directional time synchronization process comprises querying master clock.

Citation Information

Patent Citations

  • Time scaling of multi-channel audio signals

    US20080114606A1

  • Method and System for Mobile Trajectory Based Services

    US20090061903A1

  • Device, system, and method of three-dimensional spatial user authentication

    US20160300054A1

  • Binaural Sound in Visual Entertainment Media

    US20180109900A1

  • System and method for modifying gameplay according to user geographical location

    US20180272235A1