System and method for real-time multi-object tracking, synchronization, and spatialization of media content using multimodal ai system and object recognition ai application

The multimodal AI system and object recognition application addresses real-time tracking and spatialization challenges by using time delays and sound pressure level differences to generate synchronized and personalized binaural audio, ensuring accurate and immersive audio experiences across varied environments.

WO2026039572A1PCT designated stage Publication Date: 2026-02-19FERRER JULIO
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/041885
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-13
Filing Date
2025-08-13
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Existing technologies face challenges in efficiently tracking, synchronizing, and spatializing audio and media content in real-time across multiple environments and devices, particularly in venues with varying conditions, leading to desynchronization and suboptimal spatial audio experiences.

Method used

A multimodal AI system and object recognition AI application that leverages time delays and sound pressure level differences to determine sound source orientation and position, using spatial transfer functions and learning-based models to generate personalized and synchronized binaural audio, while integrating multiple sensor data for enhanced accuracy and adaptability.

Benefits of technology

Enables real-time, context-aware, and synchronized multi-object tracking and spatialization of audio and media content across diverse environments, providing immersive and accurate audio experiences by compensating for acoustic and processing delays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025041885_19022026_PF_FP_ABST
    Figure US2025041885_19022026_PF_FP_ABST
Patent Text Reader

Abstract

A system and method for synchronizing and spatializing media between client device and server device. Client-side artificial intelligence application determines the position of server device by fusing data from at least two sources: visual data obtained from a camera and acoustic data obtained from a predetermined reference signal emitted by server device. This data fusion enables continuous spatial tracking, including during periods of visual occlusion. A timestamp-based protocol is used to synchronize media playback. Based on the tracked position, client device applies directional audio filters, such as Head-Related Impulse Response (HRIR), to render audio that appears to originate from server device's physical location. The method further compensates for acoustic propagation and signal processing delays to ensure accurate spatiotemporal alignment.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEM AND METHOD FOR REAL-TIME MULTI-OBJECT TRACKING, SYNCHRONIZATION, AND SPATIALIZATION OF MEDIA CONTENT USING MULTIMODAL Al SYSTEM AND OBJECT RECOGNITION Al APPLICATION

[0001] The present application claims the priority benefit of U.S. provisional patent applications No. 63 / 682,465, No. 63 / 682,449, and No. 63 / 682,456 filed Aug. 13, 2024 and entitled “System and Method for Real-Time Multi-Object Tracking, Synchronization, and Spatialization of Media Content Using Multimodal Al System and Object Recognition Al Application,” the disclosure of which is incorporated herein by scope of the appended claims.

[0002] This invention relates to medium, method, and system for the multiobject tracking, synchronization, and spatialization of one or more of media content using one or more of Multimodal Al (artificial intelligence) system and Object Recognition Al Application with one or more electronic device.Particularly, the invention relates to a medium, method, and system for the multiobject tracking, synchronization, and spatialization of one or more of audio and media content using one or more of Multimodal Al System and Object Recognition Al Application with one or more electronic device. More particularly, the invention relates to a medium, method, and system for real-time multi-object tracking, synchronization, and spatialization of one or more of audio and media content using one or more of Multimodal Al System and Object Recognition Al Application with one or more electronic device. Even more particularly, the invention relates to a medium, method, and system for real-time multi-object tracking, synchronization, and spatialization of one or more of audio and media content using one or more of Multimodal Al System and Object Recognition Al Application with one or more electronic device within and around one or more of venue, public, private, office, open space, educational institution, store, storefront, home, vehicle, chat, gaming, virtual, and an other environment. Specifically, the invention relates to a medium, method, and system for real-time conversational, transactional, and ambient multiobject tracking, synchronization, and spatialization of one or more of audio and media content using one or more of Multimodal Al System and Object Recognition Al Application with one or more electronic device within and around one or more of venue, public, private, office, open space, educational institution, store, storefront, home, vehicle, chat, gaming, virtual, and an other environment. More specifically, the invention relates to a medium, method, and system for real-time conversational, transactional, and ambient multi-object tracking, synchronization, and spatialization of one or more of audio and media content using one or more of Multimodal Al System and Object Recognition Al Application with one or more electronic device within and around one or more of venue, public, private, office, open space, educational institution, store, storefront, home, vehicle, chat, gaming, virtual, and an other environment over a communications network. The invention provides alternative methods for real-time conversational, transactional, and ambient multi-object tracking, synchronization, and spatialization of static and non-static of one or more client-side device, server-side device, and Object Recognition Al Application.

[0003] According to embodiments of the invention, platform is provided for real-time multi-object tracking, synchronization, and spatialization of one or more of one or more audio content, media content, mixed audio content, and other data input during user’s interaction with one or more of one or more Multimodal Al System and Object Recognition Al Application. For example, multi-object tracking may determine sound source by leveraging the time delay between signals received by one or more Object Recognition Al Application. For example, multiobject tracking may determine the sound pressure level (SPL) difference between user’s head and static and non-static object. For example, multi-object tracking may determine the SPL difference between static and non-static user’s head and one or more object. For example, one or more time delay, and SPL difference may determine the orientation and position between user’s head and object.

[0004] For example, data input may comprise text, audio, video, images, physics-informed, metadata tags, language models, modalities, sensory inputs, and other data input.

[0005] For example, audio content may comprise of one or more object-based audio, channel-based audio, scene-based audio, metadata, and an other audio signal input. For example, the metadata may include the position and gain of one or more of one or more object-based, scene-based, and channel-based audio signals, and other rendering metadata. For example, audio content may comprise audio generated by Multimodal Al System. For example, audio content may comprise synchronized audio generated by Multimodal Al System. For example, audio content may comprise spatialized audio generated by Multimodal Al System. For example, audio content may comprise one or more of one or more synchronized and spatialized audio generated by Multimodal Al System. For example, audio content may comprise one or more of one or more audio content, mono audio content, stereo audio content, multiple audio content, binaural audio content, ambisonic audio content, and alternative audio content.

[0006] For example, binaural or spatial audio content may be produced by processing an audio signal with a spatial transfer function. Said spatial transfer function may comprise, without limitation, a head-related impulse response (HRIR), a head-related transfer function (HRTF), a binaural room impulse response (BRIR), or a Spatial Room Impulse Response (SRIR). A key objective of this processing is often to improve the sense of externalization for the listener. To this end, a filter may be used to modify the interaural coherence (IC), a critical perceptual cue which can be calculated from said spatial transfer function. Furthermore, said HRIR, HRTF, or SRIR may be personalized, for example, measured for or adapted to a specific user, or may be non-personalized, for example, based on a generic acoustical model or mannequin.

[0007] In a particular embodiment, a target coherence value is established to serve as a benchmark for a realistic, externalized sound field. This target, known as the diffuse-field interaural coherence, is calculated by averaging the interaural coherence values derived from a multiplicity of HRTFs, wherein each HRTF corresponds to a different direction of sound arrival. Said BRIR, SRIR, or afunctionally similar representation of a reverberant environment, may be derived from one or more methods, including: a) direct physical measurement in a real environment; b) physics-based numerical simulation, such as methods based on virtual 3D models of an environment; c) microphone array processing and rendering techniques, such as the Spatial Decomposition Method (SDM), Spatial Impulse Response Rendering (SIRR), or Rendering from Arbitrary Microphone Array (REPAIR); or d) Learning-based modeling, wherein a model is trained to generate or predict an acoustic response.

[0008] For example, said learning-based modeling may include, but is not limited to, models such as Differentiable Feedback Delay Networks (FDNs) which are optimized to a target response, Generative FDNs (GFDNs) which generate a response based on a conditioning input, Neural Acoustic Fields (NAFs) which learn a continuous representation of a sound field, or Spatialized Scattering Delay Networks. Notably, certain parametric or generative models, such as Directional Feedback Delay Networks (DFDNs), can directly synthesize a spatially reverberant audio signal in response to an input, thereby obviating the need to explicitly generate, store, and convolve with a complete BRIR or SRIR data set.

[0009] Once derived or otherwise obtained, the application of said one or more spatial transfer functions or the execution of said generative models to produce one or more binaural or spatial audio content may be performed by one or more real-time operations. Such operations include not only the primary convolution with an HRTF, BRIR, or SRIR, but may also include the direct synthesis from a parametric model or the application of said filter designed to steer the signal’s interaural coherence towards the target diffuse-field value to boost externalization. These operations may include, but are not limited to, convolution or one or more matrix operations.

[0010] For example, a position of client-side device relative to server-side device may be calculated using multilateration operation, angle of arrival, and angle of departure methods, and an other geometric positioning technique. For example, audio content may utilize one or more of one or more bandpass transfer function, headphone transfer function (HpTF), compensation filters, speed- adaptive equalization filters, loudness normalization, noise-cancellation, speech enhancement, voice isolator, Al-powered volume management, Al-based filters and an other DSP algorithm, so as to regularize audio signals, and adjust one or more timber, spectral cue, localization, and an other attribute of sound. For example, audio content may comprise one or more of one or more speech, ambient, music, sound effects, silence, noise, and an other sound.

[0011] For example, media content may comprise one or more of one or more audio content and media content. For example, media content may comprise generated by Multimodal Al System. For example, media content may be digital content. For example, media content may be an other type of content other than digital content. For example, Multimodal Al Media Application may produce audio and media content. For example, Multimodal Al Media Application may produce compressed and uncompressed audio and media content. For example, Multimodal Al Media Application may translate media content from foreign and non-foreign languages. For example, media content may comprise facialanimation generated by Multimodal Al Media Application. For example, media content may comprise of embed digital signatures. For example, alternative media content may comprise of media content with associated metadata describing one or more sound source positioning, sound pressure, multiplicity of audio objects and channel sound signals, and an other spatial property.

[0012] For example, audio and media content may be generated by one or more server-side device. For example, audio and media content may be generated by one or more client-side device. For example, audio and media content may be accessed by user using one or more client-side device and server-side device.

[0013] For example, audio and media content may be accessed by user using wire earbuds, wireless earbuds, wire headphones, wireless headphones, brain on a chip device, ear transducer, bone conductor transducer, and an other apparatus with one or more speakers. For example, audio and media content may be accessed by user using wire virtual reality (VR) headset, wireless VR headset, wire spatial computing headset, wireless spatial computing headset, wire computing glasses, wireless computing glasses, and an other apparatus with one or more speakers. For example, audio content may be accessed by user using speaker system. For example, audio content may be accessed by user using one or more of one or more server-side device and speaker system. For example, audio content may be accessed by user using one or more of one or more client-side device and speaker system.

[0014] For example, one or more camera may access and localize audio and media content using multimodal Al system and Object Recognition Al Application. For example, one or more camera may access and localize audio and media content using single camera, two camera, instance segmentation, object recognition model, or an other computer vision instrument. For example, single camera object recognition may comprise of single camera, deep neural networks (DNN), location Al sensor data, Al based algorithm and an other data input that predicts ground truth and distances in complex environments. For example, two camera object recognition may mimic human vision by calculating the disparity between corresponding points and determine depth information. For example, instance segmentation may measure the coordinates and pixel dimensions of specific objects. For example, Object Recognition Al Application may generate one or more estimates of distance to one or more objects using object boundary boxes and other localizing techniques. For example, Object Recognition Al Application may calibrate, determine detailed vision disparities, provide known and unknown baseline distances, determine object velocity, determine focal lengths of camera, determine lighting conditions, determine object weight and size, and an other function of object recognition.

[0015] For example, Object Recognition Al Application may estimate the velocity of an object by tracking its position across multiple frames in a video stream. For example, Object Recognition Al Application may determine the frame rate and scale of a scene to compute speed. For example, Object Recognition Al Application may detect the object, maintain its identity across frames, and calculate the change in object’s position over time. For example, Object Recognition Al Application may mitigate occlusion and perspective distortion that affects objectrecognition accuracy using perspective transformation, and an other technique for object recognition integrity.

[0016] For example, one or more Object Recognition Al Application may access and localize audio and media content using SPL. For example, two or more Object Recognition Al Application may be located in a camera apparatus on user’s head. For example, one or more Object Recognition Al Application may be at a location other than on user’s head. For example, one or more Object Recognition Al Application may improve the accuracy and resolution of direction in 3D space. For example, the resolution and sampling rate of Multimodal Al Media System recording may increase to improve accuracy. For example, one or more Object Recognition Al Application may be located on user’s person, headset, client-side device, and server-side device, using measurement compensation techniques.

[0017] For example, audio and media content may be processed, analyzed, and modified using one or more of one or more application sensor. For example, an application sensor may comprise microphone, visual sensor, thermal sensor, vibratory sensor, camera, LIDAR, radar, position sensor, biomarker sensor, ambient light sensor (ALS), electrocardiogram sensor (ECG), and an other sensor. For example, position sensor may comprise one or more of one or more head motion tracking sensors, infrared camera to recognize hand gestures, understand depth or map a room, eye motion tracking sensors, haptic sensors, magnetic sensors, and an other position sensor. For example, ALS may comprise time-of- day data. For example, ECG may comprise electrical activity of the heart. For example, application sensor may apply data to Multimodal Al System as data input for pattern recognition, predictive analysis, automated response, real-time monitoring, alerts and notifications, performance optimization, safety monitoring, user personalization, multi-sensor integration, behavior analysis, decision-making, deep learning, predictive health, and an other automated task.

[0018] For example, mixed audio content may expand, combine, edit, delete, fuse, transcribe, and the like one or more of one or more audio content track, ambient audio track, and data input to create one or more mixed audio content. For example, the mixed audio content may adjust the SPL, panning, speed- adaptive equalization (EQ), Al-powered volume management, effects, and dynamics of individual audio elements to create balance and polish.

[0019] For example, audio and media content may be accessed by user over a network. For example, audio and media content may be accessed by user over one or more of Internet, private virtual network, extranet, fiber optic network, wide area network (WAN), local area network (LAN), Bluetooth network, ultra-wideband network (UWB), RF, wired network, wireless network, near field magnetic inductance, sonic waves, ultrasound waves, infrared waves, wireless mesh protocol, satellite internet constellation system, and an other type of network. For example, audio and media content may be accessed by user via TCP, UDP, WebSocket, WebRTC connection, and an other real-time communication. For example, content may be accessed by user within virtual boundaries around a geographical area (geo-fencing). For example, the virtual boundaries may be coordinates of a real-world geographical area.

[0020] Embodiments are described for medium, method, and system for realtime multi-object tracking, synchronization, and spatialization using Multimodal Al System to process and integrate in real-time audio content and media content in, to, and from one or more Object Recognition Al Application and electronic device. For example, Multimodal Al System comprises multiple sensor data, applications, models, and data input to generate rich contextual outputs that enhance complex, comprehensive, unifying, and nuanced understanding and accuracy. For example, Multimodal Al System may generate more adaptive, natural, context-aware, and human-like interactions using Multimodal Al System Application. For example, Multimodal Al System Application may comprise of one or more of client-side voice-activated Al application, voice-to-text Al application, audio-to-client-device application, authentication Al application, on-device Al foundation model, and an other application. For example, Multimodal Al System Application may comprise of one or more of server-side text-to-speech Al application, text-to-audio Al application, audio-to-dialog Al application, location Al application, dialog supervisor Al application, privacy guardrail Al application, marketing Al application, server-based Al model, and an other application.

[0021] For example, Multimodal Al System Application is architected for low- latency, bidirectional data streaming, and may comprise one or more of streaming text-to-speech (TTS) engine for dynamically generating audio from text inputs as they are received, and real-time speech-to-text (STT) transcription engine for converting live audio inputs into a corresponding text stream. For example, in another real-time embodiment, the Multimodal Al System Application may use a Multimodal transformer model that processes one or more of one or more physics-informed text, verbal, gesture, touch, vital, facial, visual, audio, and an other data inputs natively.

[0022] As used herein, the term "natively Multimodal Al model" refers to a single, unified artificial intelligence architecture that processes input streams of multiple data modalities (such as audio and sensor data) and directly generates an output stream without conversion to an intermediate, human-readable format like transcribed text. This end-to-end processing is distinct from prior art "cascaded" systems that use a pipeline of separate unimodal models (e.g., speech-to-text followed by text-to-speech). The native model achieves this efficiency by converting input modalities into a shared mathematical latent space, where cross-modal reasoning occurs before the output is directly decoded into the desired modality, thereby preserving nuances like emotional tone that are lost in text conversion.

[0023] For example, Multimodal Al System Application may generative Al, natural language processing (NLP), Transformers, Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), Diffusion Models, Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), Transformers, Physics Informed Neural Networks (PINN), Large Behavior Models (LBM), Large Language Models(LLM), mixture-of-experts (MoE), quantized model, machine learning Al application and the like, to generate informed, integrated, cohesive, insightful, and leveraged data across various modalities.

[0024] In particular, according to further embodiments of the invention, one or more of one or more audio content, media content, mixed audio content, and an other data input synchronize in real-time with static and non-static on-screen media produced by Multimodal Al Media Application. For example, on-screen media may be projected using one or more digital projector, 3D digital display, 3D projection mapping, multi-view stream display, hologram projector, led screen, billboard, robotics, vehicle, computer, television, console, headset, and an other projection device. For example, on-screen media produced may be using one or more lipsyncing general-purpose robotic humanoid, an other robot with a video display, and an other reception device. For example, one or more of one or more speaker and microphone may be integrated into various locations on a humanoid robot, said locations comprising the head and other non-head portions of the robot's chassis. For example, synchronization is performed by one or more of one or more computer algorithm, clock, and sensor. Alternatively, or additionally, Multimodal Al System Application receives synchronization input from one or more of one or more user and designer.

[0025] According to further embodiments of the invention, Multimodal Al Media Application may monitor, expand, combine, edit, delete, fuse, transcribe, and the like, audio content, media content, mixed audio content, and an other data input. For example, Multimodal Al System Application generates communication between server-side and client-side device to produce real-time on-going dialog within and around one or more of venue, public, private, office, open space, educational institution, store, storefront, home, vehicle, chat, gaming, virtual, and an other environment. For example, Multimodal Al System Application may comprise one or more location Al application. For example, location Al application tracks multitude of client-device to determine the near-precise location of each client device controlled by a user in public environment.

[0026] In certain embodiments, a location Al application may determine a geospatial and environmental context for one or more server-side and client-side device. The system may be configured to process a plurality of input signals to determine said context, said input signals comprising data from one or more of: a cellular network, a Global Positioning System (GPS), a satellite internet constellation system, Ultra-Wideband (UWB) radios, a gyroscope, a barometer, an accelerometer, raw sensor data and an other position locator to assist in determining location of server-side and client-side device. According to yet other embodiments of the invention, the communication environment may include multiple venues within different locations and / or time zones viewing the same media content simultaneously, maybe even synchronously.

[0027] Furthermore, the input signals may comprise data from a wireless local area network (WLAN), including signal characteristics received from a plurality of radio links across different frequency bands, such as 2.4 GHz, 5 GHz, and 6 GHz, managed through a protocol that enables Multi-Link Operation (MLO). The unique signature of signal strengths, latencies, and interference levels across these multiple concurrently managed links can be used as a distinct input to enhance the accuracy and reliability of the context determination.

[0028] According to other embodiments of the invention, the communication environment is configured to leverage said MLO to provide high-performance data transmission between device. MLO enables a client device and a network access point to establish and concurrently utilize multiple radio links. This capability is harnessed by the system to facilitate demanding use cases, such as the synchronized delivery and presentation of media content to multiple venues, which may be situated in different physical locations and / or time zones. The system may be configured to dynamically manage data flows across these multiple links to optimize for specific performance objectives, such as throughput, latency, or reliability.

[0029] In further embodiments, the system is configured to select and operate in one of a plurality of MLO modes based on application requirements or real-time network analysis. These modes may include: a) a Simultaneous Multi-Link (SML) mode, wherein the system is configured to aggregate the bandwidth of at least two of the established radio links by transmitting and / or receiving data packets simultaneously across them. This mode is preferentially employed for applications requiring maximum data throughput and minimal latency, such as the synchronized streaming of ultra-high-definition, for example, 8K or higher, media content, b) an Alternating Multi-Link (AML) mode, wherein the system is configured to monitor the performance of the established radio links and dynamically select a single optimal link for data transmission at any given moment. The system can perform seamless, near-instantaneous switching between links to avoid transient interference or congestion. This mode is preferentially employed to ensure robust and uninterrupted connectivity, which is critical for maintaining near-precise synchronization of media playback across participating venues, thereby preventing desynchronization events caused by network degradation on any single frequency band.

[0030] For example, Multimodal Al System Application may comprise privacy guardrail Al application monitored by one or more of application, employees, contractors, and the like within and around one or more of venue, public, private, office, open space, educational institution, hospital, museum, stadium, arena, concert, sport event, convention center, hotel, store, storefront, home, vehicle, chat, gaming, virtual, and an other environment. For example, Multimodal Al System Application may comprise one or more authentication Al application. For example, authentication Al application may comprise one or more of one or more physical biometric, behavior biometric, face identification, finger identification, eye identification, and an other identifier.

[0031] For example, marketing Al application may generate real-time multiobject tracking, synchronization, and spatialization of advertisement and promotion to integrate with audio and media content. For example, marketing Al application may comprise one or more of one or more audience segmentation monitoring, behavior targeting, personalization and customization Al application, programmatic Al application, analytics Al application, fraud prevention Al application, and an other function of marketing. For example, marketing Al application determines user demographics, interests, sentiments, behaviors and other criteria to predict the effectiveness of advertisement and promotion for adesired outcome. For example, marketing Al application automates ad inventory management, strategy optimization, churn prevention, A / B testing, product review, social media posting, marketing accounting, and an other task automation.

[0032] According to further embodiments of the invention, audio content, media content, mixed audio content, and an other data input may be generated by user’s text commands, verbal commands, gesture-based commands, touch-based commands, vital signs, facial expressions, eye position, and an other signal. According to yet other embodiments of the invention, audio content, media content, mixed audio content, and an other data input may be processed, saved, stored, and used to train and finetune Multimodal Al System, Multimodal Al System Application, Multimodal Al Media Application, Object Recognition Al Application, Al Audio Mixing System and Multimodal Deconstruction Al System. These features are available in one or more in-venue settings, on-edge settings, remote settings, and others. These features are available on multiple platforms and device. For example, the on-edge setting may receive and process the commands on the same electronic device. For example, the remote settings may receive and process the commands on different electronic device. For example, client-side device and server-side device may receive and process commands on the same electronic device. For example, communication between client-side device and server-side device may comprise one or more of one or more unicast transmission, multicast transmission, and broadcast transmission.

[0033] According to yet further embodiments of the invention, Al Audio Mixing System is configured to receive and demultiplex a data stream originating from a remote server-side device, said stream comprising a plurality of discrete audio subtracks and a corresponding stream of semantic event markers associated with a participant in a live event. The engine utilizes said processor to execute instructions that continuously process the stream of semantic event markers to interpret a real-time narrative context of the participant's actions and physiological state. In response to the interpreted narrative context, the engine is further configured to dynamically and automatically adjust a set of mixing and rendering parameters for each of the discrete audio sub-tracks, said parameters comprising one or more of gain, equalization, audio compression, and the selective application of spatialization filters. The engine thereby generates a final, coherent binaural audio output that is not a static representation of the captured audio, but is instead a context-aware auditory experience that is cinematically adapted in real-time to heighten the dramatic effect of the actions and states described by the semantic event markers.

[0034] According to yet further embodiments of the invention, a Multimodal Deconstruction Al System may be one or more of a specialized deep learning architecture, for example, a transformer-based or a hybrid convolutional-recurrent neural network (CRNN) architecture, specifically designed for real-time, on-device processing. The model is configured to receive and process a plurality of heterogeneous, time-synchronized data streams simultaneously, said streams comprising: the continuous audio stream of a user's verbal command received from a client device, the raw audio stream captured by server-side device's own microphones, a time-series data stream from an integrated Inertial MeasurementUnit (IMU), and a time-series data stream from one or more integrated biometric sensors, such as an Electrocardiogram (ECG) sensor. A key aspect of the model is its ability to directly map the user's vocal command waveform to an internal intent vector without an intermediate text transcription step, thereby reducing latency and preserving vocal nuance. Concurrently, the model performs a real-time deconstruction of the participant’s reality by utilizing two integrated sub-networks: a source separation sub-network that deconstructs the participant’s raw microphone audio into a plurality of discrete audio content sub-tracks, such as ‘voice,’ ‘footsteps,’ and ‘impacts’; and an event detection sub-network that processes the IMLJ and biometric sensor data to identify and classify significant occurrences, generating a corresponding stream of semantic event markers. The model then utilizes the derived user intent vector to selectively package or prioritize the generated sub-tracks and event markers into the final, high-information-density data package that constitutes server-side response.

[0035] According to yet further embodiments of the invention, memory data store information may comprise one or more of one or more audio content, media content, mixed audio content, and other data input before, during, and after user’s interaction with one or more Multimodal Al System. For example, memory data store may comprise server-side master application, client-side master application, server-side streaming application, client-side Multimodal Al System Application, client-side Multimodal Al Media Application, server-side Multimodal Al System Application, server-side Multimodal Al Media Application, client-side Object Recognition Al Application, server-side Object Recognition Al Application, clientside Al Audio Mixing System, server-side Multimodal Deconstruction Al System and an other application. For example, memory data store may be located in one or more client-side device. For example, memory data store may be located in one or more server-side device. For example, memory data store may be located in one or more of one or more client-side device and server-side device.

[0036] For example, server-side master application and client-side master application are designed for synchronizing and managing data network packages and acting as the central coordinator in a distributed system. Server-side master application and client-side master application ensure data consistency, reliability, and efficient communication across client device, servers, and other device. For example, server-side master application and client-side master application may manage the cache, encoding, decoding, compression, and encryption of data packages to optimize network usage and secure data in transit.

[0037] For example, server-side streaming application may be responsible for ingesting audio content, media content, mixed audio content, and other data inputs, processing, and streaming to client device while ensuring synchronization. For example, server-side streaming application may encode, buffer, transcode, distribute content, and generate real-time synchronization. For example, serverside streaming application may stream multiple copies of the data for redundancy.

[0038] According to further embodiments of the invention, the memory data store may track and store the state and action of the user. For example, the state represents a user’s current situation or configuration using a client device at a given time. For example, a state may comprise past purchase activity data, presentpurchase activity data, stock market data, medical data, foreign language command, or an other state. For example, present purchase activity data may comprise strokes and clicks a client device controlled by a user records as the user selects, stores, and manages items in an e-commerce cart intended for purchase. For example, a state may comprise environmental conditions data, location tracking data, health bio signs data, and text sentiment data of the user. For example, an action may comprise data regarding the decision and the steps taken by the user using a client device as a result of a prompt presented by the Multimodal Al System application. For example, a dialog supervisor Al application manages the context, strategy, and dialog flow with the user. For example, a dialog supervisor Al application may comprise intent recognition, context management, response generation, decision making, flow control, error handling, policy management, customer support, natural language understanding (NLU), natural language generation (NLG), personalization, e-commerce, and other dialog functions. For example, a dialog supervisor Al application may recommend that the user use a client device to relocate to an alternative location, close a purchase order, encourage the user’s mood, or take another action to achieve a specific objective. For example, a dialog supervisor Al application may offer a reward to a user who achieves a desired outcome influenced by a dialog supervisor Al application prompt. For example, a dialog supervisor Al may process one or more user preferences to automatically select a language for the media content, wherein said selection comprises performing a real-time translation of the media content or selecting a version of the media content in a preferred language.

[0039] According to further embodiments of the invention, dialog supervisor Al application may comprise generative Al, natural language processing (NLP), Transformers, Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), Transformers, Physics Informed Neural Networks (PINN), Large Behavior Models (LBM), Large Language Models(LLM), mixture-of-experts (MoE), quantized model, machine learning Al application, and the like that manages conversation of state and actions, and an other algorithm that manages conversation. For example, a dialog supervisor Al application may ask to determine user intent, query the database, provide recommendations based on profile and memory data, analyze sensor data, identify data required to act, track, execute transactions, provide answers, and perform other tasks. For example, a dialog supervisor Al application may comprise multiple agents that make multiple calls to the same dialog supervisor Al application.

[0040] In one embodiment, the first delay is determined indirectly via a modelbased calculation that leverages the device's location coordinates. Client-side device first calculates the straight-line Euclidean distance (D) between its own realtime coordinates (Xc,Yc,Zc) and the coordinates of server-side device (Xs,Ys, Zs ). The first delay is then calculated by dividing this distance by the speed of sound (vsound). While some implementations use a standard, fixed value for the speed of sound, such as 343 meters per second, more advanced implementations achieve greater precision. In these more sophisticated versions, client-side device uses one or more onboard environmental sensors to continuously monitor localconditions. By tracking both the ambient air temperature and the relative humidity, the device can calculate a much more accurate, real-time value for the speed of sound. This comprehensive calculation, which accounts for how both temperature and water vapor affect the air's physical properties, significantly reduces errors and provides a more precise delay calculation compared to using a standard or temperature-only value. In the context of the present disclosure, the "first delay" represents the acoustic propagation time, which is a calculated simulation of the real-world time required for a sound wave to travel through a physical medium, such as air, from the location of a sound source, such as server-side device, to the location of a listener, such as client-side device. The primary purpose of calculating and applying this first delay is to provide a powerful psychoacoustic cue for distance perception, thereby creating a realistic and immersive spatial audio experience. The determination of this first delay may be accomplished by one or more of the following methods.

[0041] In another embodiment, the first delay is determined by a direct active measurement using a reference signal. In this method, server-side device is configured to emit a dedicated, known reference sound signal, which may be an audible tone or an inaudible ultrasonic pulse. At the moment of emission, serverside device records an emission timestamp (Temit) and transmits this timestamp to client-side device over a low-latency data channel. Client-side device's microphone continuously listens for said reference signal. Upon detection, clientside device records an arrival timestamp (Tarrive) on its own synchronized clock. The first delay is then calculated directly as the difference between the arrival and emission timestamps (FirstDelay=Tarrive-Temit). This method provides a direct measurement of the actual acoustic travel time within the specific physical environment.

[0042] In a further embodiment, the first delay is determined by a direct passive measurement using cross-correlation. In this method, client-side device compares the "clean" audio track received from server-side device (the source signal) with the audio that is contemporaneously recorded by its own microphone from the physical environment (the received signal). Client-side device's processor is configured to perform a mathematical cross-correlation between the source signal and the received signal. The output of the cross-correlation function reveals the time offset at which the two signals are most similar. This time offset corresponds directly to the acoustic propagation time, or the first delay.

[0043] Regardless of the method used for its calculation, the resulting first delay value is subsequently used in the audio processing pipeline. It is applied to the final binaural stereo output, often in conjunction with loudness compensation signals, to render a believable sense of distance. Furthermore, the first delay is a critical component of a final compensating delay, wherein it is combined with a calculated second delay (corresponding to processing latency) to ensure the final audio playback is aligned with the system's synchronized event timeline, thereby creating an audio experience that is both physically realistic and technically synchronized.

[0044] According to further embodiments of the invention, server-side device for real-time multi-object tracking, synchronizing, and spatializing of one or moreaudio and media content using a Multimodal Al System and Object Recognition Al Application includes: a processor; data storage operably connected with the processor; memory operably connected with the processor, the memory comprising one or more of server-side master application, server-side streaming application, server-side Multimodal Al Media Application, server-side Multimodal Al Media System, server-side Object Recognition Al Application and server-side Multimodal Deconstruction Al System. Server-side playback device and serverside sensor system operably connected with the processor; server-side computing device operably connected with the processor and configured to communicate over one or more networks and input / output (I / O) device with one or more sensor systems, speaker systems, and another audio transducer; server-side local interface operably connected with the processor and configured to communicate over a network with client-side device under a user's control. Server-side master application is configured to receive a message or packet comprising media content selected by the user over the network from client-side device. Server-side master application is configured to obtain the user’s preferences. Server-side master application is configured to generate personalized audio and media content using user preferences. Server-side master application is configured to receive and process real-time data packets. Server-side local interface is configured to transmit to client-side device via the network both the generated audio and media content and server-side timing information, wherein said timing information and content enable client-side device to substantially synchronize and spatialize its playback with a corresponding playback by server-side playback device.

[0045] It is to be understood that the embodiments described herein are illustrative and not restrictive. The scope of the invention is defined by the appended claims rather than the foregoing descriptions, and changes that come within the meaning and range of equivalency of the claims are therefore intended to be embraced therein. The various functions and processes described herein may be implemented in hardware, software, or a combination thereof. When implemented in software, the functions may be stored on or transmitted over as one or more instructions or code on a non-transitory computer-readable medium.DESCRIPTION OF THE DRAWINGS

[0046] FIGURE 1 is a schematic block diagram of a networked environment for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and Object Recognition Al Application.

[0047] FIGURE 2 is a schematic block diagram of server-side computing device in an alternative configuration of a networked environment for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and Object Recognition Al Application.

[0048] FIGURE 3 illustrates a system comprising: one or more client-side device, wherein at least one of the one or more client-side device is embodied as wireless glasses configured to be worn by a human user; and server-side device embodied as a humanoid robot, wherein the one or more client-side device are configured to synchronize with server-side device.

[0049] FIGURE 4 is a flowchart for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and Object Recognition Al Application. FIGURE 4 applies to a method using Multimodal Al, as viewed from client-side.

[0050] FIGURE 5 is a flowchart for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and Object Recognition Al Application. FIGURE 5 applies to a method using Multimodal Al, as viewed from server-side.

[0051] FIGURE 6 is a flowchart for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and Object Recognition Al Application. FIGURE 6 applies to a method using a single, natively Multimodal Al, as viewed from client-side.

[0052] FIGURE 7 is a flowchart for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and Object Recognition Al Application. FIGURE 7 applies to a method using a single, natively Multimodal Al, as viewed from server-side.

[0053] FIGURE 8 is a flowchart for client-side device to perform tracking, synchronization, and spatialization of media content of client-side playback with server-side playback using a Multimodal Al System and Object Recognition Al Applications. Server-side device is configured to generate audio content track and ambient audio track, and one or more mixed audio content. FIGURE 8 applies to a method using Multimodal Al, as viewed from server-side.

[0054] FIGURE 9 is a flowchart for client-side device to perform tracking, synchronization, and spatialization of media content of client-side playback with server-side playback using a Multimodal Al System and Object Recognition Al Applications. This method enables client-side device to synchronize its media playback with server-side device in both time and 3D space. It calculates network latency and the physical positions of both device, then uses this data to apply directional audio filters. The result is a timed, spatialized audio experience that sounds as if it is coming from the server's precise physical location after compensating for acoustic and processing delays. FIGURE 9 applies to a method using Multimodal Al, as viewed from client-side.

[0055] FIGURE 10 is a flowchart for client-side device to perform tracking, synchronization, and spatialization of media content of client-side playback with server-side playback using a Multimodal Al System and Object Recognition Al Applications. This method enables server-side device to synchronize its media playback with client-side device in both time and 3D space. It calculates network latency and the physical positions of both device, then uses this data to apply directional audio filters. The result is a near-perfectly timed, spatialized audio experience that sounds as if it is coming from the server's precise physical location after compensating for acoustic and processing delays. FIGURE 10 applies to a method using Multimodal Al, as viewed from server-side.

[0056] FIGURE 11 is a flowchart for a method for applying a spatialization filter to a plurality of audio system formats includes. FIGURE 11 applies to a method using Multimodal Al, as viewed from client-side.

[0057] FIGURE 12 is a flowchart for a method for processing streamed media content on client-side device. FIGURE 12 applies to a method using Multimodal Al, as viewed from client-side.

[0058] FIGURE 13 is a flowchart for a method for processing streamed media content on server-side device. FIGURE 13 applies to a method using Multimodal Al, as viewed from server-side.

[0059] FIGURE 14 is a flowchart for a method for processing and demultiplexing a data stream. FIGURE 14 applies to a method for rendering a dynamic, context-aware binaural audio using client-side Al Audio Mixing System from a client device.

[0060] FIGURE 15 is a flowchart for a method for processing and performing semantic event detection to generate a corresponding stream of event markers describing a participant's actions. FIGURE 15 applies to a method for generating a high-information-density data stream for remote auditory rendering using serverside Multimodal Deconstruction Al System from a server device.

[0061] FIGURE 16 is a flowchart for a method for rendering a synchronized, multi-sensory haptic and auditory experience. FIGURE 16 applies to a method for rendering a synchronized, multi-sensory haptic and auditory experience using client-side Al Sensory Experience Application, from a client device.

[0062] FIGURE 17 is a flowchart for a method for generating a high- information-density data stream for enabling a remote multi-sensory experience. FIGURE 17 applies to a method for generating a high-information-density data stream for enabling a remote multi-sensory experience using server-side Multimodal Al System Application from a server device.

[0063] FIGURE 18 is a flowchart for a method for determining the real-time location coordinates of a client device using an acoustic positioning system. FIGURE 18 applies to a method for enabling location-aware applications, such as spatialized audio, by using mathematical cross-correlation to process signals from a plurality of fixed acoustic beacons and calculate the device's position via Time Difference of Arrival (TDoA).

[0064] FIGURE 19 is a flowchart for a method for determining the real-time location coordinates of a client device using an Al-enhanced acoustic positioning system. FIGURE 19 applies to a method for enabling location-aware applications, such as spatialized audio, by using client-side deep learning model to process signals from a plurality of fixed acoustic beacons and calculate the device's position via Time Difference of Arrival (TDoA).

[0065] FIGURE 20 is a flowchart for a method for rendering a real-time, spatialized, and voice-preserving translated audio experience. FIGURE 20 applies to a method for using client-side Multimodal Al Media Application to receive an audio stream in a source language, perform a live translation into a target language while preserving the vocal identity of the original speaker, and render the resulting translated audio as a fully synchronized and spatialized experience on a client device.

[0066] FIGURE 21 is a flowchart for a method for enabling a remote, real-time, spatialized audio translation. FIGURE 21 applies to server-side method for generating and transmitting a data stream to a client device, wherein said datastream comprises a primary audio track in a source language, corresponding realtime spatialization metadata, and the necessary timing data to allow the remote client device to perform a synchronized, spatialized, and translated rendering.DETAILED DESCRIPTION

[0067] While the present invention is susceptible of embodiment in many different forms, there is shown in the drawings and will herein be described in detail one or more specific embodiments, with the understanding that the present disclosure is to be considered as exemplary of the principles of the invention and not intended to limit the invention to the specific embodiments shown and described. In the following description and in the several figures of the drawings, like reference numerals are used to describe the same, similar or corresponding parts in the several views of the drawings.

[0068] The system and method for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and Object Recognition Al Application includes a plurality of components such as one or more of electronic components, hardware components, and computer software components. A number of such components can be combined or divided in the system. An example component of the system includes a set and / or series of computer instructions written in or implemented with any of a number of programming languages, as will be appreciated by those skilled in the art.

[0069] The system and method for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and Object Recognition Al Application includes a plurality of components such as one or more input and output (I / O) device, analogue to digital and digital to analogue device (AD / DA), amplifier device, active and passive speaker systems, and an other electronic device.

[0070] The system in one example employs one or more computer-readable signal-bearing media. The computer-readable signal bearing media store software, firmware and / or assembly language for performing one or more portions of one or more implementations of the invention. The computer-readable signalbearing medium for the system in one example comprises one or more of a magnetic, electrical, optical, biological, and atomic data storage medium. For example, the computer-readable signal-bearing medium comprises floppy disks, magnetic tapes, CD-ROMs, DVD-ROMs, hard disk drives, downloadable files, files executable “in the cloud,” and electronic memory.

[0071] FIGURE 1 is a schematic block diagram of a networked environment for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and Object Recognition Al Application via one or more of one or more electronic device, that comprises client-side networked environment 105, server-side networked environment 110, and network 115. Network 115 comprises one or more of one or more speaker system 135, microphone system 137, wireless communication sensor 165, input and output (I / O) device 198, Internet, private virtual network, extranet, fiber optic network,wide area network (WAN), local area network (LAN), wired network, wireless network, a satellite internet constellation system and an other type of network.

[0072] Client-side networked environment 105 comprises client-side device 120 and client-side playback device 125 that is operably connected with client-side device 120. Client-side device 120 comprises, for example, one or more of one or more tablet 120, phone 120, smart device 120, virtual reality headset 120, wireless virtual reality headset 120, augmented reality headset 120, wireless augmented reality headset 120, wire computing glasses 120, wireless computing glasses, computer program 120, computer browser 120, media player 120, game console 120, virtual device 120, neuro link device 120, wearable device 120 and an other computing device 120.

[0073] Client-side device 120 runs one or more applications. Client-side device 120 deploys over network 115.

[0074] Client-side playback device 125 is configured to play media content. For example, client-side playback device 125 plays media content received from client-side device 120. Alternatively, or additionally, client-side playback device 125 plays media content received directly over network 115. For example, clientside playback device 125 comprises one or more of one or more headphone 125, earphone 125, earbud 125, earworn wearable 125, screen 125, television 125, monitor 125, in-venue projector 125, home theater 125, three-dimensional digital projector 125, and an other client-side playback device 125. For example, clientside playback device 125 comprises one or more of one or more open headphone 125, semi-open headphone 125, closed headphone 125, and an other type of headphone 125.

[0075] Client-side playback device 125 operates in an environment with clientside sensor system 130. For example, sensor system 130 comprises one or more of one or more visual sensor 130, microphone 130, thermal sensor 130, vibratory sensor 130, position sensor 130, camera 130, LIDAR 130, radar 130, biomarker sensor 130, ambient light sensor (ALS), electrocardiogram sensor (ECG), and an other sensor 130. For example, the sensor system 130 comprises one or more of one or more head motion tracking sensor 130, eye motion tracking sensor 130, infrared camera to recognize hand gestures, understand depth or map a room 130, haptic sensor 130, magnetic sensor 130 and an other position sensor 130. For example, the sensor system 130 may comprise a spatial resolution of three degrees of freedom in directions for head tracking and six degrees or more for motion tracking so as to provide one or more smooth audio transitions and stable sound imaging.

[0076] Client-side device 120 is configured to generate a dynamically updated spatial audio experience based on the relative positions of the client and a server. A key aspect of this process is the calculation of an operative direction vector, which dictates the perceived direction of the audio source. This operative vector is intelligently derived to accommodate scenarios both with and without the use of head tracking. The system first establishes a baseline direction vector representing the direct line-of-sight between the client and server device. This baseline vector may be used directly as the operative direction vector for spatialization. Alternatively, in embodiments where the user's head orientation ismonitored by a sensor, the system applies a rotation transformation to the baseline vector based on the real-time head orientation data. The resulting rotated vector then becomes the operative direction vector. This dual-path approach ensures a robust and seamless spatial audio experience whether the head tracking feature is active, disabled, or unavailable.

[0077] Client-side playback device 125 is configured to operate in an environment having a speaker system 135. Said speaker system 135 may be any of a plurality of types, including, but not limited to, a channel-based system, an object-based system, a scene-based system, or another type of three-dimensional (3D) audio system. For instance, a channel-based system may be a single-channel or multi-channel configuration comprising one or more speakers placed in designed positions within a physical space. An object-based system may utilize audio objects, such as sound sources, that contain metadata describing the intended position and other spatial properties of the object. A scene-based system may utilize a sound field representation, such as one or more spherical harmonic basis functions, that describes how sound pressure changes as a function of time and direction. The physical speaker system 135 may comprise one or more audio transducers, such as woofers, to reproduce a low-frequency range, and tweeters, to reproduce a high-frequency range.

[0078] Client-side playback device 125 may operates in an environment with microphone system 137. For example, microphone system 137 may comprise one or more of one or more microphone array, spot microphone, and an other microphone system.

[0079] Client-side device 120 comprises one or more client-side memory 140 and client-side data storage 145.

[0080] Client-side memory 140 is defined herein as including both volatile and nonvolatile memory and data storage components. For example, client-side memory 140 may comprise one or more sequential access. For example, clientside memory 140 comprises one or more client-side buffers. Volatile components are those that do not retain data values upon loss of power. Nonvolatile components are those that retain data upon loss of power. For example, clientside memory 140 may comprise one or more random access memory (RAM), read-only memory (ROM), hard disk drive, solid-state drive, USB flash drive, memory card, floppy disk, optical disc such as compact disc (CD) or digital versatile disc (DVD), magnetic tape, and other memory components. For example, RAM may comprise one or more static random access memory (SRAM), dynamic random access memory (DRAM), magnetic random access memory (MRAM), and other forms of RAM. For example, ROM may comprise one or more programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and other forms of ROM.

[0081] Client-side memory 140 comprises one or more client-side master application 150, client-side Multimodal Al Media Application 152, client-side Multimodal Al System Application 160, client-side Object Recognition Al Application 162, client-side Al Audio Mixing System 163 and client-side Al Sensory Experience Application 164.

[0082] Client-side memory 140 further comprises client-side device unique identifier. Client-side device unique identifier is a number unique to this particular device. In other words, each device in the world will have its own number that no other such device will have. A copy of client-side device’s unique identifier, known as client-side unique identifier, will be transmitted by client-side device 120 in a message or packet to server-side computing device 170. Then, a copy of clientside transmitted unique identifier, known as server-side unique identifier, will be transmitted back from server-side computing device 170 to client-side device 120. Server-side unique identifier received by client-side device 120 will then be compared with client-side device’s unique identifier to help determine the integrity of the messages and as a security check.

[0083] Optionally, client-side memory 140 further comprises an other clientside application (not pictured). The other client-side application comprises one or more additional client-side application, additional client-side service, additional client-side process, and additional client-side functionality. For example, an other client-side application runs background services. For example, an other client-side application runs boot processes. For example, an other client-side application runs other client-side applications.

[0084] Client-side data storage 145 comprises one or more single database, multiple database, cloud application platform, relational database, no-sequel database, flash memory, solid state memory, and an other client-side data storage device. Client-side data storage 145 may be located in a single installation that may be local to server-side computing device 170. Alternatively, client-side data storage 145 may be located in a single installation that may be local to client-side device 120. Alternatively, client-side data storage 145 may be distributed in a plurality of locations. Client-side data storage 145 may be distributed in a plurality of geographical locations. Client-side data storage 145 may be distributed in a plurality of geographical locations located in the same time zone. Client-side data storage 145 may be distributed in a plurality of geographical locations, wherein not the geographical locations are located in the same time zone.

[0085] Client-side data storage 145 comprises one or more of item prices, order information, media content, and other information. For example, the media content comprises one or more of a live broadcast, simultaneously recorded broadcast, audio track, a video track, an other media track, a motion picture, a commercial, a motion picture trailer, a demonstration (“demo”), a commentary, extra content, and an other form of additional content. The media content comprises one or more of media data, media content files, and other media content. The motion picture comprises one or more of a feature-length theatrical production, short-film production, an animated production, a broadcast television production, live events, a pay television production, a documentary, a commercial, a trailer, and an other motion picture. The media data comprises one or more audio track, multi-channel track, commentary, and other media data. The audio track comprises one or more English language audio track, audio track in a language other than English, and customized audio track. The commentary comprises one or more commentary by one or more directors of a motion picture, a commentary by one or more actors in a motion picture, a commentary by contributors to amotion picture other than the directors and actors, and commentary by persons other than contributors to a motion picture.

[0086] Client-side master application 150 is configured to store playable media content, such as segmented or non-segmented media tracks, in client-side data storage 145. Optionally, the application 150 also performs media processing on the playable content. This processing may involve passing the content through one or more digital signal processing (DSP) algorithms — such as bandpass transfer functions, headphone transfer functions (HpTFs), compensation filters, speed- adaptive equalization filters, or Al-based filters — to regularize audio signals or to adjust sound attributes including timbre, spectral cues, and localization. Specifically, an HpTF may be applied to smooth, accentuate, or cancel frequency response fluctuations for a particular pair of headphones. The application 150 may automatically upload a predetermined HpTF to a connected pair of headphones or, upon receiving headphone data such as a model, serial number, or other identifying information, may select and upload a customized HpTF from client-side data storage 145. Furthermore, the application 150 may determine a processing delay associated with these functions and adjust client-side playback timeline to compensate.

[0087] For example, client-side master application 150 parses the playable media content into a chronological sequence that substantially matches the sequence of the motion picture. For example, client-side master application 150 writes the playable media content to one or more of client-side data storage 145 and client-side memory 140. For example, client-side master application 150 writes the playable media content to a media content file located in one or more clientside data storage 145 and client-side memory 140.

[0088] For example, client-side Multimodal Al Media Application 152 may produce audio and media content.

[0089] For example, client-side Multimodal Al System Application 160 may comprise one or more client-side voice-activated Al applications, voice-to-text Al applications, audio-to-client-device applications, authentication Al applications, on-device Al foundation models, real-time Al application, and another application.

[0090] For example, client-side Object Recognition Al Application 162 may produce spatialization of audio and media content by accessing and localizing specific objects within a given environment. To achieve this, the application may employ a variety of computer vision instruments and models. For instance, in a single-camera configuration, client-side Object Recognition Al Application 162 may utilize a deep neural network (DNN) that processes the camera feed along with other inputs, such as location Al sensor data, to predict ground truth and estimate distances to objects even in complex settings. Alternatively, a two-camera object recognition setup may be used to mimic human binocular vision, whereby the application calculates depth information by measuring the disparity between corresponding points in the images captured by each camera.

[0091] Furthermore, client-side Object Recognition Al Application 162 may leverage instance segmentation to measure the coordinates and pixel dimensions of specific objects, going beyond simple bounding boxes to understand their exact shape. The application can generate one or more estimates of distance usingthese object boundary boxes and other advanced localizing techniques. The comprehensive functions of client-side Object Recognition Al Application 162 extend to calibrating system parameters, determining detailed vision disparities between different viewpoints, and establishing both known and unknown baseline distances for measurement. It can also determine the focal lengths of the camera(s), analyze ambient lighting conditions, and even infer an object's potential weight and size based on its classification and dimensions.

[0092] The application is also capable of dynamic analysis, such as estimating the velocity of an object. This is accomplished by detecting an object, maintaining its identity across multiple frames in a video stream, and calculating the change in its position over time, taking into account the video's frame rate and the scale of the scene. To ensure the integrity of its analysis, client-side Object Recognition Al Application 162 can actively mitigate common computer vision challenges like occlusion (where one object blocks another) and perspective distortion. It may employ techniques such as perspective transformation to correct for these issues and maintain accurate recognition.

[0093] Client-side Object Recognition Al Application 162 may operate within a broader multimodal Al media system, where it can also access and use audio information, such as Sound Pressure Level (SPL), to aid in the localization of media content. The physical deployment of the system is flexible; one or more applications may be housed in a camera apparatus worn on a user’s head, located on the user's person, or run on client-side device or server-side device. To ensure consistent and accurate results across these different deployment scenarios, specialized measurement compensation techniques are utilized. The overall accuracy and resolution of the system's direction-finding in 3D space can be further improved by increasing the resolution and sampling rate of the media recording itself.

[0094] Client-side master application 150 is configured to connect with serverside networked environment 110 so as to substantially synchronize between server-side networked environment 110 and client-side device 120 media content played on client-side playback device 125. Client-side playback device 125 comprises one or more of one or more screen 125, television 125, monitor 125, cellular phone 125, laptop computer 125, desktop computer 125, notebook 125, tablet 125, headset 125, channel-based playback system 125, object-based playback system 125, scene-based playback system 125, 3d audio playback system 125, and an other client-side playback device 125. Client-side playback device 125 plays for the user one or more audio media content, video media content, and an other form of media content. For example, networked environment 115 may be synchronized with other sensory experiences such as, for example, one or more smoke effects, fire effects, lasers, fireworks, drones and water droplets, moving chairs, and the like. For example, more than one client-side playback device 125 may be used simultaneously.

[0095] As explained below in greater detail, particularly in Figure 9, client-side master application 150 is configured to perform one or more of sampling and recording client-side running media play time (CRMPT) at which client-side media player plays the media on the client. The CRMPT is defined as the elapsed runningtime for customized media content that is being played by client-side media player on the client. If no customized media content is being played by client-side media player, the CRMPT is defined as zero. The CRMPT recorded by client-side master application 150 represents a real world time value based on the host system clock of the client. Then, client-side master application 150 creates client-side message or packet that it transmitres to server-side networked environment 110.

[0096] As explained below in greater detail, particularly in Figure 4, client-side Multimodal Al System Application 160 is configured to determine by client-side device, its own position (xc, yc, zc) using client-side Object Recognition Al Application 162.

[0097] Server-side networked environment 110 comprises server-side data storage 155, server-side computing device 170 that is operably connected with server-side data storage 155, and server-side playback device 196 that is operably connected with server-side data storage 155. Server-side data storage 155 is a second location where, as mentioned above in relation to client-side data storage 145, client-side master application 150 may store the playable media content.

[0098] Server-side data storage 155 comprises one or more of item prices, order information, media content, and other information. For example, the media content files comprise one or more of one or more audio track, video track, an other media track, motion picture, commercial, motion picture trailer, demonstration (“demo”), commentary, extra content, and an other form of additional content. The media content comprises one or more media data, media content files, and other media content. The motion picture comprises one or more of one or more feature-length theatrical production, short-film production, animated production, broadcast television production, live events, pay television production, documentary, commercial, trailer, and an other motion picture. The media data comprises one or more of one or more audio track, multi-channel track, commentary, and other media data. The audio track comprises one or more of one or more English language audio track, audio track in a language other than English, and customized audio track. The commentary comprises one or more commentary by one or more director of a motion picture, commentary by one or more actor in a motion picture, commentary by contributors to a motion picture other than the directors and actors, and commentary by persons other than contributors to a motion picture.

[0099] Server-side data storage 155 comprises one or more single database, multiple database, cloud application platform, relational database, no-sequel database, flash memory, solid state memory, and an other server-side data storage device. Server-side data storage 155 may be located in a single installation that may be local to client-side device 120. Alternatively, server-side data storage 155 may be located in a single installation that may be local to server-side computing device 170. Alternatively, server-side data storage 155 may be distributed in a plurality of locations. Server-side data storage 155 may be distributed in a plurality of geographical locations. Server-side data storage 155 may be distributed in a plurality of geographical locations located in the same time zone. Server-side data storage 155 may be distributed in a plurality of geographicallocations, wherein not the geographical locations are located in the same time zone.

[0100] Server-side computing device 170 comprises one or more server, computer, cloud-computing device, and distributed computing system.

[0101] Server-side computing device 170 may be located in a single installation. Alternatively, server-side computing device 170 may be distributed in a plurality of geographical locations. For example, server-side computing device 170 may be distributed in a plurality of geographical locations located in the same time zone. For example, server-side computing device 170 may be distributed in a plurality of geographical locations wherein not the geographical locations are located in the same time zone.

[0102] Server-side playback device 196 is configured to play media content. For example, server-side playback device 196 plays media content received from server-side computing device 170. Alternatively, or additionally, server-side playback device 196 plays media content received directly over network 115. For example, server-side playback device 196 comprises one or more of one or more in-venue projector 196, home theater 196, television 196, monitor 196, three- dimensional digital projector 196, and an other device 196.

[0103] Server-side playback device 196 is configured to operate in an environment having a speaker system 135. Said speaker system 135 may be any of a plurality of types, including, but not limited to, a channel-based system, an object-based system, a scene-based system, or another three-dimensional (3D) audio system. For instance, a channel-based system may be a single-channel or multi-channel configuration comprising one or more speakers placed in designed positions within a physical space. An object-based system may utilize audio objects, such as sound sources, that contain metadata describing the intended position and other spatial properties of the object. A scene-based system may utilize a sound field representation, such as one or more spherical harmonic basis functions, that describes how sound pressure changes as a function of time and direction. The physical speaker system 135 may comprise one or more audio transducers, such as woofers, to reproduce a low-frequency range, and tweeters, to reproduce a high-frequency range.

[0104] Server-side playback device 196 operates in an environment with server-side sensor system 197. For example, sensor system 197 comprises one or more of one or more visual sensor 197, microphone 197, thermal sensor 197, vibratory sensor 197, position sensor 197, camera 197, LIDAR 197, radar 197, biomarker sensor 197, ambient light sensor (ALS), electrocardiogram sensor (ECG), and an other sensor 197. For example, the sensor system 197 comprises one or more of one or more head motion tracking sensor 197, eye motion tracking sensor 197, infrared camera to recognize hand gestures, understand depth or map a room 197, haptic sensor 197, magnetic sensor 197 and an other position sensor 197. For example, the sensor system 197 may comprise a spatial resolution of three degrees of freedom in directions for head tracking and six degrees or more for motion tracking so as to provide one or more smooth audio transitions and stable sound imaging.

[0105] Server-side playback device 196 is configured to communicate with server-side computing device 170. For example, server-side playback device 196 communicates with server-side computing device 170 using one or more of one or more satellite, antenna, cable, network 115, and an other communication method. Server-side playback device 196 comprises one or more of one or more digital projector 196, hologram projector 196, led screen 196, screen 196, television 196, monitor 196, cellular phone 196, laptop computer 196, notebook 196, tablet 196, headset 196, multi-channel playback system 196, and an other server-side playback device 196.

[0106] Server-side computing device 170 comprises server-side memory 175. Server-side memory 175 is defined herein as including both volatile and nonvolatile memory and data storage components. For example, server-side memory 175 may comprise one or more sequential access. For example, server-side memory 175 comprises one or more server-side buffers. Volatile components are those that do not retain data values upon loss of power. Nonvolatile components are those that retain data upon loss of power. For example, server-side memory 175 may comprise one or more random access memory (RAM), read-only memory (ROM), hard disk drive, solid-state drive, USB flash drive, memory card, floppy disk, optical disc such as compact disc (CD) or digital versatile disc (DVD), magnetic tape, and other memory components. For example, RAM may comprise one or more static random access memory (SRAM), dynamic random access memory (DRAM), magnetic random access memory (MRAM), and other forms of RAM. For example, ROM may comprise one or more programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and other forms of ROM.

[0107] Server-side computing device 170 comprises one or more of serverside master application 180, server-side streaming application 185, server-side Multimodal Al media application 190, server-side Multimodal Al System Application 191 , server-side Object Recognition Al System 192 and server-side Multimodal Deconstruction Al System 193. Server-side master application 185 is configured to provide synchronization timing information to one or more of clientside master application 150, and server-side streaming application 185.

[0108] Optionally, server-side computing device 170 further comprises an other server-side application (not pictured). The other server-side application comprises one or more of additional server-side application, additional server-side service, additional server-side process, and additional server-side functionality.

[0109] For example, the other server-side application runs background services. For example, the other server-side application runs boot processes. For example, the other server-side application runs other server-side applications.

[0110] As explained below in greater detail, particularly in Figure 10, serverside master application 180 is configured to perform one or more of sampling and recording server-side running media play time (SRMPT) at which server-side media player plays the media on the server. The SRMPT is defined as an elapsed running time for customized media content that is being played by server-side media player on the server. If no customized media content is being played by server-side media player, the SRMPT is defined as zero. For example, the motionpicture’s SRMPT time might clock in at 6 minutes, 10 seconds, and 10 frames. The SRMPT recorded by server-side master application 180 represents a real world time value based on the host system clock of the server. Then, server-side master application 180 creates server-side message or packet that it transmits to clientside master application 150.

[0111] As explained below in greater detail, particularly in Figure 5, server-side master application 180 is configured to store media content. For example, serverside master application 180 stores playable media content in server-side data storage 155. The playable media content comprises one or more segmented media content track, non-segmented media content track, audio stimulus, and an other playable media content. Optionally, server-side master application 180 performs media processing of the playable media content.

[0112] Server-side master application 180 is configured to store playable media content to be played by server-side playback device 196. For example, server-side master application 180 stores the playable media content in serverside data storage 155. Optionally, server-side master application 180 performs media processing of the playable media content. For example, server-side master application 180 passes the playable media content through one or more bandpass transfer function, compensation filters, speed-adaptive equalization filters, Al- based filters and an other DSP algorithm, to regularize audio signals, and adjust one or more timber, spectral cue, localization, and an other attribute of sound. For example, server-side master application 180 parses the playable media content into a chronological sequence that substantially matches the sequence of the motion picture. For example, server-side master application 180 writes the playable media content to one or more server-side data storage 155 and serverside memory 175. For example, server-side master application 180 writes the playable media content to a media content file located in one or more server-side data storage 155 and server-side memory 175.

[0113] For example, server-side Multimodal Al Media Application 190 may produce audio and media content.

[0114] For example, server-side Multimodal Al System Application 191 may comprise of one or more of server-side text-to-speech Al application, text-to-audio Al application, audio-to-dialog Al application, location Al application, dialog supervisor Al application, privacy guardrail Al application, marketing Al application, server-based Al model, real-time Al application and an other application.

[0115] Server-side streaming application 185 segments media content for deployment via network 115 to client-side device 120. Server-side streaming application 185 supports multiple alternate data streams, two or more of which can have different bit rates from each other. For example, the multiple alternate data streams might have same bit rates from each other. Server-side streaming application 185 also allows for client-side device 120 to switch streams intelligently as network bandwidth changes. Server-side streaming application 185 also provides for media encryption and user authentication over encrypted connections.

[0116] Speaker system 135 may receive audio via one or more server-side computing device 170, network 115, an analogue connection, and a digitalconnection. For example, speaker systems 135 may be one or more actively and passively used with amplifiers.

[0117] Microphone system 137, may receive audio via speaker system 137. Microphone system 137 may be placed on the wall, on the speaker, or within the venue environment. Microphone system 137 may used to calibrate speaker system 135 and verify the speaker conditions.

[0118] Input and output (I / O) device 198 may receive audio via one or more directly from server-side computing device 170, over the network 115, an analogue connection, and a digital connection. For example, input and output (I / O) device 198 may perform one or more digital to analogue and analogue to digital conversions. For example, input and output (I / O) device 198 may route audio to one or more of one or more amplifier, speaker and an other electronic device.

[0119] As explained below in greater detail, particularly in Figure 5, server-side Multimodal Al System Application 191 is configured to determine by client-side device, its own position (xs, ys, zs) via multilateration using three or more serverside Object Recognition Al Application 192.

[0120] FIGURE 2 is a schematic block diagram of server-side computing device 170 in an alternative configuration of a networked environment for real-time synchronization of media content via one or more of one or more electronic device and speaker system.

[0121] Server-side computing device 170 comprises one or more server-side data storage 155, server-side playback device 196 (not pictured), server-side memory 175, server-side processor 210, and server-side local interface 220. Server-side local interface 220 is operationally connected with one or more serverside data storage 155, server-side memory 180, and server-side processor 210. Server-side memory comprises one or more server-side master application 180, server-side streaming application 185, server-side Multimodal Al Media Application 190, and server-side Multimodal Al System Application 191. For example, server-side processor 210 comprises server-side computer. For example, server-side local interface 220 comprises a bus. For example, serverside local interface 220 comprises a bus and further comprises one or more of an accompanying address / control bus or other bus structure.

[0122] Software components stored in one or more server-side memory 175 and server-side data storage 155 are executable by server-side processor 210. In this respect, the term executable means a program file that is in a form that can ultimately be run by server-side processor 210. For example, a compiled program is executable if it may be translated into machine code in a format that can be loaded into a random access portion of server-side memory 175 and run by serverside processor 210. For example, source code is executable if it may be expressed in a proper format, such as object code, that may be loaded into a random access portion of server-side memory 175 and run by server-side processor 210. For example, source code is executable if it may be interpreted by another executable program to generate instructions in a random access portion of server-side memory 175 and run by server-side processor 210. An executable program may be stored in one or more portions or components of server-side memory 175. For example, server-side memory 175 comprises one or more random access memory(RAM), read-only memory (ROM), hard disk drive, solid-state drive, USB flash drive, memory card, floppy disk, optical disc such as compact disc (CD) or digital versatile disc (DVD), magnetic tape, and other memory components.

[0123] One or more data and components stored in one or more server-side memory 175 and server-side data storage 155 are executable by server-side processor 210. For example, server-side processor 210 can execute one or more server-side master application 180, server-side streaming application 185, serverside Multimodal Al Media Application 190, and server-side Multimodal Al System Application 191.

[0124] For example, as an alternative to the setup in FIGURE 1 with serverside data storage 155 separate from server-side computing device 170, serverside data storage 155 may be located in server-side computing device 170. For example, server-side data storage 155 may be located in server-side memory 175.

[0125] Server-side processor 210 comprises one or more processors. Serverside memory 175 comprises one or more memories. For example, server-side memory 175 comprises at least one memory configured to operate in a parallel processing circuit. In such a case, server-side local interface 220 may serve as network 115. For example, server-side local interface 220 may facilitate communication between two processors. For example, server-side local interface 220 may facilitate communication between a processor and a memory. For example, server-side local interface 220 may facilitate communication between two memories. Server-side local interface 220 may comprise additional systems designed to coordinate this communication. For example, server-side local interface 220 may comprise a system to perform load balancing. Server-side processor 210 may comprise an electrical processor. Alternatively, or additionally, server-side processor 210 may comprise a non-electrical processor.

[0126] Any logic or application described herein, including but not limited to server-side master application 180, server-side streaming application 185, serverside Multimodal Al Media Application 190, and server-side Multimodal Al System Application 191 that comprises software or code can be embodied in any non- transitory computer-readable medium for use by or in connection with an instruction execution system such as, for example, server-side processor 210 in a computer system or other system. In this sense, the logic may comprise, for example, statements including instructions and declarations that can be fetched from the computer-readable medium and can be executed by the instruction execution system. In the context of the present disclosure, a computer-readable medium can be any medium that can contain, store, or maintain the logic or application described herein for use by or in connection with the instruction execution system. For example, the computer-readable medium may comprise one or more RAM, ROM, hard disk drive, solid-state drive, USB flash drive, memory card, floppy disk, optical disc such as a CD or a DVD, magnetic tape, and other memory components. For example, RAM may comprise one or more SRAM, DRAM, MRAM, and other forms of RAM. For example, ROM may comprise one or more PROM, EPROM, EEPROM, and other forms of ROM.

[0127] FIGURE 3 illustrates a system wherein a user 305 wears client-side device 310 that synchronizes with server-side device 315. In some embodiments, client-side device 310 may be, for example, a smart wireless glass comprising components such as a binocular camera 312, a speaker 313, and one or more sensors 314. Said sensors 314 may include, but are not limited to, a microphone, an accelerometer, a gyroscope, a magnetometer, a head tracking sensor, or a thermal sensor. In the embodiment shown, server-side device 315 is a humanoid robot, which may feature a digital screen 320 and a head comprised of mechanized or biological components, such as moving eyes, eyelids, a nose, and lips. However, in other embodiments, server-side device 315 may take other forms, for example, a vehicle, an autonomous vehicle, a wearable, another mechanical or biological device, or an intelligent environment. Said intelligent environment may act as a digital counterpart to a physical location, such as a hospital, museum, or stadium, and may be configured to process information and communicate intelligently with client-side device 310. In any embodiment, client-side device 310 may synchronize with any combination of the digital, mechanical, or biological components of server-side device 315. Furthermore, the system accommodates multiple states of relative motion, wherein either client-side device 310 or serverside device 315 may be in motion while the other is static, or both may be in motion concurrently in a two-dimensional (2D) or three-dimensional (3D) plane.

[0128] FIGURE 4 is a flowchart of method 400 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and Object Recognition Al Application. FIGURE 4 applies to a method using Multimodal Al, as viewed from client-side.

[0129] The order of the steps in the method 400 is not constrained to that shown in Figure 4 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0130] According to this method, client-side device performs tracking, synchronization, and spatialization of media content of client-side playback with server-side.

[0131] In block 405, receiving one or more user commands by client-side device, obtaining user attributes via an authentication Al application, and processing a verbal command from the user to generate a user text prompt. Block 405 transfers control to block 410.

[0132] Next, in block 410, sending, by client-side device, a request packet to server-side device, said packet comprising client-side unique identifier, the user text prompt, and environmental data such as application sensor data and location Al application data. Block 410 transfers control to block 415.

[0133] Next, in block 415, receiving, by client-side device, a response packet from server-side device, said packet comprising one or more audio content tracks, an ambient audio track, server-side unique identifier, and one or more audible or inaudible reference sound signals. Block 415 transfers control to block 420.

[0134] Next, in block 420, creating, by client-side device, client-side packet comprising client-side unique identifier and client-side start host time (CSHT). Block 420 transfers control to block 425.

[0135] Next, in block 425, sending, by client-side device, said client-side packet to server-side device to initiate a time synchronization handshake. Block 425 transfers control to block 430

[0136] Next, in block 430, receiving and processing, by client-side device, server-side packet comprising server-side unique identifier, the original CSHT, server-side start host time (SSHT), server-side end host time (SEHT), and serverside running media play time (SRMPT). Block 430 transfers control to block 435.

[0137] Next, in block 435, synchronizing in real-time, by client-side device, a local playback timeline with server-side playback timeline using said received timestamps. Block 435 transfers control to block 440.

[0138] Next, in block 440, calculating, by client-side Object Recognition Al Application, a final, authoritative set of real-time location coordinates. This calculation comprises concurrently determining client-side device's own position (Xc, Yc, Zc) and determining server-side device's position (Xs, Ys, Zs) by adaptively fusing data from a plurality of sensor streams, said streams comprising visual data from onboard cameras and acoustic data derived from analyzing the reference sound signals received from the server. Said fusion is weighted based on real-time confidence scores to maintain continuous tracking during periods of visual occlusion. Block 440 transfers control to block 445.

[0139] Next, in block 445, calculating, by client-side device, a current azimuth angle 6 using 6 = arctan2((Ys - Yc) I (Xs - Xc)) and a current elevation angle cp using cp = arctan((Zs - Zc) / sqrt((Xs - Xc)2+ (Ys - Yc)2)). Block 445 transfers control to block 450.

[0140] Next, in block 450, comparing the current server-side device location coordinates (Xs, Ys, Zs) to a set of previous coordinates stored in a memory buffer to calculate a displacement magnitude. If said displacement magnitude exceeds a pre-determined threshold, generating a new filter tracking response based on the current orientation angles (0, cp). Block 450 transfers control to block 455.

[0141] Next, in block 455, applying, by client-side device, a smooth transition algorithm, such as a cross-fade, between a previous filter tracking response retrieved from the memory buffer and the new filter tracking response to prevent abrupt audible changes in the spatialized audio output. Block 455 transfers control to block 460.

[0142] Next, in block 460, calculating, by client-side device, a first delay corresponding to the acoustic propagation time from the current server-side device location (Xs, Ys, Zs) to client-side device location (Xc, Yc, Zc). Block 460 transfers control to block 465.

[0143] Next, in block 465, generating, by client-side device, one or more mixed audio content using the audio content track and the ambient audio track. Block 465 transfers control to block 470.

[0144] Next, in block 470, applying one or more audio filters to the mixed audio content by selecting, from a database, a pair of Head-Related Impulse Responses (HRIR) corresponding to the calculated azimuth and elevation angles (0, cp), said pair comprising a left-ear HRIR and a right-ear HRIR. Block 470 transfers control to block 475.

[0145] Next, in block 475, generating a left channel of a binaural stereo output by convolving the mixed audio content with the left-ear HRIR, and generating a right channel of the binaural stereo output by convolving the mixed audio content with the right-ear HRIR. Block 475 transfers control to block 480.

[0146] Next, in block 480, applying the calculated first delay and one or more loudness compensation signals to the resulting binaural stereo output to render distance. Further processing the binaural stereo output using one or more of a Binaural Room Impulse Response (BRIR), a Headphone Transfer Function (HpTF), and a voice isolator, wherein said BRIR may be a pre-measured response or may be generated by rendering a more general spatial representation, such as a Spatial Room Impulse Response (SRIR), for the current user orientation. Block 480 transfers control to block 485.

[0147] Next, in block 485, calculating, by a processor on client-side device, a second delay corresponding to the total signal processing latency introduced by the application of said spatialization filters. Block 485 transfers control to block 490.

[0148] Next, in block 490, applying, by client-side device, a final compensating delay to client-side playback, said final compensating delay accounting for the first delay (acoustic propagation) and the second delay (processing latency) to ensure the processed audio is aligned with the synchronized playback timeline. Block 490 transfers control to block 495.

[0149] Next, in block 495, playing back, by client-side device, the fully synchronized and spatialized mixed audio content and media content. The process may then return to block 440 to continuously update the location of server-side device and the corresponding spatialized audio parameters. Block 495 then terminates the process.

[0150] It is to be understood that while the embodiment described in blocks 405 through 495 specifies the selection and application of Head-Related Impulse Responses (HRIR) and Binaural Room Impulse Responses (BRIR), the scope of the invention is not limited to these specific filter formats. The spatialization filters and responses applied by client-side device may be derived, synthesized, or rendered from any suitable spatial audio representation. Such representations include, but are not limited to, higher-order Ambisonics, spherical harmonic coefficients, or Spatial Room Impulse Responses (SRIRs). The process of calculating orientation angles (0, cp) and using them to generate a specific binaural output from these broader spatial representations is contemplated as being within the scope of the present invention.

[0151] FIGURE 5 is a flowchart of method 500 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and Object Recognition Al Application. FIGURE 5 applies to a method using Multimodal Al, as viewed from server-side.

[0152] The order of the steps in the method 500 is not constrained to that shown in Figure 5 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0153] According to this method, client-side device performs tracking, synchronization, and spatialization of media content of client-side playback withserver-side playback using a Multimodal Al System and Object Recognition Al Application.

[0154] In block 505, receiving, by server-side device, a request packet from client-side device, said packet comprising client-side unique identifier, a user text prompt, and environmental data such as application sensor data and location Al application data. Block 505 transfers control to block 510.

[0155] Next, in block 510, processing, by server-side device, the user text prompt to generate one or more audio content tracks, and concurrently using the received environmental data to generate a context-aware ambient audio track. Block 510 transfers control to block 515.

[0156] Next, in block 515, generating, by server-side device, one or more audible or inaudible reference sound signals, said signals being configured for detection and analysis by client-side device's Object Recognition Al Application. Block 515 transfers control to block 520.

[0157] Next, in block 520, transmitting, by server-side device, a first response packet to client-side device, said packet comprising the generated audio content track, the ambient audio track, the reference sound signals, and server-side unique identifier. Block 520 transfers control to block 525.

[0158] Next, in block 525, subsequently receiving, by server-side device from client-side device, a time synchronization packet, said packet comprising clientside start host time (CSHT). Block 525 transfers control to block 530.

[0159] Next, in block 530, in response to the synchronization request, creating and transmitting, by server-side device, a second response packet to client-side device, said packet containing a set of timing data, such as the original CSHT, SSHT, SEHT, and SRMPT, required for the client to perform playback synchronization. Block 530 transfers control to block 535.

[0160] Next, in block 535, initiating, by server-side device, local playback of the audio content and continuous emission of the reference sound signals, said playback being internally synchronized with the timing information dispatched to client-side device to maintain a coherent shared experience. The process may then await a subsequent request.

[0161] FIGURE 6 is a flowchart of method 600 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and Object Recognition Al Application. FIGURE 6 applies to a method using a single, natively Multimodal Al, as viewed from client-side.

[0162] The order of the steps in the method 600 is not constrained to that shown in Figure 6 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0163] According to this method, client-side device performs tracking, synchronization, and spatialization of media content of client-side playback with server-side playback using a single, natively Multimodal Al System and Object Recognition Al Applications. Client-side device is configured to communicate with server-side device via a real-time audio and data stream.

[0164] In block 605, capturing, by client-side device controlled by a user, a verbal command from the user as a continuous audio stream, and concurrently obtaining application sensor data from one or more onboard sensors, such as acamera, a microphone array, or an inertial measurement unit. Block 605 transfers control to block 610.

[0165] Next, in block 610, transmitting, by client-side device, a request packet comprising the continuous audio stream to server-side device for real-time processing, said packet comprising client-side unique identifier, the continuous audio stream capturing a user's verbal command, and environmental data such as application sensor data and location Al application data. Block 610 transfers control to block 615.

[0166] Next, in block 615, receiving, by client-side device, a response packet from server-side device, said packet comprising one or more audio content tracks, an ambient audio track, server-side unique identifier, and one or more audible or inaudible reference sound signals. Block 615 transfers control to block 620.

[0167] Next, in block 620, creating, by client-side device, client-side packet comprising client-side unique identifier and client-side start host time (CSHT). Block 620 transfers control to block 625.

[0168] Next, in block 625, sending, by client-side device, said client-side packet to server-side device to initiate a time synchronization handshake. Block 625 transfers control to block 630

[0169] Next, in block 630, receiving and processing, by client-side device, server-side packet comprising server-side unique identifier, the original CSHT, server-side start host time (SSHT), server-side end host time (SEHT), and serverside running media play time (SRMPT). Block 630 transfers control to block 635.

[0170] Next, in block 635, synchronizing in real-time, by client-side device, a local playback timeline with server-side playback timeline using said received timestamps. Block 635 transfers control to block 640.

[0171] Next, in block 640, calculating, by client-side Object Recognition Al Application, a final, authoritative set of real-time location coordinates. This calculation comprises concurrently determining client-side device's own position (Xc, Yc, Zc) and determining server-side device's position (Xs, Ys, Zs) by adaptively fusing data from a plurality of sensor streams, said streams comprising visual data from onboard cameras and acoustic data derived from analyzing the reference sound signals received from the server. Said fusion is weighted based on real-time confidence scores to maintain continuous tracking during periods of visual occlusion. Block 640 transfers control to block 645.

[0172] Next, in block 645, calculating, by client-side device, a current azimuth angle 0 using 0 = arctan2((Ys - Yc) I (Xs - Xc)) and a current elevation angle c using cp = arctan((Zs - Zc) / sqrt((Xs - Xc)2+ (Ys - Yc)2)). Said calculation comprises: a) first, determining a baseline direction vector between server-side device's location coordinates (xs,ys,zs) and client-side device's own location coordinates (xc,yc,zc); b) second, applying a rotation transformation to the baseline direction vector based on real-time orientation data received from the user's head tracking sensor; c) finally, calculating the azimuth and elevation angles from the resulting rotated vector. Block 645 transfers control to block 650.

[0173] Next, in block 650, comparing the current server-side device location coordinates (Xs, Ys, Zs) to a set of previous coordinates stored in a memory buffer to calculate a displacement magnitude. If said displacement magnitude exceeds apre-determined threshold, generating a new filter tracking response based on the current orientation angles (0, cp). Block 650 transfers control to block 655.

[0174] Next, in block 655, applying, by client-side device, a smooth transition algorithm, such as a cross-fade, between a previous filter tracking response retrieved from the memory buffer and the new filter tracking response to prevent abrupt audible changes in the spatialized audio output. Block 655 transfers control to block 660.

[0175] Next, in block 660, calculating, by client-side device, a first delay corresponding to the acoustic propagation time from the current server-side device location (Xs, Ys, Zs) to client-side device location (Xc, Yc, Zc). Block 660 transfers control to block 665.

[0176] Next, in block 665, generating, by client-side device, one or more mixed audio content using the audio content track and the ambient audio track. Block 665 transfers control to block 670.

[0177] Next, in block 670, applying one or more audio filters to the mixed audio content by selecting, from a database, a pair of Head-Related Impulse Responses (HRIR) corresponding to the calculated azimuth and elevation angles (0, cp), said pair comprising a left-ear HRIR and a right-ear HRIR. Block 670 transfers control to block 675.

[0178] Next, in block 675, generating a left channel of a binaural stereo output by convolving the mixed audio content with the left-ear HRIR, and generating a right channel of the binaural stereo output by convolving the mixed audio content with the right-ear HRIR. Block 675 transfers control to block 680.

[0179] Next, in block 680, applying the calculated first delay and one or more loudness compensation signals to the resulting binaural stereo output to render distance. Further processing the binaural stereo output using one or more of a Binaural Room Impulse Response (BRIR), a Headphone Transfer Function (HpTF), speed-adaptive equalization filters, noise-cancellation, and a voice isolator, wherein said BRIR may be a pre-measured response or may be generated by rendering a more general spatial representation, such as a Spatial Room Impulse Response (SRIR), for the current user orientation. Block 680 transfers control to block 685.

[0180] Next, in block 685, calculating, by a processor on client-side device, a second delay corresponding to the total signal processing latency introduced by the application of said spatialization filters. Block 685 transfers control to block 690.

[0181] Next, in block 690, applying, by client-side device, a final compensating delay to client-side playback, said final compensating delay accounting for the first delay (acoustic propagation) and the second delay (processing latency) to ensure the processed audio is aligned with the synchronized playback timeline. Block 690 transfers control to block 695.

[0182] Next, in block 695, playing back, by client-side device, the fully synchronized and spatialized mixed audio content and media content. The process may then return to block 640 to continuously update the location of server-side device and the corresponding spatialized audio parameters. Block 695 then terminates the process.

[0183] It is to be understood that while the embodiment described in blocks 605 through 695 specifies the selection and application of Head-Related Impulse Responses (HRIR) and Binaural Room Impulse Responses (BRIR), the scope of the invention is not limited to these specific filter formats. The spatialization filters and responses applied by client-side device may be derived, synthesized, or rendered from any suitable spatial audio representation. Such representations include, but are not limited to, higher-order Ambisonics, spherical harmonic coefficients, or Spatial Room Impulse Responses (SRIRs). The process of calculating orientation angles (0, cp) and using them to generate a specific binaural output from these broader spatial representations is contemplated as being within the scope of the present invention.

[0184] FIGURE 7 is a flowchart of method 700 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and Object Recognition Al Application. FIGURE 7 applies to a method using a single, natively Multimodal Al, as viewed from server-side.

[0185] The order of the steps in the method 700 is not constrained to that shown in Figure 7 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0186] According to this method, client-side device performs tracking, synchronization, and spatialization of media content of client-side playback with server-side playback using a single, natively Multimodal Al System and Object Recognition Al Application. Server-side device is configured to communicate with client-side device via a real-time audio and data stream.

[0187] In block 705, receiving, by server-side device, a request packet from client-side device, said packet comprising client-side unique identifier, a continuous audio stream capturing a user's verbal command, and environmental data such as application sensor data and location Al application data. Block 705 transfers control to block 710.

[0188] Next, in block 710, processing, by server-side device, the continuous audio stream using a single, natively Multimodal Al model to directly interpret the verbal command and generate one or more audio content tracks, and concurrently using the received environmental data to generate a context-aware ambient audio track. Block 710 transfers control to block 715.

[0189] Next, in block 715, generating, by server-side device, one or more audible or inaudible reference sound signals, said signals being configured for detection and analysis by client-side device's Object Recognition Al Application. Block 715 transfers control to block 720.

[0190] Next, in block 720, transmitting, by server-side device, a first response packet to client-side device, said packet comprising the generated audio content track, the ambient audio track, the reference sound signals, and server-side unique identifier. Block 720 transfers control to block 725.

[0191] Next, in block 725, subsequently receiving, by server-side device from client-side device, a time synchronization packet, said packet comprising clientside start host time (CSHT). Block 725 transfers control to block 730.

[0192] Next, in block 730, in response to the synchronization request, creating and transmitting, by server-side device, a second response packet to client-sidedevice, said packet containing a set of timing data, such as the original CSHT, SSHT, SEHT, and SRMPT, required for the client to perform playback synchronization. Block 730 transfers control to block 735.

[0193] Next, in block 735, initiating, by server-side device, local playback of the audio content and continuous emission of the reference sound signals, said playback being internally synchronized with the timing information dispatched to client-side device to maintain a coherent shared experience. The process may then await a subsequent request.

[0194] FIGURE 8 is a flowchart of method 800 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and Object Recognition Al Application. FIGURE 8 applies to a method using a Multimodal Al, as viewed from server-side.

[0195] The order of the steps in the method 800 is not constrained to that shown in Figure 8 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0196] According to this method, client-side device performs tracking, synchronization, and spatialization of media content of client-side playback with server-side playback using a Multimodal Al System and Object Recognition Al Applications. Server-side device is configured to generate audio content track and ambient audio track, and one or more mixed audio content.

[0197] In block 805, processing, by server-side device, audio content track, ambient audio track, and data input; generating, by server-side device, using audio content track and ambient audio track, one or more mixed audio content; processing, by server-side device, mixed audio content and data input. Block 805 transfers control to block 810.

[0198] Next, in block 810, uploading, by server-side device, mixed audio content to server-side device or cloud database. Block 810 transfers control to block 815.

[0199] Next, in block 815, receiving, by client-side device, from server-side device or cloud database, mixed audio content; and processing, by client-side device, mixed audio content. Block 815 then terminates the process.

[0200] FIGURE 9 is a flowchart of method 900 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and Object Recognition Al Application. FIGURE 9 applies to a method using a Multimodal Al, as viewed from client-side.

[0201] The order of the steps in the method 900 is not constrained to that shown in Figure 9 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0202] According to this method, client-side device performs tracking, synchronization, and spatialization of media content of client-side playback with server-side playback using a Multimodal Al System and Object Recognition Al Applications. This method enables client-side device to synchronize its media playback with server-side device in both time and 3D space. It calculates network latency and the physical positions of both device, then uses this data to apply directional audio filters. The result is a near-perfectly timed, spatialized audio experience that sounds as if it is coming from the server's near-precise physicallocation after compensating for acoustic and processing delays.

[0203] In block 905, receiving, by client-side device, server-side message or packet from server-side master application, server-side message or packet comprising one or more of server-side unique identifier, client-side start host time (CSHT), server-side start host time (SSHT), server-side end host time (SEHT), and server-side running media play time (SRMPT). Block 905 transfers control to block 910.

[0204] Next, in block 910, reading, by client-side device, server-side message or packet into one or more client-side buffers. Block 910 transfers control to block 915.

[0205] Next, in block 915, recording, by client-side device, client-side end host time (CEHT) upon reception of the packet. Block 915 transfers control to block 920.

[0206] Next, in block 920, evaluating and verifying, by client-side device, the logical consistency of the timing data within server-side message or packet. Calculating a half-round-trip time (HRT), by client-side device, using the formula: HRT = ((CEHT - CSHT) - (SEHT - SSHT)) / 2; calculating a playback time offset (TPO), representing the clock offset between the client and server, by clientside device, using the formula: TPO = ((SSHT - CSHT) + (SEHT - CEHT)) I 2. Block 920 transfers control to block 925.

[0207] Next, in block 925, computing, by client-side device, a synchronized client-side running media play time (CRMPT) to align with the server's media timeline, using one or more of the SRMPT and the TPO. Block 925 transfers control to block 930.

[0208] Next, in block 930, concurrently sending, by client-side device to serverside device, a request to obtain server-side device location coordinates. Block 930 transfers control to block 935.

[0209] Next, in block 935, receiving, by client-side device, said server-side device location coordinates (xs, ys, zs); calculating, by client-side device, its own position (xc, yc, zc) via one or more Object Recognition Al Application with known location coordinates. Block 935 transfers control to block 940.

[0210] Next, in block 940, calculating, by client-side device, an azimuth angle 0, using 0 = arctan2((ys - yc) / (xs - xc)). Calculating, by client-side device, an elevation angle cp using cp = arctan((zs - zc) / sqrt((xs - xc)2+ (ys - yc)2)). Block 940 transfers control to block 945.

[0211] Next, in block 945, calculating, by client-side device, a first delay corresponding to the acoustic propagation time from server-side device location to client-side device location. Block 945 transfers control to block 950.

[0212] Next, in block 950, applying, by client-side device, one or more audio filters to an audio content, wherein said application comprises: selecting a Head- Related Transfer Function (HRTF) or Head-Related Impulse Response (HRIR) signal corresponding to the calculated azimuth and elevation angles. Applying the selected HRTF / HRIR signal to render directionality. Block 950 transfers control to block 955.

[0213] Next, in block 955, applying, by client-side device, loudness compensation signals to adjust for the relative distance between the device.Further processing the audio using one or more of a Binaural Room Impulse Response (BRIR), a Headphone Transfer Function (HpTF), speed-adaptive equalization filters, noise-cancellation, and a voice isolator, wherein said BRIR may be a pre-measured response or may be generated by rendering a more general spatial representation, such as a Spatial Room Impulse Response (SRIR), for the current user orientation. Block 955 transfers control to block 960.

[0214] Next, in block 960, calculating, by a processor on client-side device, a second delay corresponding to the signal processing latency introduced by the application of said audio filters. Block 960 transfers control to block 965.

[0215] Next, in block 965, applying, by client-side device, a final compensating delay to client-side playback, said final compensating delay accounting for the first delay and the second delay to ensure the processed audio is aligned with the synchronized timeline. Block 965 transfers control to block 970.

[0216] Next, in block 970, synchronizing, by client-side device, using the CRMPT, client-side playback of synchronized and spatialized mixed audio content and media content to server-side playback of audio and media content. Block 970 then terminates the process.

[0217] It is to be understood that while the embodiment described in blocks 905 through 970 specifies the selection and application of Head-Related Impulse Responses (HRIR) and Binaural Room Impulse Responses (BRIR), the scope of the invention is not limited to these specific filter formats. The spatialization filters and responses applied by client-side device may be derived, synthesized, or rendered from any suitable spatial audio representation. Such representations include, but are not limited to, higher-order Ambisonics, spherical harmonic coefficients, or Spatial Room Impulse Responses (SRIRs). The process of calculating orientation angles (0, <p) and using them to generate a specific binaural output from these broader spatial representations is contemplated as being within the scope of the present invention.

[0218] FIGURE 10 is a flowchart of method 1000 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and Object Recognition Al Application. FIGURE 10 applies to a method using a Multimodal Al, as viewed from server-side.

[0219] The order of the steps in the method 1000 is not constrained to that shown in Figure 10 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0220] According to this method, client-side device performs tracking, synchronization, and spatialization of media content of client-side playback with server-side playback using a Multimodal Al System and Object Recognition Al Applications. This method enables server-side device to synchronize its media playback with client-side device in both time and 3D space. It calculates network latency and the physical positions of both device, then uses this data to apply directional audio filters. The result is a near-perfectly timed, spatialized audio experience that sounds as if it is coming from the server's near-precise physical location after compensating for acoustic and processing delays.

[0221] In block 1005, receiving, by server-side computing device, client-side message or packet from client-side master application at one or more pre-determined intervals during playback of content, client-side message or packet comprising one or more of client-side unique identifier and client-side start host time (CSHT). Block 1005 transfers control to block 1010.

[0222] Next, in block 1010, recording, by server-side device, server-side start host time (SSHT) upon reception of the packet. Concurrently receiving, by serverside device from client-side device, client-side device location coordinates (xc, yc, zc); calculating, by server-side device, its own position (xs, ys, zs) via one or more Object Recognition Al Application. Block 1010 transfers control to block 1015.

[0223] Next, in block 1015, generating, by server-side device, an audio content track based on a user prompt or other data. Generating, by server-side device, an ambient audio track based on application data. Creating, by server-side device, a mixed audio content by combining said audio content track and said ambient audio track. Block 1015 transfers control to block 1020.

[0224] Next, in block 1020, calculating, by server-side device, an azimuth angle 0 and an elevation angle cp based on the relative positions of server-side device and client-side device. Block 1020 transfers control to block 1025.

[0225] Next, in block 1025, calculating, by server-side device, a first delay corresponding to an acoustic propagation time between server-side device and client-side device. Block 1025 transfers control to block 1030.

[0226] Next, in block 1030, calculating, by a processor on server-side device, a second delay corresponding to the signal processing latency introduced by said filtering. Block 1030 transfers control to block 1035.

[0227] Next, in block 1035, recording, by server-side device, server-side end host time (SEHT) prior to transmitting a response; calculating a half-round-trip time (HRT), by server-side device, using the formula: HRT = ((SEHT - SSHT) - (CEHT - CSHT)) / 2, where CEHT is client-side end host time received in a subsequent packet; calculating a playback time offset (TPO) representing clock drift, by serverside device, using the formula: TPO = ((SSHT - CSHT) + (SEHT - CEHT)) I 2. Block 1035 transfers control to block 1040.

[0228] Next, in block 1040, computing, by server-side device, an authoritative server-side running media play time (SRMPT) adjusted for said TPO. Creating, by server-side device, server-side response packet comprising the authoritative SRMPT, the final spatialized audio content, and the calculated second delay. Block 1040 transfers control to block 1045.

[0229] Next, in block 1045, transmitting, by server-side device, said serverside response packet to client-side device, said packet to be used by client-side device to apply a compensating delay and synchronize its playback to server-side playback. Block 1045 transfers control to block 1050.

[0230] Next, in block 1050, playing back, by server-side device, its own mixed audio content and media content through one or more speaker systems, said playback being synchronized with the playback occurring on client-side device. Block 1050 then terminates the process.

[0231] FIGURE 11 is a flowchart of method 1100 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and Object Recognition Al Application. FIGURE 11 applies to a method using a Multimodal Al, as viewed from client-side.

[0232] The order of the steps in the method 1100 is not constrained to that shown in Figure 11 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0233] In more embodiments of the invention, a method for applying a spatialization filter to a plurality of audio system formats includes.

[0234] Next, in block 1105, calculating, by client-side device, a filter tracking response, wherein said response comprises a set of parameters representing the real-time spatial relationship between client-side device and one or more sound sources, said parameters including one or more of an azimuth angle, an elevation angle, and a propagation delay. Block 1105 transfers control to block 1110.

[0235] Next, in block 1110, applying, by client-side device, the filter tracking response to a selected audio content format, wherein the method of application is adapted based on the architecture of the selected format. Block 1910 transfers control to block 1115.

[0236] Next, in block 1115, for an object-based audio system: applying the filter tracking response to one or more individual audio objects, thereby rendering each object at the specific spatial position defined by the filter parameters. Block 1115 transfers control to block 1120.

[0237] Next, in block 1120, for embodiments featuring a channel-based audio system, such as a stereo, 5.1, or 7.1 system, the process comprises utilizing the filter tracking response to generate a binaural Tenderer or a virtualizer. Said virtualizer is configured to map the fixed channels of the system to a three- dimensional soundfield to simulate an immersive audio experience. Block 1120 transfers control to block 1125.

[0238] Next, in block 1125, for embodiments featuring a scene-based audio system, such as an Ambisonics system, the process comprises utilizing the filter tracking response to rotate the entire captured soundfield in real-time, thereby aligning the orientation of the audio with the user's head orientation. Block 1125 transfers control to block 1130.

[0239] Next, in block 1130, for a parametric binaural signal: utilizing the filter tracking response to modify the spatial parameters of the parametric binaural signal, thereby adjusting the perceived location of the audio content. Block 1130 transfers control to block 1135.

[0240] Next, in block 1135, for a pre-rendered binaural audio track: applying the filter tracking response to rotate the entire pre-rendered binaural soundfield, thereby aligning the orientation of the audio with the user's head orientation in realtime. Block 1135 then terminates the process.

[0241] FIGURE 12 is a flowchart of method 1200 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and Object Recognition Al Application. FIGURE 12 applies to a method using a Multimodal Al, as viewed from client-side.

[0242] The order of the steps in the method 1200 is not constrained to that shown in Figure 12 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0243] According to a streaming embodiment of the invention, a method for processing streamed media content on client-side device.

[0244] In block 1205, establishing a streaming session between client-side device and server-side streaming application, wherein a real-time transport protocol is selected based on the low-latency requirements of the interactive Al response loop. Block 1205 transfers control to block 1210.

[0245] Next, in block 1210, continuously receiving, by client-side device, a stream of data packets, wherein said packets comprise one or more audio content tracks and one or more audible or inaudible reference sound signals emitted by server-side device for tracking purposes. Block 1210 transfers control to block 1215.

[0246] Next, in block 1215, placing, by client-side device, the received data packets into an adaptive jitter buffer. Said buffer is configured to dynamically adjust its size based on real-time network conditions, such as latency and packet arrival variance, and the latency tolerance of the current application state. Block 1215 transfers control to block 1220.

[0247] Next, in block 1220, periodically creating and sending, by client-side device, client-side packet containing client-side unique identifier and client-side start host time (CSHT) to server-side device to initiate a synchronization event. Block 1220 transfers control to block 1225.

[0248] Next, in block 1225, receiving and processing, by client-side device, server-side packet comprising server-side unique identifier and various timestamps (CSHT, SSHT, SEHT, SRMPT), and synchronizing, in real-time, clientside playback timeline with server-side playback timeline using said received timestamps. Block 1225 transfers control to block 1230.

[0249] Next, in block 1230, calculating, by client-side Object Recognition Al Application, a final, authoritative set of real-time location and orientation data. Said calculation comprises: determining client-side device's own position (Xc,Yc,Zc); determining server-side device's position (Xs,Ys,Zs) by adaptively fusing data from a plurality of sensors, such as cameras and microphones analyzing reference sound signals, using real-time confidence scores; and calculating the current relative azimuth (0) and elevation (4>) angles between the client and server device, wherein said calculation comprises: a) determining a baseline direction vector between server-side device's location coordinates (Xs,Ys,Zs) and client-side device's own determined location coordinates (Xc,Yc,Zc); b) applying a rotation transformation to the baseline direction vector based on real-time orientation data received from a user's head tracking sensor; and c) calculating the azimuth and elevation angles from the resulting rotated vector. Block 1230 transfers control to block 1235.

[0250] Next, in block 1235, after extracting data packets from the buffer according to the synchronized timeline, detecting any lost packets and applying a spatially-aware Packet Loss Concealment (PLC) algorithm. Said algorithm is made more effective by using predictive data, such as server-side device's velocity and trajectory as determined by the Object Recognition Al Application, to intelligently prioritize the reconstruction of audio data most critical for maintaining a stable perceived spatial location. Block 1235 transfers control to block 1240.

[0251] Next, in block 1240, decoding, by client-side device, the synchronized and repaired data to generate one or more clean audio channels formatted for direct input into a spatialization engine. Block 1240 transfers control to block 1245.

[0252] Next, in block 1245, applying, by client-side device, one or more audio filters to the decoded audio, such as convolving the audio with a selected pair of Head-Related Impulse Responses (HRIR) corresponding to the real-time azimuth (0) and elevation (cp) angles calculated in block 1230, to generate a binaural stereo output. Block 1245 transfers control to block 1250.

[0253] Next, in block 1250, calculating, by client-side device, a first delay corresponding to the acoustic propagation time based on the distance between server-side device location (Xs, Ys, Zs) and client-side device location (Xc, Yc, Zc) determined in block 1230. Block 1250 transfers control to block 1255.

[0254] Next, in block 1255, calculating, by client-side device, a second delay corresponding to the total signal processing latency introduced by the application of the spatialization filters in the preceding step. Block 1255 transfers control to block 1260.

[0255] Next, in block 1260, applying, by client-side device, a final compensating delay to client-side playback timeline, said final delay accounting for the calculated first delay (acoustic propagation) and the second delay (processing latency). Block 1260 transfers control to block 1265.

[0256] Next, in block 1265, playing back, by client-side device, the fully processed mixed audio content, which is now aligned with server-side playback timeline and is rendered as a fully spatialized audio experience that accurately reflects the real-time position of server-side device as determined by the Object Recognition Al Application. Block 1265 then terminates the process.

[0257] FIGURE 13 is a flowchart of method 1300 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and Object Recognition Al Application. FIGURE 13 applies to a method viewed from server-side.

[0258] The order of the steps in the method 1300 is not constrained to that shown in Figure 13 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0259] According to a streaming embodiment of the invention, a method for processing streamed media content on server-side device.

[0260] In block 1305, establishing, by server-side device, a real-time streaming session with client-side device in response to a content request, wherein a low- latency transport protocol is selected to support an interactive Al response loop. Block 1305 transfers control to block 1310.

[0261] Next, in block 1310, generating or retrieving from a database, by serverside device, one or more audio content tracks corresponding to the content request. Block 1310 transfers control to block 1315.

[0262] Next, in block 1315, concurrently generating, by server-side device, one or more audible or inaudible reference sound signals, said signals having known characteristics designed for detection and analysis by client-side device's Object Recognition Al Application. Block 1315 transfers control to block 1320.

[0263] Next, in block 1320, commencing the continuous transmission of data packets to client-side device, wherein said stream comprises both the audio content tracks and the reference sound signals. Block 1320 transfers control to block 1325.

[0264] Next, in block 1325, receiving, by server-side device, a synchronization event request packet from client-side device, said packet comprising client-side unique identifier and client-side start host time (CSHT). Block 1325 transfers control to block 1330.

[0265] Next, in block 1330, in response to the synchronization request, creating and transmitting, by server-side device, server-side synchronization packet to client-side device. Said packet comprises the original CSHT, server-side start host time (SSHT), server-side end host time (SEHT), and server-side running media play time (SRMPT), thereby providing client-side device with necessary information to align its playback timeline. Block 1330 transfers control to block 1335.

[0266] Next, in block 1335, commencing or continuing, by server-side device, its own local playback of the audio content tracks, said playback being aligned with server-side running media play time (SRMPT) to ensure synchronization with client-side device. The process may then continue to stream content and respond to subsequent synchronization requests.

[0267] FIGURE 14 is a flowchart of method 1400 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and Object Recognition Al Application. FIGURE 14 applies to a method viewed from client-side.

[0268] The order of the steps in the method 1400 is not constrained to that shown in Figure 14 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0269] According to an embodiment of the invention, a method for rendering a synchronized and dynamically directed binaural audio experience comprises receiving, from a remote server-side device, a data package containing a plurality of separated audio sub-tracks, a corresponding stream of semantic event markers, and at least one reference signal for spatial tracking; subsequently determining, by client-side device, a real-time spatial relationship (including azimuth, elevation, and distance) to server-side device by processing said reference signal; while concurrently processing the stream of semantic event markers with client-side Al Audio Mixing System to generate a set of dynamic mixing parameters. The method further comprises applying audio filters to the sub-tracks, wherein the application of said filters is dictated by both the calculated spatial relationship and the dynamic mixing parameters, and continuously synchronizing playback timelines while compensating for acoustic and processing latencies, thereby rendering a final audio experience that is not only spatially accurate but also dynamically reflects the live event's narrative context.

[0270] In block 1405, capturing, by client-side device controlled by a user, a verbal command to initiate a focused-listening session, and in response, establishing a streaming session between client-side device and server-side streaming application, wherein a real-time transport protocol is selected based onthe low-latency requirements of the interactive Al response loop. Block 1405 transfers control to block 1410.

[0271] Next, in block 1410, continuously receiving, by client-side device, a high-information-density data package from server-side device. Said package is demultiplexed into its constituent parts: a plurality of discrete audio subtracks, such as 'commentator,' 'crowd,' and 'player impacts'; a stream of semantic event markers; and one or more audible or inaudible reference sound signals emitted by server-side device for tracking purposes. Block 1410 transfers control to block 1415.

[0272] Next, in block 1415, processing, by client-side Al Audio Mixing System, the stream of semantic event markers to understand the real-time narrative context of the event and to dynamically generate a set of mixing parameters for each discrete audio sub-track designed to heighten dramatic effect. Block 1415 transfers control to block 1420.

[0273] Next, in block 1420, placing, by client-side device, the received and demultiplexed data packets (comprising the sub-tracks and reference signals) and the generated mixing parameters into an adaptive jitter buffer. Said buffer dynamically adjusts its size based on real-time network conditions and the latency tolerance of the application. Block 1420 transfers control to block 1425.

[0274] Next, in block 1425, periodically creating and sending, by client-side device, client-side packet containing client-side unique identifier and client-side start host time (CSHT) to server-side device to initiate a synchronization event. Block 1425 transfers control to block 1430.

[0275] Next, in block 1430, receiving and processing, by client-side device, server-side packet comprising server-side unique identifier and various timestamps (CSHT, SSHT, SEHT, SRMPT), and synchronizing, in real-time, clientside playback timeline with server-side playback timeline using said received timestamps. Block 1430 transfers control to block 1435.

[0276] Next, in block 1435, calculating, by client-side Object Recognition Al Application, a final, authoritative set of real-time location and orientation data. Said calculation comprises: determining client-side device's own position (Xc,Yc,Zc); determining server-side device's position (Xs,Ys,Zs) by adaptively fusing data from a plurality of sensors, such as cameras and microphones analyzing reference sound signals, using real-time confidence scores; and calculating the current relative azimuth (0) and elevation ( ) angles between the client and server device. Block 1435 transfers control to block 1440.

[0277] Next, in block 1440, after extracting data packets for each sub-track from the buffer according to the synchronized timeline, detecting any lost packets and applying a spatially-aware Packet Loss Concealment (PLC) algorithm. Said algorithm uses predictive data, such as server-side device's velocity and trajectory as determined by the Object Recognition Al Application, to intelligently prioritize the reconstruction of audio data most critical for maintaining a stable perceived spatial location. Block 1440 transfers control to block 1445.

[0278] Next, in block 1445, decoding, by client-side device, the synchronized and repaired data to generate one or more clean audio channels for each discretesub-track, formatted for direct input into a spatialization engine. Block 1445 transfers control to block 1450.

[0279] Next, in block 1450, applying one or more audio filters to each of the one or more audio content sub-tracks, wherein said application is dictated by both the spatialization parameters (9, cp) calculated in block 1435 and the dynamic mixing parameters determined by the Al mixing system in block 1415, to generate a final mixed and spatialized audio content. For example, the ‘impacts’ sub-track is rendered from the calculated azimuth and elevation, and its gain is momentarily increased in response to the Al mixing system processing a [EVENT: TACKLE] marker. Block 1450 transfers control to block 1455.

[0280] Next, in block 1455, calculating, by client-side device, a first delay corresponding to the acoustic propagation time based on the distance between server-side device location (Xs, Ys, Zs) and client-side device location (Xc, Yc, Zc) determined in block 1435. Block 1455 transfers control to block 1460.

[0281] Next, in block 1460, calculating, by client-side device, a second delay corresponding to the total signal processing latency introduced by the application of the spatialization and mixing filters in the preceding step. Block 1460 transfers control to block 1465.

[0282] Next, in block 1465, applying, by client-side device, a final compensating delay to client-side playback timeline, said final delay accounting for the calculated first delay (acoustic propagation) and the second delay (processing latency). Block 1465 transfers control to block 1470.

[0283] Next, in block 1470, playing back, by client-side device, the fully processed mixed audio content, which is now aligned with server-side playback timeline and is rendered as a dynamically directed, narrative-driven auditory experience that accurately reflects the real-time position of server-side device and enhances the dramatic context of the live event. The process then continues by receiving the next data package at block 1410.

[0284] FIGURE 15 is a flowchart of method 1500 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and Object Recognition Al Application. FIGURE 15 applies to a method viewed from server-side.

[0285] The order of the steps in the method 1500 is not constrained to that shown in Figure 15 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0286] According to an embodiment of the invention, a method for generating a high-information-density data stream for remote auditory rendering comprises applying server-side Multimodal Deconstruction Al System to captured sensor data to perform real-time source separation on raw audio to create distinct audio sub-tracks, concurrently performing semantic event detection to generate a corresponding stream of event markers describing a participant's actions, and transmitting said sub-tracks and event markers as a single data packet.

[0287] In block 1505, receiving, by server-side streaming application, a request from client-side device to initiate a focused-listening session, said request comprising client-side unique identifier. In response, establishing a streaming session with client-side device, wherein a real-time transport protocol is selectedbased on the low-latency requirements of the interactive Al response loop. Block 1505 transfers control to block 1510.

[0288] Next, in block 1510, capturing, by server-side device, multi-channel audio from the surrounding environment using a plurality of microphones. Block 1510 transfers control to block 1515.

[0289] Next, in block 1515, processing, by server-side Multimodal Deconstruction Al System, the captured multi-channel audio to perform real-time source separation, thereby generating a plurality of discrete audio sub-tracks, such as 'commentator,' 'crowd,' and 'player impacts'. Block 1515 transfers control to block 1520.

[0290] Next, in block 1520, concurrently analyzing, by server-side Multimodal Deconstruction Al System, the captured audio and / or corresponding video data to generate a continuous stream of semantic event markers that describe the narrative context of the live event, such as [EVENT: TACKLE] or [EVENT: GOAL], Block 1520 transfers control to block 1525.

[0291] Next, in block 1525, generating one or more audible or inaudible reference sound signals designed to be emitted by server-side device's speakers, wherein said reference signals are structured to enable near-precise spatial tracking by client-side device. Block 1525 transfers control to block 1530.

[0292] Next, in block 1530, multiplexing the discrete audio sub-tracks, the stream of semantic event markers, and the reference sound signals into a single, high-information-density data package. Block 1530 transfers control to block 1535.

[0293] Next, in block 1535, continuously transmitting, by server-side device, the stream of data packages to client-side device over the established real-time transport protocol. Block 1535 transfers control to block 1540.

[0294] Next, in block 1540, concurrently listening for and receiving, by serverside device, a synchronization packet from client-side device, said packet containing client-side unique identifier and client-side start host time (CSHT). Upon reception, recording server-side host time (SSHT) is recorded. Block 1540 transfers control to block 1545.

[0295] Next, in block 1545, in immediate response to receiving client-side synchronization packet, preparing and transmitting server-side synchronization packet back to client-side device. Said server-side packet comprises the original CSHT received from the client, the recorded SSHT, a newly recorded server-side end host time (SEHT) captured just prior to transmission, and the current serverside running media play time (SRMPT). This complete set of timestamps enables client-side device to synchronize its playback timeline with the server's timeline. The process then continues by returning to block 1510 to capture, process, and stream the next segment of live event data.

[0296] FIGURE 16 is a flowchart of method 1600 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and Object Recognition Al Application. FIGURE 16 applies to a method viewed from client-side.

[0297] The order of the steps in the method 1600 is not constrained to that shown in Figure 16 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0298] According to an embodiment of the invention, a method for rendering a synchronized, multi-sensory haptic and auditory experience comprises receiving, by client-side device, a high-information-density data package comprising a plurality of discrete audio sub-tracks and a corresponding stream of semantic event markers; processing, by client-side Al Sensory Experience Application, said stream of semantic event markers to understand the real-time context of an event; dynamically generating, based on said context, a first set of audio mixing parameters and a second, corresponding set of haptic command signals; and rendering a final output by applying said audio mixing parameters to the audio subtracks while concurrently actuating said haptic command signals on a peripheral haptic device in synchronization with the audio playback.

[0299] In block 1605, capturing, by client-side device controlled by a user, a verbal command as a continuous audio stream and concurrently capturing a continuous video stream via an onboard camera. At the same time, establishing a communicative coupling with one or more peripheral haptic device. Block 1605 transfers control to block 1610.

[0300] Next, in block 1610, transmitting, by client-side device, a data transmission to server-side device, said transmission comprising client-side unique identifier, the continuous audio stream, and the continuous video stream. Block 1610 transfers control to block 1615.

[0301] Next, in block 1615, receiving, from server-side device, a high- information-density data package. Said package is demultiplexed by client-side device into its constituent parts: one or more source audio tracks; a stream of haptically-enriched semantic event markers; real-time location coordinates for one or more identified objects; and a timing data set (CSHT, SSHT, SEHT, SRMPT). Block 1615 transfers control to block 1620.

[0302] Next, in block 1620, processing, by client-side Al Sensory Experience Application, the stream of haptically-enriched semantic event markers to understand the real-time action. Based on this understanding, the application dynamically generates mixing parameters for the source audio tracks and corresponding haptic command signals, translating the event marker attributes, such as event type, intensity, and duration, into specific audio adjustments and haptic effect instructions. Block 1620 transfers control to block 1625.

[0303] Next, in block 1625, performing a time synchronization process using the received timing data set to align client-side device’s media playback timeline with server-side device. Block 1625 transfers control to block 1630.

[0304] Next, in block 1630, calculating, by client-side device, the real-time azimuth (0), elevation (cp), and a first delay (acoustic propagation) for each source audio track, using the location coordinates of the corresponding identified object relative to the user's own position. Block 1630 transfers control to block 1635.

[0305] Next, in block 1635, generating a final spatialized audio soundscape by applying audio filters, such as Head-Related Impulse Responses (HRIRs), to each source audio track. The rendering is dictated by both the spatialization parameters (0, 4>) and the dynamic mixing parameters determined by the Al Sensory Experience Application. Block 1635 transfers control to block 1640.

[0306] Next, in block 1640, concurrently generating a final haptic output command. This comprises selecting a haptic effect from a library based on the haptic command signals and spatializing said effect by directing it to a location on the peripheral haptic device that corresponds to the on-screen event's direction, as determined by the object location coordinates. Block 1640 transfers control to block 1645.

[0307] Next, in block 1645, calculating a plurality of compensating delays, including a second delay for audio processing latency and a third delay for haptic system latency, to ensure all sensory outputs are near-perfectly synchronized with the live broadcast timeline established in block 1625. Block 1645 transfers control to block 1650.

[0308] Next, in block 1650, rendering the final multi-sensory experience by playing back the dynamic, spatialized audio soundscape while simultaneously actuating the synchronized and spatialized haptic commands on the peripheral haptic device. The process may then return to block 1615 to continuously receive and process updated data packages from server-side device. Block 1650 then terminates the process.

[0309] FIGURE 17 is a flowchart of method 1700 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and Object Recognition Al Application. FIGURE 17 applies to a method viewed from server-side.

[0310] The order of the steps in the method 1700 is not constrained to that shown in Figure 17 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0311] According to an embodiment of the invention, a method for generating a high-information-density data package for enabling a remote multi-sensory experience comprises: receiving, by server-side device, a continuous audio stream and a continuous video stream from client-side device; applying a single, natively Multimodal Al System Application that integrates an Object Recognition Al Application to the received streams; concurrently performing, by said Multimodal Al System Application, real-time analysis to: (a) generate one or more source audio tracks based on an interpretation of the audio stream; (b) identify one or more objects within the video stream and determine their real-time location coordinates; and (c) generate a stream of haptically-enriched semantic event markers that describe real-time actions and include attributes necessary for generating corresponding haptic effects; and packaging and transmitting said source audio tracks, said real-time location coordinates, and said stream of haptically-enriched semantic event markers as a single, synchronized data package to client-side device for processing and rendering of a multi-sensory experience.

[0312] In block 1705, receiving, by server-side device, a data transmission from client-side device, said transmission comprising client-side unique identifier, a continuous audio stream capturing a user's verbal command, and a continuous video stream from client-side device's camera. Block 1705 transfers control to block 1710.

[0313] Next, in block 1710, processing, by server-side device, the continuous audio and video streams using a single, natively Multimodal Al System Application. Said model integrates an Object Recognition Al Application to: directly interpret the user's verbal command from the audio stream; identify one or more objects and their spatial coordinates within the video stream; and generate a stream of haptically-enriched semantic event markers. Said markers describe real-time actions and environmental states while including attributes, such as event type, intensity, duration, and texture, necessary for client-side device to generate corresponding haptic command signals. Block 1710 transfers control to block 1715.

[0314] Next, in block 1715, generating, by server-side device based on the Multimodal Al System Application's output, a high-information-density data package. Said package comprises: one or more source audio tracks suitable for subsequent spatialization; the stream of haptically-enriched semantic event markers; and real-time location coordinates for the one or more identified objects. Block 1715 transfers control to block 1720.

[0315] Next, in block 1720, receiving, by server-side device, a time synchronization request packet from client-side device comprising client-side start host time (CSHT), and in response, recording server-side start host time (SSHT) and server-side end host time (SEHT). Block 1720 transfers control to block 1725.

[0316] Next, in block 1725, transmitting, by server-side device, the high- information-density data package to client-side device. Said package comprises: the one or more generated source audio tracks, server-side unique identifier, the stream of haptically-enriched semantic event markers, the real-time location coordinates for the identified objects, and a timing data set including the CSHT, SSHT, SEHT, and server-side running media play time (SRMPT). Block 1725 transfers control to block 1730.

[0317] Next, in block 1730, commencing, by server-side device, its own synchronized playback of the source audio tracks. Server-side device continues to receive and process the real-time audio and video streams from client-side device, transmitting updated data packages to enable the client to render a fully dynamic and interactive multi-sensory experience comprising spatialized audio and haptic feedback. The process may then await the next request from client-side device. Block 1730 then terminates the process.

[0318] FIGURE 18 is a flowchart of method 1800 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and Object Recognition Al Application. FIGURE 18 applies to a method viewed from client-side.

[0319] The order of the steps in the method 1800 is not constrained to that shown in Figure 18 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0320] According to an embodiment of the invention, a method for determining the real-time location coordinates of client-side device using acoustic Time Difference of Arrival (TDoA) comprises: receiving, by a microphone array on clientside device, one or more test sound signals from a plurality of fixed acoustic beacons with known locations, and processing said signals via cross-correlationwith a reference signal to determine an arrival time at each microphone; calculating one or more time differences of arrival between the microphones based on said arrival times and converting said time differences into distance differences using the speed of sound; formulating a system of hyperbolic equations based on the known locations of the acoustic beacons and the calculated distance differences; and solving said system of equations using a numerical solver to obtain a final, authoritative set of location coordinates for client-side device.

[0321] In block 1805, a method for determining the real-time location coordinates (Xc, Yc, Zc) of client-side device is initiated. The method operates within a physical environment equipped with a plurality of fixed acoustic beacons, wherein each beacon has known, pre-determined location coordinates and is configured to emit one or more audible or inaudible test sound signals. Block 1805 transfers control to block 1810.

[0322] Next, in block 1810, receiving, by a microphone array on client-side device comprising at least three microphones, such as microphones ma, mb, and me, the one or more test sound signals emitted from one of the fixed acoustic beacons. Block 1810 transfers control to block 1815.

[0323] Next, in block 1815, processing, by client-side device, the received test sound signals by applying a normalization and noise cancellation algorithm to improve signal clarity. Block 1815 transfers control to block 1820.

[0324] Next, in block 1820, identifying an arrival time for the test sound signal at each microphone. This is achieved by performing a mathematical crosscorrelation between the processed signal received at each microphone and a stored digital template of the corresponding reference sound signal. The peak of the correlation function indicates a definitive match and provides the arrival times ta, tb, and te for the respective microphones ma, mb, and me. Block 1820 transfers control to block 1825.

[0325] Next, in block 1825, calculating, by client-side device, the time differences of arrival between the microphones. For example, the time differences are calculated as Atab = tb - ta and Atac = tc - ta, using microphone ma as the reference. Block 1825 transfers control to block 1830.

[0326] Next, in block 1830, converting, by client-side device, the calculated time differences into distance differences by multiplying them by the speed of sound (c). For example, the distance differences are calculated as Adab = c * Atab and Adac = c * Atac. Block 1830 transfers control to block 1835.

[0327] Next, in block 1835, formulating, by client-side device, a system of hyperbolic equations based on the principle of Time Difference of Arrival (TDoA). The known locations of the fixed acoustic beacons and the calculated distance differences define a set of hyperboloids, the unique intersection of which corresponds to the location of client-side device's microphone array. Block 1835 transfers control to block 1840.

[0328] Next, in block 1840, solving, by client-side device, the system of hyperbolic equations using a numerical solver, such as a least-squares algorithm, to obtain a final, authoritative set of its own location coordinates (Xc,Yc,Zc). Block 1840 then terminates the self-localization process.

[0329] In a further embodiment, server-side device may be equipped with its own microphone array and may be configured to employ the same method as described in blocks 1810 through 1840. By receiving and processing the same test sound signals from the same fixed acoustic beacons, server-side device can independently and concurrently determine its own real-time location coordinates (Xs, Ys, Zs) within the shared physical environment.

[0330] FIGURE 19 is a flowchart of method 1900 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and Object Recognition Al Application. FIGURE 19 applies to a method viewed from client-side.

[0331] The order of the steps in the method 1900 is not constrained to that shown in Figure 19 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0332] In block 1905, a method for determining the real-time location coordinates (Xc, Yc, Zc) of client-side device is initiated. The method operates within a physical environment equipped with a plurality of fixed acoustic beacons, wherein each beacon has known, pre-determined location coordinates and is configured to emit one or more audible or inaudible test sound signals. Block 1905 transfers control to block 1910.

[0333] Next, in block 1910, receiving, by a microphone array on client-side device comprising at least three microphones, such as microphones ma, mb, and me, the one or more test sound signals emitted from the fixed acoustic beacons. Block 1910 transfers control to block 1915.

[0334] Next, in block 1915, processing, by client-side device, the received test sound signals by applying a normalization and noise cancellation algorithm to improve signal clarity. Block 1915 transfers control to block 1920.

[0335] Next, in block 1920, identifying an arrival time for the test sound signal at each microphone using a trained deep learning model. Said model, which may be a Convolutional Neural Network (CNN) trained for time-series analysis, is configured to receive the processed signal from each microphone as an input. The model analyzes the input to identify learned, characteristic features of the reference sound signal, thereby distinguishing it from ambient noise and acoustic reflections (echoes). The output of the model is the arrival times ta, tb, and tc at which the reference signal was detected by the respective microphones ma, mb, and me. Block 1920 transfers control to block 1925.

[0336] Next, in block 1925, calculating, by client-side device, the time differences of arrival between the microphones. For example, the time differences are calculated as Atab = tb - ta and Atac = tc - ta, using the arrival times provided by the deep learning model. Block 1925 transfers control to block 1930.

[0337] Next, in block 1930, converting, by client-side device, the calculated time differences into distance differences by multiplying them by the speed of sound (c). For example, the distance differences are calculated as Adab = c * Atab and Adac = c * Atac. Block 1930 transfers control to block 1935.

[0338] Next, in block 1935, formulating, by client-side device, a system of hyperbolic equations based on the principle of Time Difference of Arrival (TDoA),wherein the known locations of the fixed acoustic beacons and the calculated distance differences define the system. Block 1935 transfers control to block 1940.

[0339] Next, in block 1940, solving, by client-side device, the system of hyperbolic equations using a numerical solver to obtain a final, authoritative set of its own location coordinates (Xc, Yc, Zc). Block 1940 then terminates the selflocalization process.

[0340] In a further embodiment, server-side device may be equipped with its own microphone array and may be configured to employ the same Al-based localization method as described in blocks 1910 through 1940. By receiving and processing the same test sound signals from the same fixed acoustic beacons using its own trained deep learning model, server-side device can independently and concurrently determine its own real-time location coordinates (Xs, Ys, Zs) within the shared physical environment.

[0341] FIGURE 20 is a flowchart of method 2000 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 20 applies to a method viewed from client-side.

[0342] The order of the steps in the method 2000 is not constrained to that shown in Figure 20 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0343] According to an embodiment of the invention, a method for rendering a real-time, spatialized, and voice-preserving translated audio experience comprises: receiving, by client-side device, a data stream comprising a primary audio track in a source language and corresponding spatialization metadata; processing, by client-side Multimodal Al Media Application, the source language audio to perform a real-time translation into a target language, wherein said application concurrently synthesizes the translated audio to preserve the vocal identity and emotional tone of the original speaker; applying one or more spatialization filters to the resulting translated audio track based on the spatialization metadata to generate a binaural output; and applying one or more compensating delays to ensure the final output is aligned with a synchronized playback timeline.

[0344] In block 2005, establishing a streaming session between client-side device and server-side streaming application, wherein a real-time transport protocol is selected based on the low-latency requirements of the interactive Al response loop. Block 2005 transfers control to block 2010.

[0345] Next, in block 2010, continuously receiving, by client-side device, a stream of data packets, wherein said packets comprise a primary audio track in a source language and spatialization metadata, such as the real-time position of the audio source (Xs, Ys, Zs). Block 2010 transfers control to block 2015.

[0346] Next, in block 2015, placing the received data packets into an adaptive jitter buffer to ensure a smooth and stable stream by dynamically adjusting its size based on real-time network conditions. Block 2015 transfers control to block 2020.

[0347] Next, in block 2020, performing a time synchronization handshake with server-side device by exchanging timestamp packets, such as CSHT,SSHT, and SRMPT, to align client-side playback timeline with server-side timeline. Block 2020 transfers control to block 2025.

[0348] Next, in block 2025, after extracting data from the buffer, detecting any lost packets and applying a spatially-aware Packet Loss Concealment (PLC) algorithm to reconstruct the missing audio data from the source language track. Block 2025 transfers control to block 2030.

[0349] Next, in block 2030, decoding the synchronized and repaired data to generate a clean, source language audio channel. Block 2030 transfers control to block 2035.

[0350] Next, in block 2035, processing, by client-side Multimodal Al Media Application, the clean source language audio channel to perform real-time translation. Said application translates the speech from the source language to a target language selected by the user. In a preferred implementation, the multimodal Al concurrently analyzes the vocal characteristics, such as pitch, cadence, and emotional tone, of the original speaker to generate a final translated audio track that preserves the vocal identity of the original speaker. Block 2035 transfers control to block 2040.

[0351] Next, in block 2040, applying one or more audio filters to the final translated audio track, such as convolving the translated audio with a selected pair of Head-Related Impulse Responses (HRIR) corresponding to a calculated azimuth (6) and elevation (cp), to generate a binaural stereo output. Block 2040 transfers control to block 2045.

[0352] Next, in block 2045, calculating, by client-side device, a first delay corresponding to the acoustic propagation time from the source's location to the client's location. Block 2045 transfers control to block 2050.

[0353] Next, in block 2050, calculating, by client-side device, a second delay corresponding to the total signal processing latency introduced by both the translation Al and the spatialization filters. Block 2050 transfers control to block 2055.

[0354] Next, in block 2055, applying a final compensating delay to client-side playback timeline, said final delay accounting for both the acoustic propagation delay (first delay) and the total processing latency (second delay). Block 2055 transfers control to block 2060.

[0355] Next, in block 2060, playing back, by client-side device, the fully processed audio content, which is now rendered as a near-perfectly synchronized, spatialized, and live-translated audio experience in the user's preferred language. Block 2060 then terminates the process.

[0356] FIGURE 21 is a flowchart of method 2100 for real-time multi-object tracking, synchronization, and spatialization of media content using Multimodal Al System and wireless communication sensor. FIGURE 21 applies to a method viewed from server-side.

[0357] The order of the steps in the method 2100 is not constrained to that shown in Figure 21 or described in the following discussion. Several of the steps could occur in a different order without affecting the final result.

[0358] According to an embodiment of the invention, a method for enabling a remote, real-time, spatialized audio translation comprises: retrieving or generating,by server-side device, a primary audio track in a source language; concurrently determining the real-time location coordinates of said server-side device to serve as spatialization metadata; transmitting said primary audio track and said location coordinates to a remote client-side device via a continuous data stream; and subsequently responding to a synchronization request from client-side device by transmitting a set of server-side timing data, thereby providing the client with necessary information to perform a remote translation and synchronized, spatialized rendering.

[0359] In block 2105, establishing, by server-side streaming application, a realtime streaming session with client-side device in response to a content request. Block 2105 transfers control to block 2110.

[0360] Next, in block 2110, retrieving or generating, by server-side device, a primary audio track in its original source language, for example, a foreign language. Block 2110 transfers control to block 2115.

[0361] Next, in block 2115, concurrently determining, by server-side device, its own real-time location coordinates (Xs, Ys, Zs), which represent the position of the audio source. Block 2115 transfers control to block 2120.

[0362] Next, in block 2120, commencing the continuous transmission of data packets to client-side device via a low-latency stream. Said stream comprises both the primary audio track in the source language and the corresponding real-time spatialization metadata (Xs, Ys, Zs). Block 2120 transfers control to block 2125.

[0363] Next, in block 2125, subsequently receiving, by server-side device from client-side device, a time synchronization packet, said packet comprising clientside start host time (CSHT). Block 2125 transfers control to block 2130.

[0364] Next, in block 2130, in response to the synchronization request, creating and transmitting, by server-side device, server-side synchronization packet to client-side device. Said packet comprises the original CSHT, server-side start host time (SSHT), server-side end host time (SEHT), and server-side running media play time (SRMPT), thereby providing client-side device with necessary information to align its playback timeline. Block 2130 transfers control to block 2135.

[0365] Next, in block 2135, continuing, by server-side device, its own local playback of the source language audio track, said playback being aligned with server-side running media play time (SRMPT) to ensure synchronization with client-side device. The process may then continue to stream content and respond to subsequent synchronization requests.

[0366] According to an embodiment of the invention, a method for determining the real-time location coordinates of client-side device using an Al-enhanced acoustic Time Difference of Arrival (TDoA) system comprises: receiving, by a microphone array on client-side device, one or more test sound signals from a plurality of fixed acoustic beacons with known locations; processing said signals using a trained deep learning model to identify characteristic features of the test signals and determine an arrival time at each microphone, thereby providing robustness against ambient noise and acoustic reflections; calculating one or more time differences of arrival between the microphones based on the Al-determined arrival times; formulating a system of hyperbolic equations based on the knownlocations of the beacons and said time differences; and solving said system of equations to obtain a final, authoritative set of location coordinates for client-side device.

[0367] In a further embodiment of the present disclosure, the one or more audible or inaudible reference sound signals may be configured to function as an acoustic data carrier, containing explicitly encoded information, including but not limited to the location coordinates of server-side device.

[0368] In a preferred embodiment, the reference sound signal serves a dual purpose. The physical properties of the signal's arrival (its Time of Flight and Angle of Arrival) are used for instantaneous, low-latency relative tracking. Concurrently, the signal itself contains a slowly updated, encoded data packet with the server's absolute coordinates.

[0369] While the above representative embodiments have been described with certain components in exemplary configurations, it will be understood those skilled in the art that other representative embodiments can be implemented using one or more of different configurations and different components. For example, it will be understood by those skilled in the art that the order of certain fabrication steps and certain components can be altered without substantially impairing the functioning of the invention.

[0370] For example, one or more of audio, video, and another entertainment format can be playing on client-side. For example, one of more of audio, video, and another entertainment format can be played on server-side.

[0371] For example, the steps of the flowchart depicted in Figure 8 may be implemented by one or more of client-side networked environment 105 and clientside device 120.

[0372] For example, the steps of the flowchart depicted in Figure 11 may be implemented by one or more of server-side networked environment 110 and server-side device 170.

[0373] For example, client-side master application and server-side master application may be implemented in client-side Multimodal Al System Application and server-side Multimodal Al System Application.

[0374] For example, client-side Multimodal Al System Application and serverside Multimodal Al System Application may communicate directly to other Multimodal Al Systems located in other networks.

[0375] For example, client-side Al Audio Mixing System and server-side Multimodal Deconstruction Al System may communicate directly to other Multimodal Al Systems located in other networks.

[0376] For example, instead of being located in client-side memory 140, one or more of client-side master application 150, client-side Multimodal Al Media Application 152, client-side Multimodal Al System Application 160, client-side Object Recognition Al Application 162, client-side Al Audio Mixing System 163 and client-side Al Sensory Experience Application 164 may be located in a section of client-side device 120 other than client-side memory 140.

[0377] For example, instead of being located in server-side memory 175, one or more of server-side master application 180, server-side streaming application 185, server-side Multimodal Al Media Application 190, server-side Multimodal AlSystem Application 191 , server-side Object Recognition Al Application 192 and server-side Multimodal Deconstruction Al System 193 may be located in one or more of server-side data storage 155 and a section of server-side computing device 170 other than server-side memory 175. For example, instead of being a freestanding component of server-side networked environment 110, server-side memory 175 may be located in server-side computing device 170.

[0378] For example, client-side data storage 145 may be separate from clientside device 120 rather than being comprised in client-side device 120. For example, server-side data storage 155 may be comprised in server-side computing device 170 rather than being separate from server-side computing device 170.

[0379] For example, instead of being two separate entities, client-side device 120 and client-side playback device 125 may be combined in client-side device 120. For example, server-side device 170 and server-side playback device 196 may be included in server-side device 170.The representative embodiments and disclosed subject matter, which have been described in detail herein, have been presented by way of example and illustration and not by way of limitation. It will be understood by those skilled in the art that various changes may be made in the form and details of the described embodiments, resulting in equivalent embodiments that remain within the scope of the invention. It is intended, therefore, that the subject matter in the above description shall be interpreted as illustrative and shall not be interpreted in a limiting sense.

Claims

CLAIMSWhat is claimed is:1 . A method performed by client-side computing device having network interface, the method comprising: receiving, via network interface from server-side device, plurality of data packets comprising audio content and one or more reference sound signal; determining, by one or more processor of client-side device executing Object Recognition Al Application, respective locations of client-side device and server-side device by adaptively fusing environmental data comprising plurality of sensor streams; calculating one or more spatialization parameter based on the determined locations of client-side device and server-side device; and applying one or more audio filter to said audio content based on said one or more spatialization parameter to generate spatialized audio output.

2. The method of claim 1 , wherein said plurality of sensor streams comprises visual data from one or more camera and acoustic data derived from analysis of said one or more reference sound signal.

3. The method of claim 2, wherein said adaptive fusing is weighted based on realtime confidence score associated with said visual data and said acoustic data to maintain continuous tracking of server-side device during periods of visual occlusion.

4. The method of claim 1 , wherein calculating said one or more spatialization parameter further comprises: receiving, from head tracking sensor, real-time head orientation data of a user; and applying rotation transformation to baseline direction vector based on said realtime head orientation data.

5. The method of claim 4, wherein applying said one or more audio filter comprises generating left channel of spatialized audio output by convolving said audio content with a left-ear Head-Related Impulse Response (HRIR), and generating right channel of the spatialized audio output by convolving said audio content with a right-ear HRIR.

6. The method of claim 1 , wherein determining the location of client-side device comprises dynamically selecting a positioning protocol from a plurality of available positioning protocols based on a set of monitored conditions indicating which of the available positioning protocols has a higher accuracy.

7. The method of claim 1 , further comprising: calculating a displacement magnitude between a previous location and a current location of server-side device; and in response to said displacement magnitude exceeding a pre-determined threshold, re-calculatingsaid one or more spatialization parameter and applying smooth transition algorithm between previously applied set of spatialization filter parameter and the re-calculated spatialization filter parameter.

8. The method of claim 1 , further comprising calculating a Half-Round-Trip Time (HRT) and a Playback Time Offset (TPO) using plurality of client-side and server-side timestamps to synchronize client-side playback timeline.

9. The method of claim 1 , further comprising applying a final compensating delay to playback of said spatialized audio output, said delay based on calculated acoustic propagation time and calculated signal processing latency.

10. The method of claim 9, wherein client-side device monitors ambient air temperature and relative humidity using one or more environmental sensor, and wherein said acoustic propagation time delay is calculated using adjusted speed of sound based on monitored ambient air temperature and monitored relative humidity.

11. The method of claim 9, further comprising applying one or more loudness compensation signal to said spatialized audio output to render perceived distance of server-side device.

12. The method of claim 1 , further comprising applying spatially-aware packet loss concealment (PLC) algorithm that utilizes predictive trajectory data of server-side device as determined by Object Recognition Al Application.

13. Client-side computing system, comprising: network interface; one or more processor communicatively coupled to network interface; and memory storing instruction that, when executed by said one or more processor, cause the system to perform the method of claim 1 .

14. The system of claim 13, wherein said plurality of sensor streams comprises visual data from one or more camera and acoustic data derived from analysis of said one or more reference sound signal, and wherein said adaptive fusing is weighted based on realtime confidence score to maintain tracking during periods of visual occlusion.

15. The system of claim 13, wherein the method further comprises applying smooth transition algorithm between filter sets in response to displacement of server-side device exceeding a pre-determined threshold.

16. The system of claim 13, wherein the method further comprises applying spatially- aware packet loss concealment (PLC) algorithm that utilizes predictive trajectory data of server-side device.

17. Non-transitory computer-readable storage medium having instruction stored thereon that, when executed by one or more processor of client-side device, cause said client-side device to perform the method of claim 1 .

18. Non-transitory computer-readable storage medium of claim 17, wherein said adaptive fusing of data from said plurality of sensor streams is weighted based on confidence score to maintain tracking during periods of visual occlusion.

19. Non-transitory computer-readable storage medium of claim 17, wherein the method further comprises applying smooth transition algorithm between a previously applied set and a new set of spatialization filter parameter.

20. A method performed by server-side computing device having network interface, the method comprising: receiving, via network interface from client-side device, request packet comprising a user prompt and environmental data; processing said user prompt to generate responsive audio content track; generating one or more reference sound signal configured for analysis by Object Recognition Al Application on client-side device; transmitting, via network interface to said client-side device, response packet comprising said responsive audio content track and said one or more reference sound signal; and participating in bi-directional time synchronization process with client-side device.21 . The method of claim 20, further comprising processing said environmental data to generate context-aware ambient audio track, wherein said response packet further comprises said context-aware ambient audio track.

22. The method of claim 20, wherein said one or more reference sound signal are inaudible.

23. The method of claim 20, further comprising initiating local playback of said responsive audio content track and continuous emission of said one or more reference sound signal, wherein said local playback is internally synchronized with timing information dispatched to client-side device.

24. Server-side computing system, comprising: network interface; one or more processor communicatively coupled to network interface; and memory storing instruction that, when executed by said one or more processor, cause the system to perform the method of claim 20.

25. The system of claim 24, wherein the method further comprises processing environmental data received from client-side device to generate context-aware ambient audio track.

26. The system of claim 24, wherein said one or more reference sound signal are inaudible.

27. Non-transitory computer-readable storage medium having instruction stored thereon that, when executed by one or more processor of server-side device, cause said server-side device to perform the method of claim 20.

28. Non-transitory computer-readable storage medium of claim 27, wherein the method further comprises generating context-aware ambient audio track based on environmental data received from client-side device.

29. A system for generating responsive, spatialized audio, comprising: server-side device; client-side device communicatively coupled to server-side device, wherein clientside device includes one or more camera; wherein client-side device is configured to transmit continuous audio stream capturing a verbal command of a user and associated environmental data to server-side device; and wherein server-side device is configured to receive said continuous audio stream and said environmental data, and to process said audio stream and said environmental data using a single, natively Multimodal Al model to directly generate a responsive audio content track without an intermediate text transcription step.

30. The system of claim 29, wherein server-side device is further configured to generate one or more reference sound signal for transmission with responsive audio content track.

31. The system of claim 30, wherein client-side device is further configured to determine a location of server-side device by adaptively fusing acoustic data derived from said one or more reference sound signal with visual data from said one or more camera.

32. The system of claim 29, wherein server-side device is further configured to generate context-aware ambient audio track based on said environmental data.

33. A method for generating responsive, spatialized audio via client-server system, comprising: transmitting, by client-side device comprising one or more camera, continuous audio stream capturing a verbal command of a user and associated environmental data to server-side device; and receiving, by server-side device, said continuous audio stream and said environmental data, and processing said audio stream and said environmental data using a single, natively Multimodal Al model to directly generate a responsive audio content track without an intermediate text transcription step.

34. The method of claim 33, further comprising: generating, by server-side device, one or more reference sound signal; and transmitting, by server-side device, said one or more reference sound signal to client-side device.

35. The method of claim 34, further comprising: receiving, by client-side device, said one or more reference sound signal; and determining, by client-side device, a location of server-side device by adaptively fusing acoustic data derived from said one or more reference sound signal with visual data from said one or more camera.

36. The method of claim 33, further comprising generating, by server-side device, context-aware ambient audio track based on said environmental data.

37. Non-transitory computer-readable storage medium having instruction stored thereon that, when executed by one or more processor of server-side device, cause said server-side device to perform the method of claim 33.

Citation Information

Patent Citations

  • Audio system for spatializing virtual sound sources

    CN117981347A

  • System and method for user profile enabled smart building control

    US20170055126A1

  • System and methods for audio pattern recognition

    WO2018226359A1