Multimodal intelligent agent system for dynamic environmental monitoring and human-centered support

The multimodal intelligent agent system addresses the limitations of existing environmental monitoring systems by integrating diverse data streams for context-aware monitoring and personalized support, achieving real-time, adaptive, and scalable task execution.

DE202025106637U1Active Publication Date: 2026-02-26GOUNDER MOHAN SELLAPPA DR BENGALURU +3
View PDF 0 Cites 13 Cited by

Patent Information

Application Number
DE202025106637
Authority / Receiving Office
DE · DE
Patent Type
Utility models
Current Assignee / Owner
Filing Date
2025-11-02
Publication Date
2026-02-26
Estimated Expiration
2035-11-30

AI Technical Summary

Technical Problem

Existing environmental monitoring systems lack a unified, multimodal processing framework capable of integrating diverse data streams—audio, video, text, and sensor data—into a coherent cognitive model, failing to achieve contextual understanding and personalized support due to static data processing chains and the absence of dynamic agent-based architectures.

Method used

A hardware-integrated, multimodal intelligent agent system that employs transformer-based large language models for real-time semantic fusion, adaptive agent construction, and personalized output generation, enabling simultaneous analysis of various data modalities and dynamic task-specific agent creation.

Benefits of technology

Enables real-time, context-aware monitoring and support by dynamically constructing agents for tasks like meeting transcription, behavioral analysis, or object recognition, while ensuring data privacy and operational reliability through network-edge processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A multimodal intelligent agent system for dynamic environmental monitoring and user-centered support, consisting of: a multimodal sensor module configured to continuously acquire environmental and behavioral data from multiple input modalities, including at least one visual sensor, at least one acoustic sensor, at least one environmental conditions sensor, and at least one proximity or motion detection sensor, each generating modality-specific data streams representing visual images, audio waveforms, physical environmental parameters, and motion signatures within a monitored environment; a data preprocessing and fusion subsystem that is operationally coupled with the multimodal sensor module and configured to normalize, temporally align, and transform the modality-specific data streams into high-dimensional feature embeddings using a variety of encoders, wherein the visual encoder uses convolutional or vision transformer architectures, the audio encoder uses a spectral-temporal feature extractor, and the sensor encoder transforms raw analog data into context vectors suitable for multimodal alignment; a multimodal processing unit consisting of a transformer-based large language model (LLM) trained on paired multimodal datasets and configured to perform semantic fusion, context abstraction, and inference across the aforementioned aligned multimodal feature embeddings to generate a contextual understanding of environmental and behavioral states; an adaptive agent controller coupled to the multimodal inference processing unit and configured to instantiate, manage, and terminate a variety of task-specific intelligent agents, each agent being a software unit configured to perform a specialized function selected from meeting summarization, behavioral analysis, misplaced object detection, or environmental anomaly identification, with the agents dynamically interacting with the inference engine to retrieve contextually relevant multimodal embeddings for task execution; a personalization and adaptive learning subsystem consisting of a user preference database and a neural memory structure configured to update and refine model parameters based on user-specific interaction history, thereby enabling personalized output generation, prioritization of recommendations, and long-term behavioral adaptation; and An output generation interface is operationally connected to the adaptive agent controller and configured to produce multimodal output in textual, visual, and auditory form. The interface is capable of displaying human-readable summaries, notifications, and visual reconstructions of identified entities or environmental states.
Need to check novelty before this filing date? Find Prior Art

Description

Technical field of the invention

[0001] The present invention relates to the field of artificial intelligence and intelligent environmental monitoring systems. In particular, it relates to an integrated, multimodal, intelligent, agent-based machine for the dynamic monitoring of indoor spaces and for user-centered support. The invention utilizes multimodal, large language models (LLMs) for processing audio, video, text, and sensor data for context-aware understanding, adaptive automation, and personalized output generation. BACKGROUND OF THE INVENTION

[0002] Conventional environmental monitoring systems rely primarily on unimodal data sources such as cameras, acoustic sensors, or motion detectors to perform local analyses like presence detection or simple event logging. These traditional systems are unable to interpret intermodal relationships or derive higher-level semantic meanings from heterogeneous data streams. Consequently, existing technologies are limited to generating static reports or triggering predefined alarms without understanding the contextual dependencies between human activity, environmental conditions, and object behavior.

[0003] Recent advances in artificial intelligence, particularly the development of multimodal deep learning models and large language models (LLMs), have revealed significant potential for cross-modal understanding and reasoning. Frameworks such as CLIP, Flamingo, and GPT-4V have demonstrated their ability to jointly interpret visual and textual data. However, most systems based on these architectures remain limited to single applications such as meeting transcription or video subtitling, without achieving an integrated, adaptive monitoring framework that enables autonomous decision-making and personalized support.

[0004] Furthermore, while standalone virtual assistants like Alexa, Siri, or Google Assistant offer basic context-aware interaction, they are unable to dynamically create specialized submodules or "agents" capable of independently analyzing and summarizing complex multimodal scenarios. Thus, a crucial technological gap remains in the development of a unified architecture that integrates multimodal sensing, dynamic agent creation, context-aware reasoning, and personalized adaptive output. The present invention closes this gap with a hardware-integrated, AI-driven, multimodal agent system that autonomously monitors environmental and behavioral parameters in enclosed spaces such as offices, meeting rooms, and classrooms.

[0005] The development of intelligent systems for environmental monitoring and human-machine assistance has been significantly shaped by advances in artificial intelligence, sensor networks, computer vision, and natural language processing. Over the past two decades, researchers and companies have developed numerous frameworks for automated monitoring, workplace optimization, and digital assistance. While these systems are functional for individual tasks, they remain largely unimodal, fragmented, and context-insensitive. Traditional approaches are designed to process single modalities—such as video streams, audio streams, or sensor readings—without the ability to integrate or interpret different data types. Consequently, their results are limited to narrow interpretations and lack the contextual understanding necessary for meaningful, user-centered interaction.This discrepancy between sensory input and cognitive interpretation represents a persistent technical challenge, which the present invention directly addresses.

[0006] Conventional environmental monitoring technologies began with simple sensor-based systems in smart buildings and industrial automation. These systems typically used temperature sensors, motion detectors, and light sensors to control heating, ventilation, air conditioning (HVAC), and lighting. The control logic was rule-based, relying on preset thresholds and static decision trees. While these systems effectively controlled physical parameters, they could not adapt to dynamic human behavior, changes in context, or multimodal signals. For example, a conventional sensor-based office automation system could detect room occupancy but could not distinguish between a meeting, the presence of a single user, or a security breach.Furthermore, such systems are limited by the lack of semantic inference – that is, the ability to interpret the meaning of a captured signal – and therefore do not support adaptive or personalized actions.

[0007] The introduction of video surveillance and computer-aided image processing expanded the potential of environmental monitoring by enabling the detection of objects, movements, and events from visual data. Early video analytics utilized manually developed feature extraction methods such as Scale-Invariant Feature Transform (SIFT) and Histogram of Oriented Gradients (HOG), followed by machine learning techniques. However, these approaches were limited by their sensitivity to lighting conditions, occlusion, and camera angles. Later, deep learning-based models such as Convolutional Neural Networks (CNNs) improved the accuracy of visual recognition and enabled applications such as facial recognition, behavioral analysis, and object tracking.However, these systems remained unimodal – they worked exclusively with image or video data and could not correlate visual information with other sensory or contextual cues such as speech, temperature, or user intent. Furthermore, video-based surveillance systems often raised privacy concerns, required significant computing resources, and were difficult to interpret due to their "black box" nature.

[0008] Parallel to these developments, acoustic and speech-based systems emerged for environmental monitoring and human-machine interaction. Intelligent assistants such as Amazon Alexa, Google Assistant, and Apple Siri demonstrated the capabilities of natural language processing (NLP) and automatic speech recognition (ASR). These systems enabled users to interact via voice command, retrieve information, and control networked devices. Despite their success, these assistants have significant limitations: They are heavily reliant on predefined command structures, operate within a limited scope, and cannot interpret multimodal input. For example, they cannot integrate auditory signals with visual or sensor data to make context-aware decisions.Furthermore, their dependence on cloud computing leads to latency, data protection gaps, and limited adaptability in offline or high-security environments.

[0009] In another area, IoT-based smart environment systems combine distributed sensors, actuators, and connectivity frameworks to enable large-scale monitoring and control. Deployed in smart homes, healthcare facilities, and industrial plants, these systems leverage edge and cloud computing for real-time processing of sensor data. While IoT architectures have improved scalability and automation, they remain limited by their data representation capabilities. Each sensor modality generates data in different formats—analog signals, digital data streams, or textual metadata—which are rarely harmonized effectively. Consequently, IoT systems often deliver fragmented analyses that either oversimplify or ignore correlations between human behavior, environmental dynamics, and spatial context.Without semantic fusion of multimodal data, these systems cannot provide actionable insights beyond threshold-based warnings.

[0010] Efforts have been made to develop integrated frameworks that combine visual and textual information, primarily through multimodal machine learning. Early research in image description and visual question answering (VQA) demonstrated that deep learning models can process text and image data together. The emergence of transformer-based models, particularly the Vision Transformer (ViT) and multimodal variants such as CLIP and Flamingo, revolutionized the field by enabling crossmodal embedding and attention-based fusion. These models achieved impressive results in tasks such as image search, image description generation, and visual reasoning. However, these architectures were primarily designed for static datasets and research benchmarks. They lack the real-time adaptability, personalization, and dynamic agent construction required for real-world applications in environmental monitoring.Furthermore, current multimodal frameworks often depend on extensive cloud-based computing infrastructures, which are not practical for use at the network edge in sensitive or resource-constrained environments such as offices, hospitals, or defense facilities.

[0011] Existing AI-based systems for transcribing and summarizing meetings, such as Otter.ai, Fireflies, and Sembly AI, represent a different category of automation. These tools capture audio, convert speech to text, and summarize conversations using text-based NLP models. While they simplify documentation and increase productivity, their functionality is limited to a single modality (audio and text). Contextual cues such as speaker behavior, visual gestures, environmental conditions, or user-specific relevance are not considered. Consequently, the summaries generated are generic and lack context. Furthermore, these systems cannot dynamically adapt to new use cases or perform more complex inference tasks, such as linking meeting content to real-world actions or changes in the environment.The lack of multimodal integration limits its use to the retrospective documentation of events instead of active, context-related support.

[0012] Another related field encompasses behavioral and anomaly detection systems that rely on sensor fusion or image analysis to infer human activity. These systems are commonly used in health monitoring, security, and occupational safety applications. Traditional approaches employ statistical models such as Hidden Markov Models (HMMs) or Conditional Random Fields (CRFs) to represent temporal activity sequences. More recent deep-learning-based approaches, such as Long Short-Term Memory (LSTM) networks and Temporal Convolutional Networks (TCNs), have improved temporal modeling. However, these systems remain task-specific and require large, annotated datasets for training. They cannot generalize to different contexts, individuals, or environments.Furthermore, most behavioral detection systems do not integrate semantic reasoning or personalized adjustments, resulting in limited interpretability and frequent false alarms.

[0013] Despite these advances, all existing technologies still suffer from a crucial limitation: the lack of a unified, multimodal processing framework capable of integrating diverse data streams—audio, video, text, and sensor data—into a coherent cognitive model. Human environments are inherently multimodal; actions, language, and environmental changes occur simultaneously and influence one another. A truly intelligent monitoring system must therefore process these modalities together, not in isolation. Current architectures fail to achieve this because they rely on static data processing chains that treat each data source independently. The lack of semantic fusion leads to disjointed results, rendering systems incapable of understanding complex real-world interactions or providing personalized support based on contextual information.

[0014] Furthermore, existing solutions rarely implement dynamic, agent-based architectures. Most AI systems operate as monolithic units that execute a fixed sequence of tasks. This rigidity prevents them from adapting to changing conditions or generating specialized task agents as needed. In contrast, human intelligence is based on modular cognitive structures capable of dynamically activating specialized thought patterns for different situations. Replicating this ability in artificial systems remains an unresolved challenge. Without the dynamic creation and management of agents, multimodal AI systems cannot efficiently scale to multiple simultaneous tasks, such as tracking objects, summarizing meetings, and simultaneously analyzing human behavior in the same environment.

[0015] Another significant drawback of current systems is the lack of personalization and context integration. Most commercial AI solutions generate results solely based on current data streams, without considering historical patterns, user preferences, or previous interactions. This leads to repetitive, impersonal results that are irrelevant to specific users or contexts. For example, a meeting summary system might regularly highlight procedural details but fail to identify the discussion points relevant to each individual participant. Personalized adaptation requires a persistent learning layer that manages user models and continuously updates them based on interaction history—a capability that existing frameworks lack.

[0016] Technical limitations such as computing inefficiency, network dependency, and data privacy concerns hinder the deployment of multimodal AI systems. Many advanced models utilize cloud-based servers for training and inference, which leads to latency and data privacy risks when processing sensitive data such as audio or video recordings. Edge-based systems, while improving data privacy, often reach their limits due to hardware constraints when performing large-scale multimodal calculations. This trade-off between computing power and data security remains unresolved in most current architectures. Furthermore, the energy consumption and bandwidth requirements increase when processing continuous, large multimodal data streams, making such systems impractical for real-time monitoring over extended periods.

[0017] The accumulation of these limitations has led to an unmet need for a unified, multimodal, agent-based intelligent system capable of contextual reasoning, adaptive action, and real-time personalized output for human users. The present invention closes this technological gap by integrating multimodal, large-scale language models with agent-based automation, thereby enabling real-time semantic fusion, contextual understanding, and personalized response generation. Unlike existing unimodal or static multimodal systems, the proposed invention dynamically constructs specialized agents for different tasks—such as meeting transcription, behavioral analysis, or object recognition—while maintaining a common cognitive core for reasoning and adaptation.This approach not only overcomes the isolated structure of the state of the art, but also establishes a new paradigm for intelligent human-environment interaction through unified multimodal cognition and adaptive automation. OBJECTS OF INVENTION

[0018] The main objective of the present invention is to provide a multimodal intelligent system capable of fusing and analyzing real-time data from audio, video, sensor and text inputs to enable dynamic monitoring and personalized, user-centric support.

[0019] Another objective of the invention is the development of a device that uses multimodal LLMs for context-related real-time reasoning and the automated generation of summaries, meeting minutes, behavioral insights, and the identification of misplaced items.

[0020] Another objective of the invention is to enable the construction of adaptive agents through computing modules that dynamically instantiate task-specific intelligent agents based on recognized context or environmental conditions.

[0021] Another objective of the invention is to provide a robust, network-edge machine architecture that can function independently of cloud-based calculations while ensuring data protection, low latency and high operational reliability. SUMMARY OF THE INVENTION

[0022] The invention describes an integrated, multimodal intelligent agent system for dynamic environmental monitoring and user-centered support. The system comprises a hardware-integrated device structure with multimodal sensors, a central processing unit with GPU acceleration, a memory subsystem, and adaptive logic. The multimodal sensors acquire diverse data, including images, audio recordings, text information, and measurements from environmental sensors (temperature, humidity, motion, CO2, etc.). The acquired data is preprocessed and fused multimodally using a transformer-based encoder-decoder network. The fused representation is processed by a multimodal language model, which forms the cognitive core of the system.

[0023] The logic engine utilizes an adaptive agent construction technique that dynamically generates and manages subagents based on detected events or tasks such as meeting analysis, behavior pattern recognition, or object localization. The agents operate in parallel, each calling context-specific logic chains via the LLM to generate structured outputs. A personalization layer integrates user history, preferences, and previous interactions to tailor system responses. Outputs are displayed via a multimodal interface with textual, visual, and auditory feedback. This enables comprehensive, context-aware automation for intelligent indoor monitoring and assistance.

[0024] The present invention aims to provide an integrated, multimodal intelligent system capable of comprehensively monitoring, interpreting, and supporting human activities in dynamic indoor environments. This is achieved through the simultaneous analysis of various data modalities, including audio, video, text, and sensor data. The invention overcomes the limitations of existing unimodal or fragmented systems by employing a unified framework that combines environmental perception, semantic understanding, and contextual reasoning. This enables the system to deliver precise, personalized, real-time outputs. Furthermore, the invention aims to develop an advanced cognitive machine architecture capable of autonomously processing diverse environmental data streams, understanding human behavior, and providing intelligent summaries or recommendations for action based on observed contextual information.

[0025] Another important objective of the invention is the introduction of a novel mechanism for dynamic agent construction. This enables the system to autonomously generate, manage, and execute specialized, task-specific agents to respond to environmental changes or user-defined objectives. This allows the system to remain adaptable and scalable, and to simultaneously handle various tasks within the same operating framework, such as meeting logging, behavioral analysis, identification of misplaced objects, and detection of environmental anomalies. The invention thus aims to replicate human-like flexibility in artificial intelligence systems by introducing modular, self-instantiating agents that collaborate under a unified, multimodal core.

[0026] A further objective of the invention is seamless multimodal data fusion and processing through the integration of transformer-based large language models (LLMs) that can align and interpret relationships across heterogeneous data formats. This mechanism enables robust semantic understanding and allows the system to infer intentions, activities, and environmental states from combined visual, auditory, textual, and sensory cues. By leveraging the representational power of multimodal embeddings, the invention improves the accuracy of decision-making and enables the generation of coherent, context-aware descriptions and insights that correspond to human understanding.

[0027] A further objective of the invention is the development of a hardware-integrated device structure that embodies the multimodal intelligent agent framework in a compact, modular machine design. The invention aims to provide a self-contained system capable of performing computations directly on the device without relying on remote cloud infrastructures. This ensures data privacy, operational reliability, and real-time capability. The physical embodiment of the invention comprises multimodal sensors, computing units, and user interfaces in an ergonomic design that allows for discreet deployment in various environments such as offices, classrooms, hospitals, and industrial facilities.

[0028] A further objective of the invention is to provide personalized and adaptive user interaction through the integration of a continuous learning mechanism into the system architecture. The invention aims to enable the intelligent agents to learn incrementally from user preferences, interaction history, and behavioral patterns, thereby refining their responses and increasing the relevance of the generated outputs over time. This level of personalization allows the system to adapt meeting summaries, behavioral analyses, and environmental feedback to the individual needs, communication styles, and contextual priorities of the users, thus creating a highly user-centered interaction paradigm.

[0029] Another important objective of the invention is to increase operational efficiency, automation, and situational awareness in monitored environments. By combining multimodal data processing and adaptive inference, the invention aims to automate complex analytical tasks that traditionally require human intervention, such as identifying key discussion points in meetings, detecting misplaced objects, or recognizing behavioral anomalies. The automation achieved by the system reduces cognitive load, minimizes human error, and significantly improves company productivity by providing precise and actionable information in real time.

[0030] A further objective of the invention is to ensure reliability and fault tolerance in environmental monitoring through redundant multimodal data processing. The invention is designed so that functionality is maintained even in the event of partial data loss or limited sensor performance. For example, if no image signal is available due to poor lighting conditions, the system can rely on audio and motion sensor data to maintain environmental sensing and decision accuracy. This redundancy ensures robustness and reliability, thus making the system suitable for critical applications in healthcare, security, and industry.

[0031] The invention aims to promote inclusion and accessibility by enabling multimodal interaction modes that cater to the diverse abilities and preferences of users. Through text, speech, and visual outputs, the system ensures that people with varying sensory or physical abilities can interact effectively with the device. By interpreting environmental and behavioral signals in different modalities, the system also facilitates communication for people with disabilities, for example, through text-based transcriptions or visual activity summaries in environments where verbal communication is difficult.

[0032] A further objective of the invention is the development of a scalable and interoperable system architecture that can be seamlessly integrated into existing digital ecosystems such as enterprise software, building automation systems, and data analytics platforms. The invention aims to provide secure communication interfaces and APIs that enable interoperability without compromising data security. This allows the proposed system to function as a central intelligent control unit within broader organizational infrastructures and to synchronize with complementary systems to provide coherent and data-driven operational intelligence.

[0033] A further objective of the invention is to promote a sustainable and energy-efficient smart infrastructure through intelligent computational planning, energy-conscious sensor activation, and adaptive processing strategies. The invention minimizes redundant calculations through dynamic resource allocation based on the relevance of ongoing activities. For example, during periods of low environmental change or inactivity, the system can reduce the measurement frequency and evaluation cycles, thus saving energy without impairing situational awareness. This adaptive energy management ensures efficient and environmentally friendly operation of the system while simultaneously providing continuous intelligent monitoring.

[0034] The overarching goal of the invention is to advance artificial intelligence towards human-like multimodal cognition. This will enable machines to understand, interpret, and act upon complex environmental and social contexts. By integrating multimodal sensors, reasoning with large language models, adaptive agent development, and personalization into a single device, the invention establishes a new paradigm for intelligent monitoring and human-machine collaboration. It not only addresses existing technological limitations but also creates a scalable foundation for the next generation of cognitive automation systems. These systems are capable of intelligently and contextually perceiving, interpreting, and interacting with their environment. BRIEF DESCRIPTION OF THE IMAGE

[0035] These and other features, aspects and advantages of the present invention will be better understood if the following detailed description is read with reference to the accompanying drawing, in which the same symbols represent the same parts: Fig. Figure 1 shows a block diagram of an integrated multimodal intelligent agent system for dynamic environmental monitoring and user-centric support.

[0036] Furthermore, those skilled in the art will recognize that the elements in the drawing are simplified and not necessarily drawn to scale. For example, the flowcharts illustrate the process by highlighting the main steps to facilitate understanding of the present disclosure. With regard to the construction of the device, one or more components may be represented in the drawing by conventional symbols. The drawing may show only those specific details relevant to understanding the embodiments of the present disclosure, so as not to clutter the drawing with details that are already apparent to those skilled in the art from the description contained herein. Detailed description of the invention

[0037] To facilitate understanding of the principles of the invention, reference is made below to the embodiment shown in the drawing, which is described using specific terms. It is understood, however, that this does not limit the scope of protection of the invention. Rather, modifications and further developments of the depicted system, as well as further applications of the inventive principles shown therein, are conceivable, insofar as they would normally occur to a person skilled in the art in the field of the invention.

[0038] It will be clear to those skilled in the art that the foregoing general description and the following detailed description are exemplary and explanatory of the invention and are not to be understood as a limitation of it.

[0039] References to “an aspect”, “another aspect”, or similar phrases in this description mean that a particular feature, structure, or property described in connection with the embodiment is included in at least one embodiment of the present disclosure. Therefore, phrases such as “in one embodiment”, “in another embodiment”, and similar expressions in this description may, but do not necessarily, all refer to the same embodiment.

[0040] The terms "includes," "comprehensive," or similar expressions denote non-exclusive inclusion. Thus, a procedure or method containing a list of steps does not only include those steps but may also include further steps not explicitly listed or inherent in the procedure or method. Likewise, the statement "includes..." for one or more devices, subsystems, elements, structures, or components, without further limitations, does not preclude the existence of other devices, subsystems, elements, structures, or components.

[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meanings generally known to those skilled in the art in the field to which this invention belongs. The systems, methods, and examples described herein serve only for illustration and are not to be understood as limiting.

[0042] Embodiments of the present disclosure are described in detail below with reference to the attached drawing.

[0043] Fig.Figure 1 shows a block diagram of an integrated multimodal intelligent agent system for dynamic environmental monitoring and user-centric support. The system 100 comprises: a multimodal sensor module (102) that continuously acquires environmental and behavioral data from multiple input modalities, including at least one visual sensor, at least one acoustic sensor, at least one environmental condition sensor, and at least one proximity or motion sensor. Each of these modalities generates modality-specific data streams representing visual images, audio signals, physical environmental parameters, and motion signatures within a monitored environment. A data preprocessing and fusion system (104), coupled with the multimodal sensor module, normalizes,temporally aligns and transforms the modality-specific data streams into high-dimensional feature embeddings using multiple encoders. The visual encoder uses convolutional or vision transformer architectures, the audio encoder a spectral-temporal feature extractor, and the sensor encoder transforms raw analog data into context vectors suitable for multimodal alignment. A multimodal processing unit (106) with a transformer-based large language model (LLM) trained on paired multimodal datasets and configured for semantic fusion, context abstraction, and inference across the aligned multimodal feature embeddings to generate a contextual understanding of environmental and behavioral states; an adaptive agent controller (108) coupled to the multimodal processing unit and responsible for instantiating,Management and termination of a multitude of task-specific intelligent agents, each agent being a software unit configured to perform a specialized function selected from meeting summarization, behavioral analysis, misplaced object detection, or environmental anomaly identification, with the agents dynamically interacting with the processing unit to retrieve contextually relevant multimodal embeddings for task execution; a personalization and adaptive learning subsystem (110) with a user preference database and a neural memory structure configured to update and refine model parameters based on the user's interaction history, thereby enabling personalized output generation, recommendation prioritization, and long-term behavioral adaptation; and an output generation interface (112),which is operationally connected to the adaptive agent controller and is configured to generate multimodal outputs in textual, visual, and auditory form, the said interface being capable of displaying human-readable summaries, notifications, and visual reconstructions of identified entities or environmental states.

[0044] In one embodiment, the multimodal sensor module (102) comprises a panoramic HD camera mounted on a motorized gimbal configured for three-dimensional rotation. The gimbal is controlled by a micro-servo array that dynamically adjusts the field of view in response to acoustic localization data acquired by the directional microphone array. This enables adaptive visual tracking of active human participants or objects of interest in the monitored environment.

[0045] In one embodiment, the data preprocessing and fusion subsystem (104) further comprises a multimodal synchronization buffer and a transformer-based fusion encoder configured to temporally align asynchronous data streams using timestamp correlation and learned attention weights, so that visual frames, audio transcripts, and environmental sensor data are aligned to represent temporally coherent events, enabling accurate derivation of cause-and-effect relationships across modalities during real-time monitoring.

[0046] In one embodiment, the multimodal processing unit (106) employs a two-stage architecture comprising (i) a cross-modal alignment transformer that maps feature embeddings from visual, auditory, and sensory modalities into a unified latent representation, and (ii) a semantic inference transformer trained by supervised contrastive learning on multimodal instructional datasets to perform context-sensitive reasoning, thereby enabling the engine to generate natural language descriptions, event summaries, and predictive inferences based on environmental dynamics.

[0047] In one embodiment, the multimodal processing unit (106) additionally integrates a submodule for causal reasoning, which is configured to determine directional relationships between multimodal variables, such as the assignment of a detected motion vector to a simultaneously occurring acoustic signature or the identification of anomalous environmental fluctuations as the cause of behavioral changes, thereby improving interpretability and reducing false positive conclusions under complex environmental conditions.

[0048] In one embodiment, the adaptive agent controller (108) dynamically generates a plurality of computational agents, each instantiated in a containerized runtime environment that includes memory isolation and API access controls. These agents are configured to subscribe to data streams via an internal event broker, autonomously determine task relevance based on attentional evaluation thresholds, and request multimodal embeddings from the inference engine to perform specialized operations, including, but not limited to, semantic summarization, behavioral tagging, or anomaly evaluation.

[0049] In one embodiment, at least one meeting summary agent is configured to receive transcribed audio and corresponding visual attention maps of the participants and applies a finely tuned large-scale speech model to generate structured meeting minutes that include speaker-related action items, temporal annotations, and contextual highlights. The agent also validates the generated text using nonverbal cues such as gestural emphasis or intonation variations, which are detected by the multimodal processing unit.

[0050] In one embodiment, at least one object detection agent is configured to analyze fused visual and proximity sensor data to track static and dynamic objects, generate a reference inventory of expected positions, and detect misplaced or distant objects by comparing live embeddings with historical scene graphs stored in the personalization subsystem. The agent provides a visual reconstruction of the anomalous spatial area for user review via the output interface.

[0051] In one embodiment, a behavioral analysis agent is configured to derive human activity sequences from fused visual and motion sensor data, segment behavioral episodes using a temporal activity transformer, and generate natural language descriptions of behavioral routines or deviations therefrom, with the agent further computing a behavioral consistency index over time to detect irregularities that may indicate fatigue, stress, or anomalous events.

[0052] In one embodiment, the personalization and adaptive learning subsystem (110) comprises a recurrent memory encoder configured to manage a user-specific latent state vector that captures previous environmental conditions, user preferences, and task interactions, and wherein the subsystem uses Reinforcement Learning with Human Feedback (RLHF) to update weighting parameters of the multimodal processing unit, thereby achieving incremental adaptation of responses without retraining the core model.

[0053] The integrated multimodal intelligent agent system for dynamic environmental monitoring and user-centered support operates with a series of computationally coordinated modules that together form a unified cognitive framework. This framework is capable of perceiving, reasoning, and acting in indoor environments. The system architecture, as described in the preceding claims, functions as a self-contained intelligent machine that continuously acquires multimodal sensor data, fuses this input using advanced transformer-based encoders, makes high-level semantic inferences using a multimodal language model, and dynamically creates specialized agents to perform specific cognitive tasks such as meeting summarization, behavior recognition, or object localization.The overall operation of the system is orchestrated by an adaptive agent controller that manages the agents' lifecycle events, while a personalization subsystem ensures context continuity and user-specific customization through memory-based learning.

[0054] On a technical level, the process begins with the multimodal sensor module, which continuously acquires heterogeneous data from various sources. High-resolution cameras provide image data, directional microphones capture acoustic signals, and environmental sensors measure variables such as temperature, humidity, air quality, and motion. Each data stream is time-stamped and transferred to the data preprocessing and fusion system. The process starts with a multimodal synchronization routine that uses a temporal alignment mechanism based on cross-modality timestamp correlation and dynamic interpolation. This ensures that all modalities are temporally coherent and correspond to the same context window. After synchronization, each modality passes through its respective feature extraction pipeline.For visual data, a Convolutional Vision Transformer (ViT) encoder converts pixel-level features into a sequence of latent embeddings, preserving spatial hierarchies. The audio stream is processed by a Log-Mel spectrogram generator and subsequently by a Temporal Convolutional Network (TCN) or an attention-based speech encoder that captures phonetic and prosodic features. Environmental and motion sensors generate structured time-series data, which are normalized and transformed into context vectors using recurrent embedding networks, capturing short- and long-term temporal dependencies.

[0055] Following feature extraction, the multimodal embeddings are transferred to the fusion encoder, which forms the core of the data fusion technique. The fusion encoder is based on a cross-modal transformer framework, in which each modality is represented as a token sequence in a shared attentional space. The model learns cross-attention weights, enabling the system to recognize correlations between modalities—for example, the synchronization of an acoustic event (speech) with a visual feature (lip movement) or the association of environmental changes (movement or temperature change) with corresponding visual cues. The attentional mechanism calculates importance values ​​at the modality level, which dynamically adjust based on signal strength and contextual relevance.For example, if a visual modality is impaired due to poor lighting, the system reduces the weighting of the visual embedding and gives greater weight to audio or sensor data. The fusion process generates a unified, high-dimensional representation vector, the so-called multimodal context embedding, which serves as cognitive input for the multimodal processing unit.

[0056] The multimodal processing unit utilizes a transformer-based large language model (LLM) trained on paired multimodal datasets comprising image-text, audio-text, and sensor-text pairs. The processing unit operates in a two-stage process: cross-modal alignment and semantic inference. In the alignment stage, the fusion encoder's embeddings are projected into a shared latent space, where semantic similarity between modalities is maximized through contrastive learning objectives. This allows the model to establish latent correspondences, such as linking spoken phrases with visual gestures or associating spatial anomalies with textual scene descriptions. In the semantic inference stage, the model processes the unified contextual embedding using multiple self-attention layers that capture intermodal dependencies and generate contextualized representations.These representations are decoded into high-level semantic outputs, such as textual descriptions of environmental states, summaries of ongoing discussions, or predictions of user intent. The output layer of the reasoning engine generates either structured data (JSON objects for system use) or natural language output (summaries, notifications, or reports) for human interpretation.

[0057] A key feature of the invention is its adaptive agent controller. This acts as a dynamic orchestration layer that autonomously creates and manages specialized agents in real time based on context triggers. The controller continuously monitors the inference outputs of the reasoning engine and evaluates the contextual probability distribution of detected events. If an event exceeds a predefined relevance threshold—for example, the detection of speech activity, anomalous movement, or environmental deviation—the controller instantiates a corresponding agent in a containerized runtime environment. Each agent operates as an independent processing unit with isolated memory and API access rights. The controller's scheduling technique allocates computing resources to the agents based on real-time utilization, energy efficiency metrics, and user priorities.After instantiation, the agents communicate with the reasoning engine via a shared attention bus. This allows them to request relevant embeddings and execute domain-specific reasoning tasks.

[0058] The meeting summary agent, for example, uses a finely tuned transformer-decoder model trained on conversation datasets. It receives a transcribed version of the meeting audio from the reasoning engine's automatic speech recognition (ASR) module. This is synchronized with participant-specific visual embeddings to identify active speakers. The agent segments the conversation into dialogue contributions, applies context window attention, and identifies key discussion points, decisions, and areas for action. Using feedback learning based on reinforcement effects, it verifies the generated summaries against nonverbal cues such as gesture intensity or intonation to improve the accuracy of intonation detection.The final structured summary is then formatted through a post-processing pipeline that includes entity recognition, timestamping, and role-based tagging before being transferred to the output interface.

[0059] The object detection agent uses an image-sensor fusion technique to create a continuously updated spatial map of the monitored environment. This technique generates a scene graph where each node represents a detected object and each edge represents spatial relationships or proximity data from the sensors. During operation, the agent performs scene comparisons by comparing the current scene graph with stored baseline configurations of the personalization system. If deviations are detected—for example, if an object is missing from its intended location—the agent uses its analysis engine to generate a natural language description and a spatial visualization of the anomaly.The object tracking technique uses a Kalman filter in combination with a re-identification module based on a Siamese network to ensure object continuity across multiple frames, thus guaranteeing robust detection even in the case of partial occlusion or motion blur.

[0060] The behavioral analysis agent uses an activity detection technique based on temporal transformers and recurrent neural networks to analyze fused motion, visual, and auditory embeddings. The model segments continuous activity sequences into labeled behavioral episodes such as "standing," "sitting," "collaborating," or "presenting." Each episode is evaluated based on its confidence level, and deviations from previously learned behavioral patterns are flagged as anomalies. The agent calculates a behavioral consistency index (BCI) over a sliding time window by comparing observed activities with expected routines stored in the personalization memory. Significant deviations trigger alerts or adaptive suggestions—such as environmental adjustments or reminders—to enhance well-being and productivity.

[0061] The personalization and adaptive learning subsystem forms the system's long-term memory and cognitive adjustment layer. It maintains a persistent user model, encoded as a latent state vector, which is updated after each interaction using reinforcement learning with human feedback (RLHF). The learning process optimizes a reward function that balances task accuracy, user satisfaction, and energy efficiency. Each update modifies the attentional allocation of the inference engine for subsequent inferences, allowing the system to adapt to individual user preferences, communication styles, and behavioral patterns. To protect privacy, this learning takes place locally on the device.In installations with multiple devices, parameter updates are aggregated via a federated learning mechanism that combines encrypted gradient updates from multiple devices without exchanging raw data. This ensures consistent global improvement while maintaining local data security.

[0062] The output interface translates the conclusions and agent outputs into understandable, multimodal feedback. It consists of a graphical dashboard, an audio synthesis unit, and an optional augmented reality (AR) overlay. The rendering technique maps textual or structured outputs onto visual representations and displays detected events, summarized discussions, and behavioral insights as annotated overlays on live camera images. For the auditory feedback, a neural text-to-speech engine converts generated summaries or alerts into natural-sounding speech tailored to user preferences. The system uses adaptive latency compensation to ensure that the multimodal outputs remain synchronized with environmental events, thus enabling a seamless interactive experience.

[0063] The entire system is based on a self-monitored, continuous learning process that ensures sustained model accuracy and adaptability. It monitors inference confidence levels, error rates, and modality drift statistics to identify when retraining is needed. During periods of low utilization, the controller initiates fine-tuning cycles directly on the device using self-labeled data collected during operation. This allows the system to adapt to changing environmental conditions or evolving user behavior without requiring external datasets or manual monitoring. Over time, the system continuously refines its reasoning capabilities, ensuring that its contextual understanding remains accurate and relevant.

[0064] From a technical perspective, the described system achieves several advancements. Multimodal transformer fusion enables dynamic weighting and semantic alignment of heterogeneous modalities, thus overcoming the rigid coupling limitations of traditional data fusion. The adaptive agent controller introduces a modular and scalable approach to AI inference, mimicking biological cognitive processes through the dynamic creation of subagents. The personalization subsystem establishes an adaptive feedback loop that allows incremental learning without complete retraining, while the privacy-friendly federated update mechanism ensures ethical and secure deployment.Taken together, these technological innovations transform passive environmental monitoring into an active, inferential, and context-sensitive process that continuously interprets the dynamics of people and their environment, learns from them, and reacts in real time. The result is a self-evolving, user-centered intelligent system that integrates perception, cognition, and interaction into a single adaptive architecture.

[0065] The proposed machine comprises the following main subsystems: multimodal sensor module, data preprocessing and fusion unit, multimodal inference processing unit, agent-based automation controller, personalization and learning subsystem, and an output interface.

[0066] The multimodal sensor module integrates a range of devices, including high-resolution cameras for capturing visual data, directional microphones for audio input, environmental sensors for temperature, motion, and air quality data, and optional RFID modules for object detection. These sensors continuously collect multimodal input data from the monitored environment in real time.

[0067] The data preprocessing and fusion unit performs signal conditioning, noise reduction, and normalization across all modalities. Visual frames are extracted using convolutional transformers, while acoustic signals are converted into spectrogram embeddings. Sensor data is vectorized and temporally aligned with other modalities. The processed inputs are fused using a multimodal transformer, which aligns the representations across different modalities to create a unified contextual embedding.

[0068] The multimodal processing unit forms the cognitive core of the device. It incorporates a transformer-based multimodal LLM that has been pre-trained with paired audio-text, image-text, and sensor-text data. This model enables the system to infer semantic relationships, reason across modalities, and generate descriptive or analytical output in natural language. The processing unit includes specialized modules for activity detection, dialogue summarization, and anomaly detection.

[0069] The agent-based automation controller dynamically generates task-specific agents that interact with the logic engine to perform targeted functions. For example, a "Meeting Agent" monitors conversation histories, automatically transcribes them, and creates context-rich summaries with actionable recommendations. An "Object Agent" analyzes visual and sensor data to detect misplaced objects by comparing their current positions with stored environmental maps. A "Behavioral Agent" infers human activity from multimodal inputs, creates daily behavior logs, and detects anomalies in user routines.

[0070] The personalization and learning system stores user-specific data such as behavioral patterns, interaction history, and environmental preferences. Using reinforcement and transfer learning mechanisms, it updates the model parameters to adapt responses over time. This adaptive feedback loop ensures that outputs remain contextual and user-specific, thereby improving accuracy and user satisfaction.

[0071] The output interface consists of a graphical user interface (GUI) with integrated audio and text output. The system can display meeting summaries, visual reconstructions of object positions, and personalized recommendations. In some configurations, the device can communicate with external systems for enterprise integration via secure APIs.

[0072] The drawing and the preceding description illustrate embodiments. Those skilled in the art will recognize that one or more of the described elements can be combined to form a single functional element. Alternatively, certain elements can be divided into several functional elements. Elements of one embodiment can be added to another. For example, the process flows described here can be modified and are not limited to the manner described herein. Furthermore, the actions of a flowchart need not be performed in the sequence shown; nor do all actions necessarily need to be carried out. Actions that do not depend on other actions can be performed in parallel with the other actions. The scope of protection of the embodiments is in no way limited by these specific examples. Numerous variations, whether explicitly stated in the description or not, such as...Differences in structure, dimensions, and materials are possible. The scope of protection of the embodiments is at least as comprehensive as described by the following claims.

[0073] The advantages, other benefits, and problem solutions have been described above with reference to specific embodiments. However, the advantages, benefits, problem solutions, and any components that can effect or enhance an advantage, benefit, or solution are not to be construed as critical, necessary, or essential features or components of the claims. REFERENCES 100 An Integrated Multimodal Intelligent Agent System for Dynamic Environmental Monitoring and Human-Centered Support. 102 Multimodal sensor module 104 Data Preprocessing and Fusion Subsystem 106 Multimodal Processing Unit for Inferences 108 Adaptive Agent Controller 110 Personalization and Adaptive Learning Subsystem 112 Output generation interface

Claims

[1] A multimodal intelligent agent system for dynamic environmental monitoring and user-centric support, consisting of: a multimodal sensor module configured to continuously acquire environmental and behavioral data from multiple input modalities, including at least one visual sensor, at least one acoustic sensor, at least one environmental conditions sensor, and at least one proximity or motion detection sensor, each generating modality-specific data streams representing visual images, audio waveforms, physical environmental parameters, and motion signatures within a monitored environment; a data preprocessing and fusion subsystem that is operationally coupled with the multimodal sensor module and configured to normalize, temporally align, and transform the modality-specific data streams into high-dimensional feature embeddings using a variety of encoders, wherein the visual encoder uses convolutional or vision transformer architectures, the audio encoder uses a spectral-temporal feature extractor, and the sensor encoder transforms raw analog data into context vectors suitable for multimodal alignment; a multimodal processing unit consisting of a transformer-based large language model (LLM) trained on paired multimodal datasets and configured to perform semantic fusion, context abstraction, and inference across the aforementioned aligned multimodal feature embeddings to generate a contextual understanding of environmental and behavioral states; an adaptive agent controller coupled to the multimodal inference processing unit and configured to instantiate, manage, and terminate a variety of task-specific intelligent agents, each agent being a software unit configured to perform a specialized function selected from meeting summarization, behavioral analysis, misplaced object detection, or environmental anomaly identification, with the agents dynamically interacting with the inference engine to retrieve contextually relevant multimodal embeddings for task execution; a personalization and adaptive learning subsystem consisting of a user preference database and a neural memory structure configured to update and refine model parameters based on user-specific interaction history, thereby enabling personalized output generation, prioritization of recommendations, and long-term behavioral adaptation; and An output generation interface is operationally connected to the adaptive agent controller and configured to produce multimodal output in textual, visual, and auditory form. The interface is capable of displaying human-readable summaries, notifications, and visual reconstructions of identified entities or environmental states. [2] System according to claim 1, wherein the multimodal sensor module comprises a panoramic HD camera mounted on a motorized gimbal configured for three-dimensional rotation. The gimbal is controlled by a micro-servo array which dynamically adjusts the field of view in response to acoustic localization data acquired by the directional microphone array, thereby enabling adaptive visual tracking of active human participants or objects of interest in the monitored environment. [3] System according to claim 1, wherein the data preprocessing and fusion subsystem further comprises a multimodal synchronization buffer and a transformer-based fusion encoder configured to temporally align asynchronous data streams using timestamp correlation and learned attention weights, so that visual frames, audio transcripts and environmental sensor data are aligned to represent temporally coherent events, enabling accurate derivation of cause-and-effect relationships across modalities during real-time monitoring. [4] System according to claim 1, wherein the multimodal inference processing unit uses a two-stage architecture comprising: (i) a cross-modal alignment transformer that maps feature embeddings from visual, auditory, and sensory modalities into a unified latent representation, and (ii) a semantic inference transformer trained by supervised contrastive learning on multimodal instructional datasets to perform context-sensitive inference, thereby enabling the engine to generate natural language descriptions, event summaries, and predictive inferences based on environmental dynamics. [5] System according to claim 4, wherein the multimodal inference processing unit further integrates a causal inference submodule configured to determine directional relationships between multimodal variables, such as the assignment of a detected motion vector to a concurrently occurring acoustic signature or the identification of anomalous environmental fluctuations as causes of behavioral changes, thereby improving interpretability and reducing false positive inferences under complex environmental conditions. [6] System according to claim 1, wherein the adaptive agent controller dynamically generates a plurality of computational agents, each instantiated in a containerized runtime environment that includes memory isolation and API access controls. These agents are configured to subscribe to data streams via an internal event broker, autonomously determine task relevance based on attention score thresholds, and request multimodal embeddings from the reasoning engine to perform specialized operations, including, but not limited to, semantic summarization, behavioral tagging, or anomaly evaluation. [7] System according to claim 6, wherein at least one meeting summarization agent is configured to receive transcribed audio data and corresponding visual attention maps of the participants and applies a finely tuned large language model to generate structured meeting transcripts that include speaker-related action points, temporal annotations and contextual highlights, wherein the agent further validates the generated text with nonverbal cues such as gestural emphasis or tone variations, which are recognized by the multimodal inference processing unit. [8] System according to claim 6, wherein at least one object detection agent is configured to analyze fused visual and proximity sensor data to track static and dynamic objects, generate a reference inventory of expected positions, and detect misplaced or removed objects by comparing live embeddings with historical scene graphs stored in the personalization subsystem, wherein the agent provides a visual reconstruction of the anomalous spatial area via the output interface for user verification. [9] System according to claim 6, wherein a behavior analysis agent is configured to derive human activity sequences from fused visual and motion sensor data, segment behavior episodes using a temporal activity transformer, and generate natural language descriptions of behavior routines or deviations therefrom, wherein the agent further computes a behavior consistency index over time to detect irregularities that may indicate fatigue, stress, or anomalous events. [10] System according to claim 1, wherein the personalization and adaptive learning subsystem comprises a recurrent memory encoder configured to maintain a user-specific latent state vector that captures previous environmental conditions, user preferences and task interactions, and wherein the subsystem uses Reinforcement Learning with Human Feedback (RLHF) to update weighting parameters of the multimodal processing unit, thereby achieving incremental adjustment of responses without retraining the core model.

Citation Information

Cited By

  • Multi-monitoring equipment collaborative operation method and device based on Internet of Things

    CN121814929A

  • Online learning method and online learning platform for security training

    CN121860825A

  • Intelligent customer service interaction method and system based on AI large model

    CN121903622A

  • Multi-modal large model driven Internet of Things Agent adaptive interaction system

    CN121959493A

  • Agent operation system and method oriented to long-term task and driven by task state

    CN121960559A