Intelligent security and risk response system based on integrated analysis of multimodal data
The intelligent security system integrates multimodal data for precise detection and automated response, addressing vulnerabilities in existing systems by enhancing accuracy and situational awareness.
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- KOREA ELECTRONICS TECH INST
- Filing Date
- 2025-11-06
- Publication Date
- 2026-07-29
AI Technical Summary
Existing physical security systems rely heavily on image processing and lack integration of audio and sensor data, leading to vulnerabilities in detecting complex risks, situational understanding, and automated response capabilities.
An intelligent security and risk response system that integrates video, audio, and sensor data using deep learning-based feature extraction, fusion, and spatiotemporal analysis to accurately detect and respond to risk situations.
Enhances detection accuracy in blind spots and low-light environments, provides situational understanding, and enables automated responses tailored to risk levels, minimizing security incidents.
Smart Images

Figure 112025124178688-PAT00010_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to an intelligent security and risk response system applicable to the field of physical security and a method of operation thereof. Specifically, it relates to an intelligent security and risk response system and a method of operation thereof that recognizes and responds to risk situations based on multimodal data. Background Technology
[0002] 1. Limitations of Single-Modal Based Security Systems
[0003] Existing physical security systems primarily rely on image processing technologies based on CCTV, such as object recognition and intrusion detection. However, technologies that comprehensively utilize information from audio (e.g., screams, impact sounds) or sensors (e.g., motion, temperature, gas, etc.) in addition to CCTV footage are currently lacking.
[0004] Existing physical security systems relying on image processing technology are vulnerable to changes in lighting, visibility obstacles, and blind spots, making it highly likely that they will fail to detect critical dangers or generate false positives.
[0005] Furthermore, recognizing complex physical risks such as assault, theft, and arson requires an intelligent structure capable of integrally processing various physical signals, but existing physical security systems fail to meet this requirement.
[0007] 2. Lack of situational understanding due to non-reflection of temporal and spatial information
[0008] Existing security technologies detect risks based solely on images or sensor values at a specific point in time, failing to consider time-series characteristics such as the timing, duration, and progression of an action. For instance, to distinguish whether a "fall" is a simple slip or a state of unconsciousness caused by toxic gas poisoning, it is necessary to analyze environmental changes immediately before and after the action, as well as the object's movement path; however, existing security technologies lack the capability to analyze such temporal and spatial context.
[0009] Furthermore, in existing security technologies, spatial information is not utilized for purposes such as determining distances between objects, access paths, or identifying risk zones, making it difficult to accurately infer causes or assess risk levels.
[0011] 3. Lack of automation in response judgment and execution
[0012] Even when a dangerous situation is detected, existing security systems are limited to mechanical responses such as sending simple alarms, generating warning sounds, and saving video clips, making real-time judgment and dynamic response impossible. Furthermore, since existing systems rely on the operator's subjective judgment for situation assessment and response strategies, immediate response may be delayed or judgment errors may occur during a crisis, and they lack the intelligence of risk response strategies, such as determining action priorities based on risk levels and applying multi-stage policies. Prior art literature
[0013] KR 10-1853903 B1 (Registration Date April 25, 2018)KR 10-2576651 B1 (Registration Date September 5, 2023) The problem to be solved
[0014] Unlike conventional security technologies based on single-modal data, the present invention aims to provide a converged security intelligence generation technique capable of recognizing dangerous situations with high accuracy and responding preemptively by integrating and analyzing video, audio, and sensor-based multimodal data in real time.
[0015] Furthermore, unlike conventional security technologies that lack situational understanding due to the non-reflection of temporal and spatial information, the present invention aims to provide a spatiotemporal fused security intelligence generation technique capable of understanding high-dimensional risk situations based on context by integrating and analyzing behavior detection, object tracking, and spatial positional relationships based on time-series multimodal data.
[0016] In short, the present invention aims to provide a method for generating intelligent composite security intelligence capable of precise response judgment and automatic execution based on the results of integrating and analyzing CCTV video, audio signals, and environmental sensor data in a physical security system to detect risk situations and understand the situation while reflecting temporal and spatial contexts.
[0017] In particular, the present invention has the ultimate goal of overcoming the limitations of detection accuracy, situational misunderstanding, and response delays associated with existing single-modal-based detection technologies, and implementing complex security intelligence capable of preemptive response.
[0018] The objectives of the present invention are not limited to those mentioned above, and other unmentioned objectives will be clearly understood by those skilled in the art from the description below. means of solving the problem
[0019] A method of operation of an intelligent security and risk response system according to an embodiment of the present invention comprises: a step of generating a plurality of feature vectors by passing multimodal data through a deep learning-based feature extractor for each modality; a step of generating fused features by dynamically analyzing and processing the plurality of feature vectors using a feature fusion module, which is a deep learning-based adaptive modality fusion engine; a step of determining an event type through spatiotemporal analysis of the fused features; a step of generating situational context information including the context and cause of a situation through inference using a deep learning-based situational understanding module based on the fused features and the event type; and a step of selecting a response scenario and a measure item according to the situational context information, and executing an automated response by linking with an external physical system according to the measure item.
[0021] An intelligent security and risk response system according to one embodiment of the present invention includes: a processor; and a memory that stores one or more instructions executed through the processor.
[0022] The above one or more commands include: a command to generate multiple feature vectors by passing multimodal data through a deep learning-based feature extractor for each modality; a command to generate fused features by dynamically analyzing and processing the multiple feature vectors using a feature fusion module, which is a deep learning-based adaptive modality fusion engine; a command to determine an event type through spatiotemporal analysis of the fused features; a command to generate situational context information including the context and cause of a situation through inference using a deep learning-based situational understanding module based on the fused features and the event type; and a command to select a response scenario and an action item according to the situational context information, and to execute an automated response by linking with an external physical system according to the action item. Effects of the invention
[0024] The present invention has the following effects.
[0025] 1. Highly accurate detection of complex risk situations possible
[0026] According to the present invention, through the integrated analysis of multimodal data, dangerous situations can be recognized with high precision even in blind spots, low-light environments, and invisible abnormal situations (e.g., gas leaks, impact sounds, etc.).
[0028] 2. Implementation of a situation understanding function capable of identifying the context and cause of behavior occurrence
[0029] According to the present invention, by considering the flow of time, interactions between objects, and location information within space together, it is possible to infer the cause and determine the context of a detected action, thereby reducing the possibility of misunderstanding or omission of the action.
[0031] 3. Enhancement of real-time security through automated response execution
[0032] According to the present invention, measures tailored to the level of risk can be automatically performed according to a defined policy scenario, thereby enabling the securing of a rapid and precise response system without operator intervention. Therefore, according to the present invention, monitoring efficiency is increased and security incidents are minimized.
[0033] The effects obtainable from the present invention are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art from the description below. Brief explanation of the drawing
[0035] FIG. 1 is a block diagram showing the configuration of an intelligent security and risk response system according to one embodiment of the present invention. Figure 2 is a block diagram showing the configuration of a Security Intelligence Model (SIM) installed in the system of Figure 1. FIG. 3 is a flowchart illustrating the operation method of an intelligent security and risk response system according to an embodiment of the present invention. Specific details for implementing the invention
[0036] The present invention relates to an intelligent security and risk response system and a method of operation thereof in the field of physical security. Specifically, the present invention relates to a system and a method of operation thereof for recognizing and responding to dangerous situations by analyzing temporal and spatial characteristics using multimodal data such as video, audio, and sensors.
[0037] The present disclosure proposes an intelligent security framework that enables precise security response without human intervention by evaluating risk levels based on detected risk situations and situation analysis results, and automatically executing access control, alarm transmission, and monitoring display according to predefined response scenarios.
[0038] The features of the present invention compared to existing technology are as follows.
[0039] (1) Multimodal-based real-time integrated analysis structure
[0040] While existing technologies have remained limited to image-centric single-modal detection, this invention is based on a structure that integrates and processes heterogeneous data, such as video, audio, and sensors, in real time. This compensates for the context and conditions missed by a single signal, thereby improving the precision of risk situation detection and the rate of preventing omissions.
[0041] (2) Enhancement of context awareness based on time series and spatial information
[0042] While existing technologies were limited to static judgments at individual points in time, the present invention considers the occurrence, duration, and termination points of an action, as well as changes in the object's position, together. For example, the system according to the present invention interprets the cause of an action, such as 'falling over,' in a high-dimensional manner by combining it with environmental changes or surrounding sensor information.
[0043] (3) Automated risk assessment and response scenario execution structure
[0044] While existing technologies rely on manual or simplified logic for response after detection, the present invention determines the level of risk based on the detected dangerous situation and the results of the situation analysis. Furthermore, the present invention enables automated response measures to be taken based on the risk level derived from the analysis of the dangerous situation. For example, the system according to the present invention can immediately execute situation-specific response logic, such as access control, warning dissemination, and monitoring display, based on scenarios.
[0046] The advantages and features of the present invention and the methods for achieving them will become clear by referring to the embodiments described below in detail together with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below but may be implemented in various different forms. These embodiments are provided merely to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the invention, and the present invention is defined only by the scope of the claims. Meanwhile, the terms used in this specification are for describing the embodiments and are not intended to limit the present invention. In this specification, the singular form includes the plural form unless specifically stated otherwise in the text. The terms "comprises" and / or "comprising" as used in this specification do not exclude the presence or addition of one or more other components, steps, actions, and / or elements in addition to the mentioned components, steps, actions, and / or elements.
[0047] Terms such as "first," "second," etc., may be used to describe various components, but said components should not be limited by said terms. These terms may be used for the purpose of distinguishing one component from another. For example, without departing from the scope of the present invention, the first component may be named the second component, and similarly, the second component may be named the first component.
[0048] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. Conversely, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between. Other expressions describing the relationship between components, such as "between" and "exactly between," or "adjacent to" and "directly adjacent to," should be interpreted in the same way.
[0049] In describing the present invention, detailed descriptions of related prior art are omitted if it is determined that such descriptions may unnecessarily obscure the essence of the invention.
[0050] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. In order to facilitate an overall understanding in describing the present invention, the same reference numerals will be used for the same means regardless of the drawing number.
[0052] FIG. 1 is a block diagram showing the configuration of an intelligent security and risk response system according to one embodiment of the present invention.
[0053] The intelligent security and risk response system (100) integrates and analyzes multimodal data (video, audio, sensor data) to recognize a risk situation based on time series and spatial information, and generates intelligent security intelligence to respond thereto. The intelligent security and risk response system (100) can extract features from multimodal data received from an external terminal or server or collected through a sensor, fuse the extracted features, recognize a situation based on the fused features, respond to the recognized situation, and perform learning of the security intelligence model.
[0054] As shown in FIG. 1, an intelligent security and risk response system (100) according to one embodiment of the present invention can be implemented in the form of a computer system.
[0055] Referring to FIG. 1, an intelligent security and risk response system (100) according to one embodiment of the present invention comprises a processor (110), a communication device (120), a memory (130), a storage device (140), an input interface device (150), an output interface device (160), and a bus (170).
[0056] The intelligent security and risk response system (100) illustrated in FIG. 1 is according to one embodiment, and the components of the intelligent security and risk response system (100) according to the present invention are not limited to the embodiment illustrated in FIG. 1 and may be added, changed, or deleted as needed.
[0057] For reference, unlike FIG. 1, the intelligent security and risk response system (100) according to one embodiment of the present invention may be implemented in software or in the form of hardware such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).
[0058] The processor (110) may be a cluster comprising at least one central processing unit (CPU) and / or at least one graphics processing unit (GPU). Additionally, the processor (110) may be a semiconductor device that executes computer-readable instructions stored in memory (130) or a storage device (140).
[0059] The communication device (120) can transmit or receive wired signals or wireless signals. The communication device (120) may include both a wired communication module and a wireless communication module. The wired communication module may be implemented as a power line communication device, a telephone line communication device, a cable home (MoCA), Ethernet, IEEE1294, an integrated wired home network, and an RS-485 control device. Additionally, the wireless communication module may be composed of a module for implementing functions such as WLAN (wireless LAN), Bluetooth, HDR WPAN, UWB, ZigBee, Impulse Radio, 60GHz WPAN, Binary-CDMA, wireless USB technology and wireless HDMI technology, as well as 5G (5th generation communication), LTE-A (long term evolution-advanced), LTE (long term evolution), and Wi-Fi (wireless fidelity).
[0060] The memory (130) and storage device (140) may include various forms of volatile or non-volatile storage media. For example, the memory (130) may include ROM (read-only memory) and RAM (random access memory). In the embodiments of this description, the memory (130) may be located inside or outside the processor (110), and the memory (130) may be connected to the processor (110) through various known means. The memory (130) is various forms of volatile or non-volatile storage media, and for example, the memory (130) may include read-only memory (ROM) or random access memory (RAM).
[0061] Accordingly, embodiments of the present invention may be implemented as a method implemented on a computer or as a non-transient computer-readable medium storing computer-executable instructions. In one embodiment, when executed by a processor (110), the computer-readable instructions may perform a method according to at least one aspect of the present disclosure.
[0062] In addition, the operation method of the intelligent security and risk response system (100) according to an embodiment of the present invention may be implemented in the form of program instructions that can be executed through various computer means and may be recorded on a computer-readable medium.
[0063] The above computer-readable medium may include program instructions, data files, data structures, etc., either individually or in combination. The program instructions recorded on the computer-readable medium may be specially designed and configured for embodiments of the present invention, or they may be known and available to a person skilled in the art of computer software. The computer-readable recording medium may include a hardware device configured to store and execute program instructions. For example, the computer-readable recording medium may be magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; ROM; RAM; flash memory, etc. The program instructions may include not only machine code, such as that generated by a compiler, but also high-level language code that can be executed by a computer through an interpreter, etc.
[0064] The processor (110) executes computer-readable instructions stored in memory (130) or storage device (140) to execute the method of operation of the intelligent security and risk response system described below with reference to FIG. 3.
[0066] Figure 2 is a block diagram showing the configuration of a Security Intelligence Model (SIM) installed in the system of Figure 1.
[0067] The processor (110) inputs multimodal data (210) into a security intelligence model (SIM) to obtain contextual information (270).
[0068] As illustrated in FIG. 2, the security intelligence model (SIM) includes a data preprocessor (220), a feature extractor (230), a feature fusion module (240), an event detection module (250), and a situation understanding module (260).
[0070] The intelligent security and risk response system (100) receives multimodal data (210) in real time through a communication device (120) or an input interface device (150), or collects it from a connected or built-in sensor (not shown) and inputs it into a data preprocessor (220) of a security intelligence model (SIM). For example, the multimodal data (210) may include video (e.g., CCTV video), acoustic data (e.g., scream, explosion sound, footsteps, etc.), and sensor data (e.g., temperature, gas, vibration, access logs, etc.). Acoustic data includes voice data and may be a digitized audio signal. The processor (110) synchronizes each data constituting the multimodal data (210) in a time-series form and inputs it into the data preprocessor (220).
[0072] A data preprocessor (220) preprocesses image, audio, and sensor data constituting multimodal data (210). For example, the data preprocessor (220) can perform basic preprocessing operations such as normalizing each multimodal data (210), removing noise from the multimodal data (210), or aligning the timestamps of the multimodal data (210). Additionally, the data preprocessor (220) can adjust the resolution of the image included in the multimodal data (210), convert the unit of the sensor data, or perform spectrogram conversion of the audio signal. The processor (110) inputs the multimodal data (210) that has passed through the data preprocessor (220) into a feature extractor (230).
[0074] A feature extractor (230) generates multiple feature vectors by passing preprocessed multimodal data (210) through a modality-specific deep learning-based feature extractor (CNN, RNN, etc.). For example, the feature extractor (230) may include an image feature extractor (231), an audio feature extractor (232), and a sensor data feature extractor (233). In this case, the image feature extractor (231) receives an image as input and generates an image feature vector (F v , e.g., character action feature vector) can be generated, and the audio feature extractor (232) receives acoustic data as input and generates an acoustic feature vector (F a , e.g., frequency feature vector) can be generated. In addition, the sensor data feature extractor (233) receives sensor data as input and generates a time series feature vector (F s It can generate ). The processor (110) inputs a plurality of feature vectors generated by the feature extractor (230) into the feature fusion module (240).
[0076] The Feature Fusion Module (240) can be described as an Adaptive Modality Fusion Engine.
[0077] The feature fusion module (240) is a module that generates fused features by dynamically analyzing and processing feature vectors extracted from various multimodal data (210), such as video, audio data, and sensor data. The feature fusion module (240) includes a feature alignment module (241), a feature attention module (242), a modality memory module (243), and a dimension reduction module (244). In the security intelligence model (SIM), the feature fusion module (240) is a core processing engine that fuses multiple feature vectors into a single integrated semantic representation.
[0078] The feature fusion module (240) plays a central role in the overall security intelligence flow, from situation awareness and judgment to response, by selectively activating or deactivating each internal module (241, 242, 243, 244) according to the type of risk situation, the configuration of input multimodal data (210), and environmental conditions. For example, when all internal modules of the feature fusion module (240) are activated, the feature vector can be converted into a fused feature by sequentially passing through the feature alignment module (241), the feature attention module (242), the modality memory module (243), and the dimensionality reduction module (244).
[0080] Below, all internal modules included in the feature fusion module (240) are activated, and feature vectors (F v , F a , F s) passes through the feature alignment module (241), feature attention module (242), modality memory module (243), and dimensionality reduction module (244) to fuse features (F reduced Under the premise that it is converted into ), the function of each module of the feature fusion module (240) is explained.
[0081] The Feature Alignment Module (241) is a module that supports the integration of multiple information based on accurate time and location by correcting temporal or spatial asynchronous problems between different multimodal data (210), such as video, audio data, and sensor data. The Feature Alignment Module (241) comprises a plurality of feature vectors (F v , F a , F s By receiving ) as input, sorting based on time and physical space, and mapping them to a common mathematical space so that vectors with high semantic similarity are placed at close coordinates in the said mathematical space, the sorted features (F align , aligned feature) is generated. That is, the feature alignment module (241) generates an image feature vector (F v ), acoustic feature vector(F a ) and time series feature vector(F s By aligning ) based on time and space and then mapping them to a common representation space, the aligned features (F align It generates ). For example, the feature alignment module (241) may include deep learning models such as a Token Flow, Multimodal Perceiver, Attention, and a Transformer-based alignment network that performs positional embedding alignment for the aforementioned temporal, spatial, and semantic alignment tasks. These models become capable of aligning different modal features through a learning process.
[0083] Before describing the feature attention module (242), the operation of the feature alignment module (241) is described in detail.
[0084] The feature alignment module (241) performs the role of generating a single integrated multimodal feature representation by aligning and matching feature vectors extracted from different modalities, such as images, audio, and sensors, temporally and spatially. However, this process is not a simple concatenation, but a combination of temporal alignment, spatial calibration, and representation alignment.
[0085] When using multimodal data (e.g., video, sensor data, acoustic data) simultaneously, there is a problem in that the time points of the spatial information observed by each modal do not exactly coincide. For example, a camera captures a person falling at point A at t=0 seconds, but a hazardous gas sensor is installed at a location slightly away from point A, so a change in gas is not detected immediately after the person falls, but a rise in concentration may be observed around t=5 seconds. Therefore, for integrated analysis of multimodal data, spatiotemporal observation discrepancies between modals must be resolved. That is, a process of mapping features obtained from different modals to a common time axis and spatial axis is required. This role is performed by a feature alignment module (241). The feature alignment module (241) performs temporal alignment, spatial calibration, and representation alignment.
[0086] First, the feature alignment module (241) corrects timestamp differences, latency, and sampling period differences occurring between modals through temporal alignment. The feature alignment module (241) realigns each modal by modeling timestamp differences and latency occurring between modals (e.g., sensor response speed, communication delay). In the example described above, the increase in the concentration of the hazardous gas sensor occurred at t=5, but the feature alignment module (241) can realign this sensor data (concentration) to the situation at t=0. Additionally, as an example regarding the correction of sampling period differences, if the frame rate of the image data is 30fps, the frequency of the sensor data is 10Hz, and the frequency of the acoustic data is 16kHz, the feature alignment module (241) plays a role in correcting these sampling period differences.
[0087] Additionally, the feature alignment module (241) converts different physical coordinate systems (e.g., camera images, sensor measurement locations, etc.) to a common reference through spatial calibration. That is, the feature alignment module (241) converts different coordinate systems, such as sensor data, images, and maps, into a common reference coordinate system. For example, the feature alignment module (241) aligns the camera field of view and the gas sensor location to the same spatial frame. As another example, the feature alignment module (241) maps CCTV pixel coordinates to an actual ground coordinate system.
[0088] Additionally, the feature alignment module (241) semantically aligns the features of different modals using techniques such as cross-attention and projection during the representation alignment process. This process is not a simple vector connection, but a process of generating a refined multimodal representation by mapping multiple feature vectors into a common representation space.
[0089] In one embodiment of the present invention, aligned features (F) generated in the feature alignment module (241) align ) sequentially passes through the feature attention module (242), the modality memory module (243), and the dimensionality reduction module (244) to obtain a dimensionality-reduced fused feature (F reduced It is converted into ).
[0091] The Feature Attention Module (242) is a feature attention module that handles aligned features (F align It is a module that assigns weights to important information among ) to suppress unnecessary information and allow focus on key risk signals. The feature attention module (242) is an aligned feature (F align ) is received as input, and cross-attention and / or multi-head attention is applied to the aligned features (F align 'Attention-weighted fusion features' (F) to which weights are applied to important information included in ). att , attention-weighted fused feature) is generated. For example, the feature attention module (242) uses at least one deep learning model among Q-Former, Perceiver Resampler, or Transformer Encoder to generate attention-weighted fused features (F att Can generate ).
[0092] The modality memory module (243) is a module that enables risk perception and prediction by reflecting past states and behavioral contexts by accumulating and storing changes in feature vectors over time based on time-series data. The modality memory module (243) is an attention-weighted fused feature (F att) is received as input, and by applying Mamba, a Selective State Space-based time-series state update and long-term dependency learning technique, or RetNet (Retentive Network), a linear attention and recurrence-based Temporal Context Accumulation technique, memory enhancement features for each modality reflecting temporal context ( Creates ).
[0093] The Dimension Reduction Module (244) is a module that receives a high-dimensional multimodal feature vector as input, compresses it to a lower dimension while retaining necessary information, thereby increasing model computation efficiency and preventing overfitting. The Dimension Reduction Module (244) can maintain the spatial feature map form as a latent bottleneck (e.g., T*C*H*W) without completely compressing it into a vector. The Dimension Reduction Module (244) uses a deep learning model-based encoder to generate modality-specific memory-enhanced features ( Fusion features (F), which are the finally fused features obtained by reducing and refining the dimensions of ) reduced Generates ). The above fusion feature is a modality-specific memory enhancement feature ( It can be said that this is a feature expressed more efficiently compared to ). For example, the dimensionality reduction module (244) is an encoder that reduces the dimensionality of the feature, and can use MAE, Perceiver IO Bottleneck, etc.
[0095] As described above, the feature fusion module (240) uses each internal module (241, 242, 243, 244) to fuse a plurality of feature vectors and reduce the dimensionality to produce refined fused features (F reduced ) generates. The processor (110) inputs the fusion features generated by the feature fusion module (240) into the event detection module (250).
[0097] The Event Detection Module (250) is a fusion feature (F reduced This is a module that determines the event type through temporal-spatial analysis of ). The event detection module (250) analyzes the event flow on the time axis and spatial relationships such as distance, location, and movement path between objects. For example, the event detection module (250) includes a spatial convolutional model and a temporal convolutional model, and can learn temporal patterns using these models. Here, the spatial convolutional model analyzes the position, attitude, sensor value distribution, and spatial structure within the frame, and the temporal convolutional model can learn changes between frames (movement, state change, environment change).
[0098] The event detection module (250) has fusion features (F reduced By inputting ) into an already trained deep learning model, fusion features (F reduced Extracts spatiotemporal patterns from ) and determines event types based on spatiotemporal patterns.
[0099] For example, deep learning models that the event detection module (250) can use as event detection models include R(2+1)D, Video Swin Transformer, TimesNet+ST-GNN, etc. R(2+1)D is a model that efficiently learns spatiotemporal patterns by decomposing 3D Conv into 2D+1D, and has strengths in action recognition and complex event detection. Video Swin Transformer is a spatiotemporal window-based hierarchical attention, and has strengths in large-scale event patterns. In addition, TimesNet + ST-GNN is a model that detects events by combining time-series patterns and spatial relationships, and has strengths in detecting events by analyzing the situation based on sensors.
[0100] As described above, the event detection module (250) determines the event type (e.g., collapse, abnormal behavior) based on spatiotemporal patterns. For example, when the event detection module (250) performs spatiotemporal analysis based on fusion features, if it determines that a 'collapse' behavior has occurred, and subsequently there is no change in location and the concentration of harmful gas has risen above a threshold, it can determine the event type as 'poisoning collapse' by synthesizing these judgment results. The event detection module (250) determines the event type through the learning of spatiotemporal patterns and contextual inference. In this case, the event detection model of the event detection module (250) automatically infers the event type of 'poisoning collapse' by spatiotemporally integrating the temporal sequence (person collapses → no change in location → rise in gas concentration) and the spatial context (danger zone, sensor location, person's location).
[0101] As previously explained, R(2+1)D, ST-GNN, and Video Swin Transformer can be used as event detection models. To introduce the training methods of each model, R(2+1)D performs spatiotemporal analysis by integrating behavior and changes in the environment, while ST-GNN learns spatiotemporal relationships between sensors, locations, and objects (e.g., people) in the form of a graph. Additionally, the Video Swin Transformer performs inference by connecting events and sensor tokens with attention.
[0102] Meanwhile, an event detection model can be constructed by combining rules and model learning. For example, since it is often difficult to obtain sufficient complex situation data, a hybrid approach can be adopted in constructing the event detection model, where basic situation judgments are rule-based and exceptions and complex situations are processed based on learning. For example, a rule can be set in the event detection module (250) to determine a candidate for collapse due to poisoning if 'collapse occurs', 'no change in position for 10 seconds', or 'increase in gas concentration' is detected.
[0104] The processor (110) inputs the fused features generated by the feature fusion module (240) and the event types generated by the event detection module (250) into the situation understanding module (260).
[0105] The Situation Understanding Module (260) is a convergence feature (F reduced It infers the context and cause of the situation based on the ) and event types. That is, the situation understanding module (260) inputs the fused features and event types into an already learned deep learning model to generate situation context information (270).
[0106] In the present disclosure, situation context information (270) is information including the context, cause, and semantic interpretation of a situation, and may be referred to as 'situation context and cause representation'. For example, the situation understanding module (260) may distinguish between 'simple collapse' and 'collapse due to a hazardous environment', or distinguish between 'intrusion' and 'authorized night entry'. The situation context information (270) may include a situation class (e.g., collapse due to poisoning, stop after intrusion) and situation description text. The situation description text may be generated when using a vision language model.
[0107] For example, deep learning models that can be used by the situation understanding module (260) include ST-GNN, Transformer, and Vision Language Model (VLM). Vision Language Models that can be used include BLIP-2, Flamingo, Kosmos-2, GPT-4V, etc. ST-GNN sets people, sensors, locations, etc. as nodes and configures spatiotemporal relationships as edges to perform graph-based inference, and has strengths in interpreting complex situational contexts. Transformer has fused features (F reduced It learns contextual interactions using multi-head attention by tokenizing ) and event types (event detection results). The vision language model describes situations in natural language or performs inference about complex causes by integratively processing visual and linguistic information.
[0109] Table 1 summarizes the functions, input data, and output data of each internal module (241, 242, 243, 244) of the feature fusion module (240), the event detection module (250), and the situation understanding module (260). Table 2 explains the meaning of the input and output data in Table 1. Additionally, Table 3 shows the deep learning models applicable to each module and the computational processing method of each deep learning model.
[0111] Module name Input data Output data explanation Feature alignment module (241) Image feature vector (F v ), acoustic feature vector(F a ), time series feature vector(F s ) Aligned features (F align , Aligned feature) Create a common standard representation by aligning video, audio, and sensor features based on time and space. Feature attention module (242) F align Attention-weighted fusion feature (F att , attention-weighted fused features) Generate core fusion features by applying importance-based weights to aligned features. Modality memory module (243) F att Memory Enhancement Features ( , Memory-enhanced features) Generation of modality-specific memory-enhancing features reflecting temporal context Dimension reduction module (244) Dimensionally reduced fusion features (F reduced , Low-dimensional fused features) Generating efficient representations by reducing and refining high-dimensional multimodal features Event detection module (250) F reduced Event type (Event class) Classification of Risk Event Types Situation understanding module (260) Event type, F reduced Contextual information (F context , Situation context information) Inferring the context and causes of a situation by analyzing comprehensive patterns of time, space, behavior, and environment
[0113] data meaning F v , F a , F s Basic features extracted from video, audio, and sensor backbones F align Integrated features aligned temporally and spatially and mapped to a common representation space F att fused features based on attention Memory-enhanced features reflecting temporal context F reduced Final fusion characteristics after dimensionality reduction Event type (event class) Results of detecting and classifying events based on dimensionality-reduced fusion features
[0114] Module role Applicable deep learning models Matching operation / processing method Feature alignment module (241) Temporal and spatial asynchronous alignment between video, audio, and sensors, and common baseline transformation Token Flow (Zhang et al., 2024) Temporal Attention-based temporal alignment and delay correction between modals Multimodal Perceiver (Jaegle et al., 2022) Cross-Attention-based feature projection and spatial coordinate matching between modals Transformer-based alignment network Sampling rate correction + positional embedding alignment Feature attention module (242) Highlighting important information among aligned features, reinforcing key information related to risk situations Q-Former (BLIP-2, 2023) Selecting important features and assigning weights using query token-based Cross-Attention Perceiver Resampler (Flamingo, 2022) Multi-Head Attention-based Latent Sampling and Importance Transformer Encoder Feature reweighting using attention maps Modality memory module (243) Reflection of past states and context of change along the time axis, accumulation of the event progression Mamba (Gu & Dao, 2023) Selective State Space-based Time Series State Update and Long-Term Dependency Learning RetNet (2023) Temporal Context Accumulation based on Linear Attention + Recurrence Dimension reduction module (244) Generates an efficient representation by compressing high-dimensional fusion features into latent vectors. MAE (He et al., 2022) Efficient dimensionality reduction through Masking, Encoding, and Latent Bottleneck Projection Perceiver IO Bottleneck (2021) Transform high-dimensional features into latent bottleneck vectors (information distillation) Event detection module (250) Classification of event types and calculation of time boundaries, detection of spatiotemporal behavioral patterns TimesNet(2023) Time-series event pattern detection based on temporal 2D variation encoding TFT / Temporal CNN Temporal feature sequence classification and boundary detection R(2+1)D (Tran et al., 2018) Decomposing 3D Convolution into 2D spatial + 1D temporal to learn spatial-temporal patterns and detect behavioral events Situation understanding module (260) Synthesizing spatiotemporal patterns to infer context and cause, interpret meaning, and explain ST-GNN / Transformer Temporal-Spatial Graph Reasoning, Interpretation of situational patterns based on spatiotemporal relationships BLIP-2 (Q-Former), Flamingo Inferring situational meaning and causes with Vision-Language Querying + Cross-attention Kosmos-2 / GPT-4V Vision+Language Integrated Reasoning and Natural Language Contextualization (XAI)
[0115] The processor (110) performs a risk assessment and response judgment based on situation context information (270). That is, the processor (110) evaluates the risk level based on the situation context information (270)—that is, the detected situation and circumstances—and determines a response scenario that matches the situation and / or risk level. The processor (110) automatically selects action items according to the response scenario through policy mapping. Then, the processor (110) executes an automated response according to the selected action items. As a specific example, the processor (110) performs responses such as access control, alarms, and notifications by linking with a physical system that is embedded in or external to the intelligent security and risk response system (100). That is, the processor (110) performs actual actions (e.g., locking doors, sending warnings to the control room, administrator push notifications, linking with the fire fighting system) by linking with the physical system according to the determined response scenario. In the present disclosure, a ‘response scenario’ refers to a series of response procedures that an intelligent security and risk response system (100) must perform according to situation context information (270), and a ‘measure’ refers to an individual execution unit (task) corresponding to the response scenario.
[0116] As a specific example, if the situation context information (270) is 'occurrence of poisoning collapse', the processor (110) determines response scenarios such as issuing an alarm, requesting the dispatch of safety personnel, automatically closing the gas shut-off valve, CCTV tracking, and situation logging, and can execute 'broadcasting an emergency broadcast' and 'issuing an on-site alarm' as measures in response to the alarm issuance. This is possible through the operation of a device (e.g., alarm device, CCTV) linked to an output interface device (160) or an intelligent security and risk response system (100), or through the transmission of a message to said device as a recipient.
[0118] The processor (110) processes data ranging from multimodal data (210) to situational context information (270), and stores risk rating assessments based on the situational context information (270), response scenarios, and action results in the database of the storage device (140). Then, it continuously trains the security intelligence model (SIM) based on user feedback of the intelligent security and risk response system (100) and performs intelligence improvement. This is a key element for improving security situation response capabilities and model precision in the long term.
[0120] FIG. 3 is a flowchart illustrating the operation method of an intelligent security and risk response system according to an embodiment of the present invention.
[0121] Referring to FIG. 3, the operation method of an intelligent security and risk response system (100) according to one embodiment of the present invention consists of steps S310 to S370. The operation method of the intelligent security and risk response system (100) illustrated in FIG. 3 is according to one embodiment, and the steps of the operation method of the intelligent security and risk response system (100) according to the present invention are not limited to the embodiment illustrated in FIG. 3 and may be added, changed, or deleted as necessary.
[0123] Step S310 is the step for receiving multimodal data.
[0124] The intelligent security and risk response system (100) receives multimodal data (210) from a communication device (120), an input interface device (150), or a sensor connected to or built into the intelligent security and risk response system (100).
[0125] Multimodal data (210) may include any one of video, audio data and sensor data or a combination thereof.
[0127] Step S320 is the data preprocessing step.
[0128] The intelligent security and risk response system (100) preprocesses video, audio data, and sensor data constituting multimodal data (210). For example, the intelligent security and risk response system (100) can perform basic preprocessing operations such as normalizing each of the multimodal data (210), removing noise from the multimodal data (210), or aligning the timestamps of the multimodal data (210). Additionally, the intelligent security and risk response system (100) can adjust the resolution of the video included in the multimodal data (210), convert the unit of the sensor data, or perform spectrogram conversion of the audio signal.
[0130] Step S330 is the feature extraction step.
[0131] The intelligent security and risk response system (100) passes multimodal data through a modality-based deep learning feature extractor (230) to generate multiple feature vectors. The multiple feature vectors may be any one of image feature vectors, acoustic feature vectors, and time series feature vectors, or a combination thereof.
[0132] For example, the feature extractor (230) may include an image feature extractor (231), an audio feature extractor (232), and a sensor data feature extractor (233). In this case, the image feature extractor (231) receives an image as input and an image feature vector (F v, e.g., character action feature vector) can be generated, and the audio feature extractor (232) receives acoustic data as input and generates an acoustic feature vector (F a , e.g., frequency feature vector) can be generated. In addition, the sensor data feature extractor (233) receives sensor data as input and generates a time series feature vector (F s Can generate ).
[0134] The S340 stage is the feature fusion stage.
[0135] The intelligent security and risk response system (100) uses a feature fusion module (240), which is a deep learning-based adaptive modality fusion engine, to dynamically analyze and process multiple feature vectors to produce fused features (F reduced Creates ).
[0136] The feature fusion module (240) includes a feature alignment module (241), a feature attention module (242), a modality memory module (243), and a dimension reduction module (244).
[0137] The feature alignment module (241) aligns a plurality of feature vectors based on time and space, maps the aligned feature vectors to a common representation space, and aligns the features (F align ) generates. The feature attention module (242) uses either cross attention and multihead attention or a combination thereof to generate aligned features (F align By assigning weights to some information included in ) that is judged to be important information, attention-weighted fusion features (F att ) generates. The modality memory module (243) generates the attention-weighted fusion feature (F att Memory enhancement features per modality by reflecting the temporal context in ) ) generates. And, the dimensionality reduction module (244) uses a deep learning model-based encoder to generate modality-specific memory enhancement features ( By reducing the dimensionality of ) fusion features (F reduced Creates ).
[0139] Step S350 is the event detection step.
[0140] The intelligent security and risk response system (100) has a convergence feature (F reduced The type of event present in the modality data (210) is determined through spatiotemporal analysis of ).
[0141] The event detection module (250) of the intelligent security and risk response system (100) has a fusion feature (F reduced By inputting ) into an already trained deep learning model, fusion features (F reduced Extracts a spatiotemporal pattern from ), and determines the event type based on the spatiotemporal pattern.
[0143] Step S360 is the situation understanding step.
[0144] The intelligent security and risk response system (100) has a convergence feature (F reduced Based on the event type, situation context information (270) including the context and cause of the situation is generated through inference using a deep learning-based situation understanding module (260).
[0145] The Situation Understanding Module (260) is a convergence feature (F reduced It infers the context and cause of the situation based on ) and event types. That is, the situation understanding module (260) infers the fusion features (F reduced ) and event types are input into an already trained deep learning model to generate situation context information (270). The deep learning model may include any one of ST-GNN, Transformer, Vision Language Model (VLM), or a combination thereof.
[0147] S370 is the situation response phase.
[0148] The intelligent security and risk response system (100) selects a response scenario and action items based on situation context information (270), and executes an automated response by linking with an external physical system based on the action items. For example, the action items may include broadcasting an emergency, issuing an on-site alarm, sending a dispatch request message to a safety officer terminal, automatically closing a gas valve, tracking CCTV, and logging the situation.
[0150] The method of operation of the aforementioned intelligent security and risk response system has been described with reference to the flowchart presented in the drawings. For simplicity of explanation, the method has been illustrated and described in a series of blocks; however, the present invention is not limited to the order of said blocks, and some blocks may occur in a different order or simultaneously with other blocks as illustrated and described herein, and various other branches, flow paths, and sequences of blocks that achieve the same or similar results may be implemented. Furthermore, not all illustrated blocks may be required for the implementation of the method described herein.
[0152] Meanwhile, in the description with reference to FIG. 3, each step may be further divided into additional steps or combined into fewer steps according to an embodiment of the present invention. Also, some steps may be omitted as necessary, and the order between steps may be changed. Furthermore, even if other omitted details are included, the contents of FIG. 1 and 2 may be applied to the contents of FIG. 3. Also, the contents of FIG. 3 may be applied to the contents of FIG. 1 and 2.
[0154] Although the present invention has been described above with reference to preferred embodiments, those skilled in the art will understand that various modifications and changes can be made to the invention without departing from the spirit and scope of the invention as described in the following claims. Explanation of the symbols
[0156] 100: Intelligent Security and Risk Response System 110: Processor 120: Communication device 130: Memory 140: Storage device 150: Input interface device 160: Output interface device 170: Bus
Claims
Claim 1 A method performed by an intelligent security and risk response system, comprising: a step of generating multiple feature vectors by passing multimodal data through a deep learning-based feature extractor for each modality; a step of generating fused features by dynamically analyzing and processing the multiple feature vectors using a feature fusion module, which is a deep learning-based adaptive modality fusion engine; a step of determining an event type through spatiotemporal analysis of the fused features; a step of generating situational context information including the context and cause of a situation through inference using a deep learning-based situational understanding module based on the fused features and the event type; and a step of selecting a response scenario and a measure item according to the situational context information, and executing an automated response by linking with an external physical system according to the measure item, wherein the situational understanding module includes a vision language model. Claim 2 In claim 1, the multimodal data is one of image, audio data, and sensor data, or a combination thereof. Method of operation of an intelligent security and risk response system. Claim 3 In claim 1, the plurality of feature vectors is one of image feature vectors, acoustic feature vectors, and time series feature vectors, or a combination thereof. Method of operation of an intelligent security and risk response system. Claim 4 A method of operation of an intelligent security and risk response system, wherein the feature fusion module comprises: a feature alignment module that aligns the plurality of feature vectors based on time and space and maps the aligned feature vectors to a common representation space to generate aligned features; and a feature attention module that generates attention-weighted fusion features by assigning weights to some information included in the aligned features using either cross attention and multihead attention or a combination thereof. Claim 5 A method of operation of an intelligent security and risk response system, wherein, in paragraph 4, the feature fusion module comprises: a modality memory module that generates modality-specific memory-enhanced features by reflecting temporal context to the attention-weighted fusion features; and a dimensionality reduction module that generates fusion features by reducing the dimensionality of the modality-specific memory-enhanced features using a deep learning model-based encoder. Claim 6 A method of operation of an intelligent security and risk response system, wherein the step of determining the event type in claim 1 is to input the fusion feature into a deep learning model that has already been trained to extract a spatiotemporal pattern from the fusion feature, and to determine the event type based on the spatiotemporal pattern. Claim 7 delete Claim 8 An intelligent security and risk response system comprising: a processor; and a memory storing one or more commands executed through the processor, wherein the one or more commands include: a command to generate multiple feature vectors by passing multimodal data through a modality-specific deep learning-based feature extractor; a command to generate fused features by dynamically analyzing and processing the multiple feature vectors using a feature fusion module, which is a deep learning-based adaptive modality fusion engine; a command to determine an event type through spatiotemporal analysis of the fused features; a command to generate situational context information including the context and cause of a situation through inference using a deep learning-based situational understanding module based on the fused features and the event type; and a command to select a response scenario and action item according to the situational context information, and to execute an automated response by linking with an external physical system according to the action item, wherein the situational understanding module includes a vision language model. Claim 9 In paragraph 8, the multimodal data is any one of image, audio data, and sensor data, or a combination thereof, in an intelligent security and risk response system. Claim 10 In claim 8, the plurality of feature vectors is any one of image feature vectors, acoustic feature vectors, and time-series feature vectors, or a combination thereof, in an intelligent security and risk response system. Claim 11 An intelligent security and risk response system according to claim 8, wherein the feature fusion module comprises: a feature alignment module that aligns the plurality of feature vectors based on time and space and maps the aligned feature vectors to a common representation space to generate aligned features; and a feature attention module that generates attention-weighted fusion features by assigning weights to some information included in the aligned features using either cross attention and multihead attention or a combination thereof. Claim 12 An intelligent security and risk response system according to claim 11, wherein the feature fusion module comprises: a modality memory module that generates modality-specific memory-enhanced features by reflecting temporal context to the attention-weighted fusion features; and a dimensionality reduction module that generates fusion features by reducing the dimensionality of the modality-specific memory-enhanced features using a deep learning model-based encoder. Claim 13 An intelligent security and risk response system, wherein, in claim 8, the command for determining the event type includes a command for inputting the fusion feature into a deep learning model that has already learned the fusion feature to extract a spatiotemporal pattern from the fusion feature, and determining the event type based on the spatiotemporal pattern. Claim 14 delete