System and method for intelligent distributed CCTV using multitasking-based video analysis and visual question answering

The hybrid multitasking structure in CCTV systems optimizes performance by processing large-scale language models in the cloud and real-time video analysis on-device AI, addressing latency and cost issues while enhancing situational awareness.

KR102996983B1Active Publication Date: 2026-07-29PIASPACE CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
PIASPACE CO LTD
Filing Date
2025-01-02
Publication Date
2026-07-29

AI Technical Summary

Technical Problem

Current CCTV systems face limitations in recognizing complex situations and require significant computational resources for real-time video analysis, leading to increased costs and network latency, which hampers rapid response during emergencies.

Method used

A hybrid multitasking structure that processes large-scale language models in the cloud and real-time video analysis on-device AI, utilizing local servers for low-latency tasks, and integrating multimodal models for efficient image and text data processing.

Benefits of technology

Enables real-time, accurate, and cost-effective video analysis with reduced network load, allowing stable monitoring and advanced situational awareness in large-scale CCTV networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure R1020250000093_ABST
    Figure R1020250000093_ABST
Patent Text Reader

Abstract

An intelligent distributed CCTV system using multitasking-based video analysis and VQA according to one embodiment may include: an on-device terminal that extracts video features based on a video embedding model from a patch cropped from a frame of a video stream and extracts a video feature vector by embedding the video features; a cloud server that extracts text features based on a text embedding model from a prompt containing content to be detected in the video stream and extracts a text feature vector by embedding the text features; and a local server that calculates the similarity between the video feature vector and the text feature vector and determines a patch in which the similarity is greater than or equal to a preset threshold as a patch containing content to be detected according to the prompt.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to a CCTV system and to a distributed processing system capable of optimizing video analysis performance and reducing costs by processing a large-scale language model in the cloud and processing real-time video analysis tasks with on-device AI. Background Technology

[0003] CCTV systems are widely used today for various security purposes. In particular, CCTV systems play a role in detecting and responding to abnormal situations by capturing video in real time in public places, commercial facilities, traffic management, and industrial sites.

[0004] Meanwhile, recent CCTV systems are evolving beyond simply capturing and storing video into sophisticated systems that utilize artificial intelligence technology to automatically analyze footage and detect abnormal situations. These AI-based CCTV systems are now able to provide more sophisticated security solutions by processing complex video data, including object detection, behavior analysis, and anomaly detection.

[0005] In particular, current CCTV systems primarily utilize object detection and anomaly detection technologies based on predefined rules. Through this, they perform security tasks by detecting specific objects appearing on the screen or movements occurring within designated areas. These CCTV systems operate by detecting people, vehicles, and objects in the video in real time, generating warnings based on certain conditions, or providing notifications when predefined events occur. This approach demonstrates a certain level of effectiveness in preventing and responding to crimes or accidents that may occur within the areas monitored by the CCTV.

[0006] However, there are some limitations to the current CCTV system.

[0007] First, since object detection and anomaly detection primarily rely on fixed rules, there are limitations in recognizing complex situations or responding flexibly to various variables. For instance, existing systems often fail to function properly in unstructured and complex situations such as violence between people, accidents, or fires. This is because analysis is based solely on the appearance or movement of simple objects, resulting in a lack of capability to recognize complex situations.

[0008] Furthermore, large-scale computing resources are required for the real-time analysis of CCTV video data. In particular, processing high-resolution video in real time demands significant computational performance and memory. This issue becomes even more pronounced when operating large-scale CCTV networks; processing video from tens to hundreds of CCTV cameras in real time requires high-performance servers and large-capacity storage, which can lead to a rapid increase in network and equipment maintenance costs.

[0009] Under these circumstances, processing data using a centralized structure may result in delays in real-time anomaly detection. For instance, transmitting all video data to a central server for processing increases network load, and particularly if network latency is prolonged, immediate response in emergency situations can become difficult. This issue acts as a significant disadvantage, especially in environments requiring rapid response from CCTV systems during emergencies such as crimes or accidents.

[0010] Therefore, improvements are necessary to complement the limitations of the current CCTV system and develop it into a more sophisticated and efficient system. Prior art literature

[0012] Republic of Korea Published Patent Application No. 10-2023-0039934 The problem to be solved

[0013] The present invention aims to provide a technology that can reduce costs while efficiently performing the task of analyzing situations in real time and recognizing abnormal situations in a CCTV system.

[0014] To this end, the present invention proposes a hybrid multitasking structure in which large-scale language models requiring high-volume computation but infrequent use are processed in the cloud to efficiently utilize high-performance resources, while video analysis requiring real-time processing is handled by on-device AI. This structure focuses on optimizing system performance and minimizing network latency and costs by efficiently allocating appropriate computational resources according to the characteristics of each task.

[0015] Furthermore, the present invention aims to provide a technology that ensures the speed and accuracy of image analysis by performing latency-sensitive tasks on a local server. In particular, by processing tasks requiring significant memory and computational resources through a local server, the invention seeks to prevent processing delays that may occur in large-scale network environments and maximize cost efficiency.

[0016] In addition, the present invention aims to provide a structure that reduces unnecessary data processing time and saves resources by applying a multimodal model to simultaneously analyze image and text data to improve the accuracy of situational awareness, and by prioritizing the processing of important information through object detection and region of interest analysis.

[0017] Accordingly, the present invention aims to optimize the performance of a CCTV system and reduce operating costs, while simultaneously enabling efficient response in environments requiring complex situational awareness.

[0018] Meanwhile, the technical problems of the present invention are not limited to those mentioned above, and other unmentioned technical problems will be clearly understood by a person skilled in the art from the description below. means of solving the problem

[0020] An intelligent distributed CCTV system using multitasking-based video analysis and VQA according to one embodiment may include: an on-device terminal that extracts video features based on a video embedding model from a patch cropped from a frame of a video stream and extracts a video feature vector by embedding the video features; a cloud server that extracts text features based on a text embedding model from a prompt containing content to be detected in the video stream and extracts a text feature vector by embedding the text features; and a local server that calculates the similarity between the video feature vector and the text feature vector and determines a patch in which the similarity is greater than or equal to a preset threshold as a patch containing content to be detected according to the prompt.

[0021] Additionally, the on-device terminal is a device including computing resources capable of collecting video streams from CCTV and performing computations on an AI-based object detection model, the cloud server is a device including high-performance computing resources accessible via the Internet and capable of processing computations on a large-scale language model, and the local server may be a device operating at a location close to the on-device terminal to have relatively less latency compared to communication between the on-device terminal and the cloud server, and capable of performing computations on an AI-based event classification model.

[0022] In addition, the local server acquires the video feature vector and the text feature vector from an external source, and operates in an on-premises environment so that the computation results based on the video feature vector and the text feature vector can be processed within the on-premises environment.

[0023] Additionally, the above prompt includes a plurality of prompts containing content that defines the situation for a plurality of events, and the local server calculates the similarity between a video feature vector derived from one frame and a plurality of text feature vectors derived from the plurality of prompts, and can determine the event defined in the prompt corresponding to the text feature vector with the highest similarity to the video feature vector derived from one frame as an event for one frame.

[0024] In addition, the local server can determine events included in patches with a similarity level above a preset threshold by performing Visual Question Answering (VQA) on patches with a similarity level above a preset threshold in the form of a query preset on the cloud server.

[0025] In addition, the on-device terminal may include an operation of setting a bounding box on an object included in a frame of the video stream based on an object detection model, and cropping the bounding box to generate a patch.

[0026] In addition, the local server can calculate similarity through a scalar product between the video feature vector and the text feature vector.

[0027] In addition, the local server can calculate similarity using the cosine similarity or dot product between the video feature vector and the text feature vector.

[0028] In addition, the local server can classify events for the input patches by inputting patches whose similarity is greater than or equal to a preset threshold into a pre-trained event classification model.

[0029] In addition, the on-device terminal can generate multiple patches by cropping frames of the video stream at preset time intervals.

[0030] A method of operation of an intelligent distributed CCTV system according to one embodiment may include: an on-device terminal acquiring a video stream; an on-device terminal generating a patch by cropping a frame of the video stream; an on-device terminal extracting video features from the patch based on an image embedding model and embedding the video features to extract a video feature vector; a cloud server acquiring a prompt containing content to be detected in the video stream; a cloud server extracting text features from the prompt based on a text embedding model and embedding the text features to extract a text feature vector; a local server calculating a similarity between the video feature vector and the text feature vector; and a local server determining a patch in which the similarity is greater than or equal to a preset threshold as a patch containing content to be detected according to the prompt. Effects of the invention

[0032] According to the above-described embodiment, the present invention enables real-time video analysis and situation recognition to be performed without delay through the combination of a large-scale language model and on-device AI in a CCTV system, thereby optimizing performance and achieving the effect of reducing costs.

[0033] In particular, the present invention efficiently utilizes high-performance resources by processing computations of large-scale language models that require large-scale computations but are used infrequently in the cloud, and ensures real-time responsiveness by minimizing latency through processing video analysis that must be processed in real-time using on-device AI.

[0034] In addition, complex inference tasks requiring significant memory and performance can be computed via a local server, thereby reducing network load and maximizing system processing speed while simultaneously lowering operating costs.

[0035] Furthermore, by simultaneously processing video and text data through multimodal analysis, the present invention enables more accurate recognition and response to complex situations in CCTV footage. In particular, by reducing unnecessary data processing through object detection and region of interest analysis, and by selectively processing only necessary information, system resources can be utilized effectively. As a result, stable real-time monitoring is possible even in large-scale CCTV networks, and advanced situational awareness functions can be provided.

[0036] Therefore, the present invention can implement a high-performance security system capable of faster and more accurate detection of abnormal situations and real-time response by maximizing the performance of a large-scale CCTV system and significantly improving cost efficiency.

[0037] Meanwhile, the effects of the present invention are not limited to those mentioned above, and other unmentioned technical effects will be clearly understood by a person skilled in the art from the description below. Brief explanation of the drawing

[0039] FIG. 1 is a configuration diagram of a CCTV system according to one embodiment. FIG. 2 is a configuration diagram of an on-device terminal, a cloud server, and a local server according to one embodiment. FIG. 3 is an example diagram showing the operations performed and the data processed by the on-device terminal, cloud server, and local server of a CCTV system according to the overall embodiment in simple blocks. FIG. 4 is an example diagram showing the operations performed and the data processed by the on-device terminal, cloud server, and local server of a CCTV system according to the first embodiment in simple blocks. FIG. 5 is an example diagram showing the operations performed and the data processed by the on-device terminal, cloud server, and local server of a CCTV system according to the second embodiment in simple blocks. FIG. 6 is an example diagram illustrating the operation of dividing frames of an image to generate patches according to a Tiled Inference technique according to one embodiment. FIG. 7 is an example diagram illustrating the operation of dividing frames of an image to generate patches according to a Pipeline Inference technique according to one embodiment. FIG. 8 is a flowchart illustrating the steps of operations performed by an on-device terminal, a cloud server, and a local server constituting a CCTV system according to one embodiment. Specific details for implementing the invention

[0040] Detailed information regarding the purpose, technical configuration, and resulting effects of the present invention will be more clearly understood through the following detailed description based on the drawings attached to the specification of the present invention. An embodiment according to the present invention will be described in detail with reference to the attached drawings.

[0041] The embodiments disclosed herein should not be interpreted or used to limit the scope of the invention. It is obvious to those skilled in the art that the description including the embodiments herein has various applications. Accordingly, any embodiments described in the detailed description of the invention are illustrative for better explaining the invention and are not intended to limit the scope of the invention to the embodiments.

[0042] The functional blocks shown in the drawings and described below are merely examples of possible implementations. In other implementations, other functional blocks may be used without departing from the spirit and scope of the detailed description. Additionally, while one or more functional blocks of the present invention are shown as individual blocks, one or more of the functional blocks of the present invention may be a combination of various hardware and software configurations that perform the same function.

[0043] Furthermore, the expression that it includes certain components is an “open-ended” expression that merely refers to the existence of such components and should not be understood as excluding additional components.

[0044] Furthermore, when it is stated that one component is “connected” or “joined” to another component, it should be understood that while it may be directly connected or joined to that other component, there may also be other components present in between.

[0045] Hereinafter, various embodiments of the present invention are described with reference to the accompanying drawings. However, this is not intended to limit the present invention to specific embodiments and should be understood to include various modifications, equivalents, and / or alternatives of the embodiments of the present invention.

[0046] The present invention proposes a distributed CCTV system that combines a large-scale language model in the cloud with on-device AI. Specifically, the invention proposes a distributed CCTV system that reduces the overall network load by having the cloud process large-scale computations of the large-scale language model, having the on-device AI process real-time video analysis, and having local servers perform tasks requiring low latency and complex inference. Furthermore, the present invention proposes a technology that enables stable real-time monitoring and advanced situational awareness even in large-scale CCTV networks by simultaneously processing video and text data through multimodal analysis and efficiently utilizing system resources by excluding unnecessary data.

[0047] Below, we will examine the configuration of the CCTV system of the present invention and the operation of each component.

[0048] FIG. 1 is a configuration diagram of a CCTV system (10) (hereinafter referred to as 'system (10)') according to one embodiment.

[0049] Referring to FIG. 1, the system (10) may include an on-device terminal (100), a cloud server (200), and a local server (300).

[0050] The on-device terminal (100) is a local device connected to a video recording device such as a CCTV. For example, the on-device terminal (100) can collect a video stream in real time to perform tasks such as object detection and tracking. The on-device terminal (100) may include computing resources with specifications capable of performing computations of an artificial intelligence-based object detection model to process tasks requiring real-time analysis and to process tasks requiring low latency on-device.

[0051] The cloud server (200) is a device that stores and controls a large-scale language model based on a Transformer model and processes operations for prompts or questions requested by an external device via a communication network. For example, the cloud server (200) can process operations of the Transformer model for the requested prompt or question. The cloud server (200) may include high-performance computing resources that are accessible via the Internet and have specifications capable of processing operations of the large-scale language model.

[0052] The local server (300) is a local server (300) that operates in a location close to the on-device terminal (100) to have relatively less latency compared to communication between the on-device terminal (100) and the cloud server (200), and performs AI-based computations. For example, the local server (300) can relay the on-device terminal (100) and the cloud server (200) to process data collected in real time from the on-device terminal (100). For example, the local server (300) can process specific computations by comparing information transmitted from the on-device terminal (100) with information transmitted from the cloud server (200). The local server (300) is operated in an on-premises environment and can process computation results regarding information collected from the outside within the on-premises environment.

[0053] The specific configuration of the on-device terminal (100), cloud server (200), and local server (300) of the system (10) according to an embodiment of the present invention is as shown in FIG. 2 below.

[0054] FIG. 2 is a configuration diagram of an on-device terminal (100), a cloud server (200), and a local server (300) according to one embodiment.

[0055] Referring to FIG. 2, an on-device terminal (100), a cloud server (200), and a local server (300) according to one embodiment may each include a memory (110), a processor (120), an input / output interface (130), and a communication interface (140).

[0056] The memory (110) can store data obtained from an external device or data generated by itself. The memory (110) can store instructions that can perform operations of the processor (120). For example, the memory of the on-device terminal (100) can store video streams and object detection models. Additionally, the cloud server (200) can store a large-scale language model based on a Transformer model. Additionally, the local server (300) can store an event classification model.

[0057] The processor (120) is a computing device that controls overall operation. The processor (120) can execute instructions stored in memory (110). The operation of the on-device terminal (100), cloud server (200), and local server (300) according to an embodiment of the present invention can be understood as an operation performed by the processor (120).

[0058] The input / output interface (130) may include a hardware interface or a software interface for inputting or outputting information.

[0059] The communication interface (140) enables the transmission and reception of information through a communication network. To this end, the communication interface (140) may include a wireless communication module or a wired communication module.

[0060] The on-device terminal (100), cloud server (200), and local server (300) can be implemented as various types of devices capable of performing calculations through a processor (120) and transmitting and receiving information through a network. For example, they can be implemented in the form of a computer device, a portable communication device, a smartphone, a portable multimedia device, a laptop, a tablet PC, etc., but are not limited to these examples.

[0061] FIG. 3 is an example diagram showing the operations performed and the data processed by the on-device terminal (100), cloud server (200), and local server (300) of the system (10) according to the overall embodiment in simple blocks. If the blocks of FIG. 3 are divided according to the order of operation of the system, they can be divided into FIG. 4 according to the first embodiment and FIG. 5 according to the second embodiment.

[0062] FIG. 4 is an example diagram showing the operations performed and the data processed by the on-device terminal (100), cloud server (200), and local server (300) of the system (10) according to the first embodiment in simple blocks.

[0063] First, we will examine the operation performed by the on-device terminal (100) among the large unit blocks shown in FIG. 4.

[0064] Referring to FIG. 4, the on-device terminal (100) can acquire a video stream. For example, the video stream can be collected from a video recording device such as a CCTV.

[0065] The on-device terminal (100) can create a patch by cropping a frame included in the acquired video stream. The reason for dividing the video frame into a patch of a smaller size is as follows. If the entire video frame is processed to be analyzed at once in the operation described below, it is difficult to detect important objects or events when they are located in a local part of the frame. Accordingly, an embodiment of the present invention divides the video frame into a patch set to a smaller size, thereby enabling detailed analysis without missing important information that may occur in a local area.

[0066] Additionally, the on-device terminal (100) may generate multiple patches by cropping the video stream frame by frame at preset time intervals. These time intervals can be adjusted according to the speed of the video stream or the characteristics of the object to be analyzed, thereby enabling the on-device terminal (100) to analyze actions or events occurring at specific times more precisely. Furthermore, after multiple patches are generated, these patches can be processed in parallel in subsequent stages, thereby further improving the analysis speed of the entire system.

[0067] To this end, an embodiment of the present invention can create a patch as shown in the following FIGS. 6 and FIGS. 7.

[0068] FIG. 6 is an example diagram illustrating the operation of dividing frames of an image to generate patches according to a Tiled Inference technique according to one embodiment.

[0069] Referring to FIG. 6, the on-device terminal (100) can generate multiple patches by dividing a video frame into a predetermined size. For example, as can be seen in FIG. 6, the on-device terminal (100) can divide a video frame into six patches. When features are extracted from the patches divided in this way, important information can be extracted even in situations where an important object exists only in a small area of ​​the screen or a local event occurs.

[0070] FIG. 7 is an example diagram illustrating the operation of dividing frames of an image to generate patches according to a Pipeline Inference technique according to one embodiment.

[0071] Referring to FIG. 7, the on-device terminal (100) can first detect an object included in a video frame and secondarily crop a region of interest (ROI) containing the detected object to create a patch. FIG. 7 visually illustrates this process and shows a method of first setting a region of interest through object detection in a video frame and then extracting only the set region of interest to create a patch.

[0072] For example, an on-device terminal (100) can generate a patch by cropping the bounding box set on an object detected in a frame of a video stream based on an object detection model such as YOLO. As can be seen in FIG. 7, bounding boxes can be set for various objects within a video frame based on an object detection model, and the on-device terminal (100) can set the bounding box as a region of interest. Accordingly, the on-device terminal (100) can remove unnecessary non-interest regions, excluding the bounding box, from the analysis target. When extracting features from such a divided patch, it is more efficient than processing the entire frame, and by focusing on analyzing only the core information, detection accuracy and processing speed can be increased simultaneously.

[0073] After the patch generation process described above, the on-device terminal (100) can extract video features from the patch based on an image embedding model and extract a video feature vector by embedding the extracted video features. For example, the image embedding model can extract features of objects within the video by cutting the patch into fixed sizes and sequentially encoding each cut patch, and generate a video feature vector by embedding spatial and temporal information of each feature.

[0074] For example, an image embedding model may include a ViT (Vision Transformer) model. Here, the ViT model is a deep learning model used for image classification and feature extraction, which applies the Transformer structure to image processing. Although the Transformer model originated in natural language processing, the ViT model introduces the Transformer structure to image processing, allowing it to take patches within an image as input, extract features, and perform various analyses based on them.

[0075] The specific process by which the ViT model extracts video feature vectors from patches is as follows.

[0076] Before processing the input first patch (referring to the patch generated according to the ROI settings as the 'first patch'), the ViT model divides the first patch into smaller second patches (referring to the patch into which the ViT model further divided the first patch as the 'second patch'). For example, if the first patch of size 224x224 is divided into second patches of size 16x16, a total of 196 patches can be generated from the first patch.

[0077] Through this process, the ViT model divides the first patch into smaller second patches and processes them individually, instead of processing the entire first patch. This enables the ViT model to extract important information even from smaller areas.

[0078] Each second patch divided from the first patch contains a 2D image. To convert each second patch into a vector, ViT flattens the pixel data of the second patch to create a one-dimensional vector. Then, a learnable linear embedding is applied to each second patch to convert it into a vector of a fixed size.

[0079] For example, if the second patch size is 16x16 and is an RGB image, each second patch is converted into a vector of 16x16x3=768 dimensions. Through this, each second patch is represented as an embedding vector of a fixed size (e.g., 768 dimensions), and this embedding vector contains the visual features of the corresponding second patch.

[0080] The ViT model may additionally use position embedding to preserve the location information of the second patch. Position embedding serves to recognize where the second patch was located within the first patch.

[0081] Position embeddings are vectors representing the relative positional information of the second patch, which are added to each second patch embedding vector to preserve spatial information. Through this, the ViT model becomes able to understand the structure of the entire image while considering the position of each patch.

[0082] The ViT model inputs the second patch embedding vectors into the encoder layer of the Transformer. The encoder layer learns the interactions between each second patch through a self-attention mechanism. Self-attention learns how each second patch is associated with other second patches to highlight important information. The encoder layer consists of multiple layers, and in each layer, the information between the second patches is increasingly abstracted and features are summarized.

[0083] Accordingly, the ViT model can output a feature vector representing the overall features of the first patch by efficiently extracting and integrating visual features within each second patch. In the present invention, this feature vector is referred to as a 'video feature vector'. The video feature vector is a vector summarizing important visual information within the input first patch and can be utilized for various tasks such as image classification, object detection, and anomaly detection. The video feature vector is output as a vector of a fixed size (e.g., 768 dimensions), and each element may correspond to a specific visual attribute of the image.

[0084] As described above, the video feature vector generated by the on-device terminal (100) can be used for calculating similarity with the text feature vector generated from the prompt according to the operation to be described later, multi-label classification by an event classification model, and detection of a specific event by an event classification model. A detailed explanation of the details will be provided later.

[0085] Next, we will examine the operation performed by the cloud server (200) among the large unit blocks shown in FIG. 4.

[0086] Referring to FIG. 4, the cloud server (200) may obtain a prompt containing content to be detected in a video stream. For example, the prompt may include a text prompt in a pre-configured format according to the purpose of detection for the video stream. For example, the prompt may be configured by the designer or user of the system (10). For example, the prompt may include a text description defined to detect specific behaviors or events such as violence, fire, or unauthorized entry (e.g., "Detects violent situation").

[0087] The cloud server (200) can extract text features from a prompt based on a text embedding model and extract a text feature vector by embedding the text features. For example, the text embedding model can tokenize the prompt word by word to extract features and generate a text feature vector by embedding the semantic relationships of each word.

[0088] For example, a text embedding model may include a Transformer model. Here, the Transformer model may include a large-scale language model such as BERT (Bidirectional Encoder Representations from Transformers). The Transformer model takes text data as input, extracts its semantic features, and converts them into vectors.

[0089] The specific process by which the Transformer model extracts feature vectors from the prompt is as follows.

[0090] The prompt may contain text describing the context of the video that the system needs to analyze. For example, there may be prompts entered in natural language, such as "violence detection" or "fire detection." This prompt is input into a Transformer model, which learns the semantic relationships of the text to understand the connections between each word.

[0091] The Transformer model first divides the prompt into tokens before processing text input. This tokenization is the process of breaking down a sentence into individual words or subwords. Through this, the Transformer model is able to recognize and process each word as an independent unit.

[0092] Each word in the tokenized prompt can be vectorized through an embedding layer. This process converts each word into a fixed-size vector. The Transformer model represents each word as an array of numbers, enabling it to learn semantic relationships between words.

[0093] In particular, the Transformer model utilizes a Self-Attention mechanism to identify how each word in an input prompt interacts with other words within a sentence and what role each word plays within the context. As such, each word in the prompt can be transformed into a vector reflecting contextual information through the Self-Attention mechanism. The encoder layer of the Transformer model learns these word vectors more deeply and can generate higher-level feature vectors by abstracting semantic relationships within the context. The encoder layer consists of multiple layers, and as it passes through each layer, the semantic relationships between words can be learned more clearly.

[0094] Accordingly, the Transformer model can extract a feature vector representing the overall meaning of the prompt. In the present invention, this feature vector is referred to as a 'text feature vector'. The text feature vector is a feature vector that summarizes the semantic information of the text included in the prompt. The text feature vector is output as a vector of a fixed size (e.g., 512 dimensions), and each element can correspond to a semantic attribute of the prompt.

[0095] In this way, the text feature vector generated by the cloud server (200) can be used to detect specific events in a video stream based on the calculation of similarity with the video feature vector according to the operation to be described later. An explanation of this content will be described later in the operation of the local server (300).

[0096] Next, we will examine the operations performed by the local server (300) among the large unit blocks shown in FIG. 4.

[0097] Referring to FIG. 4, the local server (300) can calculate the similarity between a video feature vector generated by an on-device terminal (100) and a text feature vector generated by a cloud server (200). For example, the local server (300) can calculate the similarity through a scalar product between the video feature vector and the text feature vector. For example, techniques for performing a scalar product to calculate similarity may include cosine similarity and a dot product.

[0098] Through this, the local server (300) measures how similar the two feature vectors are in direction. In this case, the higher the similarity value, the more similar the situation is to the content described in the prompt. That is, the local server (300) can determine, through similarity calculation, which patch within the video frame is likely to have an event related to the prompt.

[0099] The local server (300) can determine that a patch whose calculated similarity is greater than or equal to a preset threshold is a patch containing the content that the prompt intends to detect. That is, the local server (300) can determine that a patch whose similarity value exceeds a preset threshold is likely to contain the content that the prompt intends to detect (e.g., violence, intrusion, etc.).

[0100] At this time, the local server (300) can determine more specifically what event the detection content included in the patch corresponds to. For example, the local server (300) can input the patch into an event classification model, classify the events of the patch based on the class output by the event classification model, and output a notification about an abnormal situation.

[0101] For example, an event classification model can be implemented based on the Resnet (Residual Network) structure.

[0102] Here, ResNet is a Convolutional Neural Network model that performs image classification in deep learning. ResNet is a model that introduces Residual Connections to enable effective training even in very deep networks.

[0103] The event classification model is supervised-learned based on Resnet to identify specific events from images, and can output a class for a single label or multiple labels for an input patch.

[0104] The event classification model, based on the ResNet structure, can be trained using a supervised learning approach to classify the classes of events contained within input images. Through supervised learning, the model learns how to recognize specific events (e.g., violence, fire, accidents, etc.) using a pre-trained dataset. Consequently, it becomes possible to determine which event occurred in a given patch based on the image information contained within that patch.

[0105] In this case, the event classification model can perform single-label or multi-label classification. In the single-label classification method, it predicts events corresponding to a single class for an input patch and can output the most important events within that patch. Additionally, in the multi-label classification method, considering that multiple events may occur simultaneously within a patch, it can predict events corresponding to multiple classes at the same time. Through this, effective analysis can be performed even in complex scenes or situations where multiple events occur.

[0106] For example, a multi-label classification model can simultaneously detect two events within a patch: a person running and violent behavior occurring at the same time. The event classification model can accurately classify these multiple events and output the class corresponding to each event.

[0107] Additionally, the prompt input to the cloud server (200) in FIG. 4 may include an embodiment that includes a plurality of prompts containing content that defines the situation for a plurality of events. In this case, the cloud server (200) may derive a text feature vector for each of the plurality of prompts. Accordingly, the local server (300) may calculate the similarity between a video feature vector derived from a single frame obtained from the on-device terminal (100) and a plurality of text feature vectors derived from a plurality of prompts obtained from the cloud server (200), and determine the event defined in the prompt corresponding to the text feature vector with the highest similarity to the video feature vector derived from a single frame as an event for a single frame.

[0108] The local server (300) can comprehensively analyze the situation of the patch based on the above embodiments and generate a warning message or notification according to the identified event.

[0109] FIG. 5 is an example diagram showing, in simple blocks, the operations performed and data processed by the on-device terminal (100), cloud server (200), and local server (300) of the system (10) according to the second embodiment. The second embodiment corresponds to an embodiment that can be additionally implemented after the first embodiment described above has been performed. Duplicate descriptions of the contents already explained in FIG. 4 among the blocks of FIG. 5 are omitted.

[0110] Referring to FIG. 5, the local server (300) can determine events included in patches with a similarity level greater than or equal to a preset threshold by performing Visual Question Answering (VQA) in the form of a preset query on the cloud server (200).

[0111] Here, VQA (Visual Question Answering) refers to visual question answering, an artificial intelligence technology that provides natural language answers to questions related to a given image or video.

[0112] For example, if the similarity of a specific patch in the local server (300) is greater than or equal to a preset threshold, the local server (300) can perform VQA query-response for the patch through interaction with the cloud server (200). To do this, the local server (300) can send the patch extracted from the video stream and a predefined query to the large-scale language model of the cloud server (200) to analyze events for the specific patch extracted from the video stream. The cloud server (200) can process the query using a Transformer model to determine the events contained in the patch. For example, if the question "What is the situation in the video?" is given to the cloud server (200), the Transformer model can convert the given question into a question text feature vector, and the given video (patch) can be converted into a video feature vector through the Transformer model (e.g., a Transformer model with a ViT structure), and then combine these two features to generate an answer to the question. Through this process, the local server (300) can obtain an answer such as ‘a person has fallen’ from the Transformer model of the cloud server (200).

[0113] Additionally, the local server (300) can derive a sentence describing a scene in the video by utilizing a predefined prompt through a comparison of similarity between a video feature vector and a text vector. The local server (300) can generate a sentence describing a scene in the video through the predefined prompt, and can perform VQA work by using this to ask a query about the video scene and generate an appropriate answer.

[0114] FIG. 8 is a flowchart illustrating the steps of operations performed by an on-device terminal (100), a cloud server (200), and a local server (300) that constitute a system (10) according to one embodiment. The operations of the on-device terminal (100), the cloud server (200), and the local server (300) that constitute the embodiment of FIG. 8 can be understood as operations performed by a processor (120) provided in each.

[0115] Each step disclosed in FIG. 8 is merely a preferred embodiment for achieving the purpose of the present invention, and some steps may be added or deleted as needed, and any one step may be included in another step. The order of each operation disclosed in FIG. 8 is arranged for ease of understanding only, and such order is not limited to a chronological order, and the order may be changed and operated differently according to the designer's choice.

[0116] Referring to FIG. 8, in step S1010, the on-device terminal (100) can acquire a video stream.

[0117] In step S1020, the on-device terminal (100) can generate a patch of cropped frames of the video stream.

[0118] In step S1030, the on-device terminal (100) can extract video features from a patch based on a ViT (Vision Transformer) model and embedding the video features to extract a video feature vector.

[0119] In step S1040, the cloud server (200) can obtain a prompt containing content to be detected in the video stream.

[0120] In step S1050, the cloud server (200) can extract text features from the prompt based on the Transformer model and embed the text features to extract a text feature vector.

[0121] In step S1060, the local server (300) can calculate the similarity between the video feature vector and the text feature vector.

[0122] In step S1070, the local server (300) can determine that a patch with a similarity greater than or equal to a preset threshold is a patch containing content to be detected according to the prompt.

[0123] Meanwhile, since the specific operations performed by the on-device terminal (100), cloud server (200), and local server (300) in each step of FIG. 8 have been described in detail together with FIG. 1 to 7, a redundant description will be omitted.

[0124] According to the above-described embodiment, the present invention enables real-time video analysis and situation recognition to be performed without delay through the combination of a large-scale language model and on-device AI in a CCTV system, thereby optimizing performance and achieving the effect of reducing costs.

[0125] In particular, the present invention ensures real-time responsiveness by processing large-scale language models that require large-scale computations but are used infrequently in the cloud, thereby efficiently utilizing high-performance resources, and by processing video analysis that must be processed in real-time using on-device AI, thereby minimizing latency. Additionally, complex inference tasks that require a lot of memory and performance can be computed through a local server (300), thereby reducing network load and maximizing the processing speed of the system while simultaneously reducing operating costs.

[0126] Furthermore, by simultaneously processing video and text data through multimodal analysis, the present invention enables more accurate recognition and response to complex situations in CCTV footage. In particular, by reducing unnecessary data processing through object detection and region of interest analysis, and by selectively processing only necessary information, system resources can be utilized effectively. As a result, stable real-time monitoring is possible even in large-scale CCTV networks, and advanced situational awareness functions can be provided.

[0127] Therefore, the present invention can implement a high-performance security system capable of faster and more accurate detection of abnormal situations and real-time response by maximizing the performance of a large-scale CCTV system and significantly improving cost efficiency.

[0128] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more items unless the relevant context clearly indicates otherwise.

[0129] In this document, each of the phrases such as “A or B,” “at least one of A and B,” “at least one of A or B,” “A, B or C,” “at least one of A, B and C,” and “at least one of A, B, or C” may include all possible combinations of items listed together in the corresponding phrase. Terms such as “1,” “2,” or “first” or “second” may be used simply to distinguish a component from another component and do not limit the components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as “coupled” or “connected” to another (e.g., 2nd) component, with or without the terms “functionally” or “communicationly,” it means that the component may be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0130] As used in this document, the term "module" may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be a component formed integrally, or a minimum unit of a component or part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0131] Various embodiments of this document may be implemented as software (e.g., a program) comprising one or more instructions stored in a storage medium (e.g., memory) that can be read by a device (e.g., an electronic device). The storage medium may include random access memory (RAM), a memory buffer, a hard drive, a database, erasable programmable read-only memory (EPROM), electrically erasable read-only memory (EEPROM), read-only memory (ROM), and / or the like.

[0132] Additionally, the processor of the embodiments of this document may call at least one instruction among one or more instructions stored from a storage medium and execute it. This enables the device to operate to perform at least one function according to at least one called instruction. Such one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The processor may be a general-purpose processor, a Field Programmable Gate Array (FPGA), an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), and / or the like.

[0133] A device-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.

[0134] Methods according to the various embodiments disclosed in this document may be provided as part of a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as a manufacturer's server, an application store's server, or the server's memory.

[0135] According to various embodiments, each component (e.g., module or program) of the described components may include a singular or multiple entities. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the components of the multiple components in the same or similar manner as they were performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically; one or more of the operations may be executed in a different order; omitted; or one or more other operations may be added. Explanation of the symbols

[0137] 10: CCTV System 100: On-device terminal 200: Cloud Server 300: Local Server 110: Memory 120: Processor 130: Input / Output Interface 140: Communication interface

Claims

Claim 1 In an intelligent distributed CCTV system utilizing multitasking-based video analysis and VQA, an on-device terminal extracts video features based on a video embedding model from a patch cropped from a frame of a video stream and extracts a video feature vector by embedding said video features, wherein the on-device terminal generates two or more patches from a single frame; and a cloud server extracts text features based on a text embedding model from a prompt containing content to be detected in said video stream and extracts a text feature vector by embedding said text features; A system comprising a local server that calculates the similarity between the video feature vector and the text feature vector and determines, among the two or more patches, a patch in which the similarity is greater than or equal to a preset threshold as a patch containing content to be detected according to the prompt, wherein the local server is operated in a location close to the on-device terminal to have relatively less latency compared to communication between the on-device terminal and the cloud server, performs calculations of an artificial intelligence-based event classification model, acquires the video feature vector and the text feature vector from an external source, and operates in an on-premises environment to process calculation results based on the video feature vector and the text feature vector within the on-premises environment. Claim 2 A system according to claim 1, wherein the on-device terminal is a device comprising computing resources capable of collecting video streams from a CCTV and performing computations on an artificial intelligence-based object detection model, and the cloud server is a device comprising high-performance computing resources accessible via the Internet and capable of processing computations on a large-scale language model. Claim 3 delete Claim 4 A system according to claim 1, wherein the prompt includes a plurality of prompts containing content that defines the situation for a plurality of events, and the local server calculates the similarity between a video feature vector derived from one frame and a plurality of text feature vectors derived from the plurality of prompts, and determines the event defined in the prompt corresponding to the text feature vector with the highest similarity to the video feature vector derived from one frame as an event for one frame. Claim 5 A system according to claim 1, wherein the local server performs Visual Question Answering (VQA) in the form of a query pre-set on the cloud server for patches whose similarity is above a pre-set threshold, and determines events included in patches whose similarity is above the pre-set threshold. Claim 6 A system according to claim 1, wherein the on-device terminal includes the operation of setting a bounding box on an object included in a frame of the video stream based on an object detection model, and cropping the bounding box to generate a patch. Claim 7 In claim 1, the local server is a system that calculates similarity through a scalar product between the video feature vector and the text feature vector. Claim 8 In claim 7, the local server is a system that calculates similarity using the cosine similarity or dot product between the video feature vector and the text feature vector. Claim 9 A system according to claim 1, wherein the local server includes the operation of classifying events for an input patch by inputting a patch having a similarity greater than or equal to a preset threshold into a pre-trained event classification model. Claim 10 In claim 1, the on-device terminal generates a plurality of patches by cropping frames of the video stream at preset time intervals. Claim 11 A method of operation for an intelligent distributed CCTV system comprises: an operation in which an on-device terminal acquires a video stream; an operation in which the on-device terminal generates a patch by cropping a frame of the video stream, thereby generating two or more patches from a single frame; an operation in which the on-device terminal extracts video features from the patch based on an image embedding model and embeds the video features to extract a video feature vector; an operation in which a cloud server acquires a prompt containing content to be detected in the video stream; an operation in which the cloud server extracts text features from the prompt based on a text embedding model and embeds the text features to extract a text feature vector; and an operation in which a local server calculates the similarity between the video feature vector and the text feature vector. The method comprises an operation in which the local server determines, according to the prompt, that a patch among the two or more patches in which the similarity is greater than or equal to a preset threshold is a patch containing content to be detected, wherein the local server is operated at a location close to the on-device terminal to have relatively less latency compared to communication between the on-device terminal and the cloud server, performs calculations of an artificial intelligence-based event classification model, and acquires the video feature vector and the text feature vector from an external source, and operates in an on-premises environment to process the calculation results based on the video feature vector and the text feature vector within the on-premises environment.