Apparatus and method for generating alarms for text-based images using a video language model
Patent Information
- Application Number
- KR1020240023347
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-02-19
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-02-19
Smart Images

Figure 112024018759549-PAT00001_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to an alarm generation technology, and more specifically, to an apparatus for generating a text-based video alarm using a video language model and a method for the same. Background Technology
[0002] Video alarm technology refers to a device or technology that provides a warning signal to a video monitor or control system when situations such as fire, fighting, or falling occur in recorded or streamed video.
[0003] Existing video alarm technologies have the following disadvantages and limitations. To build a video alarm algorithm, it is necessary to acquire a large amount of true / false data to train the algorithm. This process incurs significant effort, time, and cost in securing such data. For example, to generate a fire alarm from video, a large amount of both fire and non-fire data must be obtained.
[0004] In the case of typical deep learning-based video alarm technology, when a user wants to add a specific alarm, both the training data for the new alarm and the existing alarm training data must be retrained. For example, to add a fight detection algorithm to an existing deep learning-based video fall detection algorithm, the algorithm must be retrained using data for each function. Alternatively, to avoid such retraining, individual algorithms must be built and executed for each alarm occurrence. Consequently, running each algorithm requires an additional amount of computation equal to the number of alarms, which poses difficulties in building real-time video alarms. For instance, there is the problem of having to provide multiple algorithms, such as fall detection, fight detection, and fence-climbing detection algorithms. Prior art literature
[0005] Korean Patent Publication No. 2023-0096464 (Published June 30, 2023) The problem to be solved
[0006] The objective of the present invention is to provide a device for generating a text-based video alarm using a video language model and a method for doing the same. means of solving the problem
[0007] The method for generating a video alarm according to the present invention comprises the steps of: when an alarm text including background text and object text is input by a text input unit, performing embedding on the background text through a video language model to derive a background text embedding vector and performing embedding on the object text to derive an object text embedding vector; when a surveillance video is input by a video input unit, setting the input surveillance video as a background video and using an object detection model to derive an object video in which an area occupied by an object in the input surveillance video is detected; when a surveillance video is input by a video input unit, performing embedding on the background video through a video language model to derive a background video embedding vector and performing embedding on the object video to derive an object video embedding vector; a similarity evaluation unit calculating the similarity between the background text embedding vector and the background video embedding vector and calculating the similarity between the object text embedding vector and the object video embedding vector; and when the calculated similarity exceeds a preset threshold value, an alarm unit generating an alarm corresponding to the alarm text.
[0008] The step of calculating the similarity is characterized in that the similarity evaluation unit calculates the similarity through the vector dot product between the background text embedding vector and the background image embedding vector, and calculates the similarity through the vector dot product between the object text embedding vector and the object image embedding vector.
[0009] The above method comprises, prior to the step of deriving the object text embedding vector, a step in which a text processing unit derives a text embedding vector for a plurality of candidate texts through an image language model; a step in which an image processing unit derives an image embedding vector for a plurality of frames of a plurality of candidate images including a plurality of suitable candidate images and a plurality of unsuitable candidate images through an image language model; a step in which a similarity calculation unit calculates the similarity between the image embedding vector and the text embedding vector; a step in which a classification unit performs clustering on a plurality of frames of the plurality of suitable candidate images according to the similarity to classify them into a harmonious cluster, which is a group with relatively high similarity, and a low-similarity cluster, which is a group with relatively low similarity; a step in which the classification unit classifies a plurality of frames of the plurality of unsuitable candidate images into the low-similarity cluster; a step in which the classification unit performs similarity-based labeling to assign a positive label to frames belonging to the harmonious cluster and a negative label to frames belonging to the low-similarity cluster; and a step in which an analysis unit [represents] that has a positive label assigned to each of the plurality of candidate texts The method further includes a step of deriving a distribution of similarity between frames and frames with assigned voice labels, a step in which the analysis unit selects a threshold value corresponding to a target false alarm rate for classification between frames with assigned positive labels and frames with assigned voice labels based on the distribution corresponding to each of the plurality of candidate texts, and a step in which the analysis unit selects a candidate text among the plurality of candidate texts that exhibits the lowest non-announcement rate relative to the target false alarm rate as an alarm text.
[0010] The above method further includes, prior to the step of deriving text embedding vectors for the plurality of candidate texts through an image language model, a step in which, when an input text is input, the text processing unit derives a plurality of candidate texts associated with the input text through a large language model.
[0011] The device for generating a video alarm according to the present invention comprises: a text input unit that, when an alarm text including background text and object text is input, performs embedding on the background text through a video language model to derive a background text embedding vector and performs embedding on the object text to derive an object text embedding vector; a video input unit that, when a surveillance video is input, sets the input surveillance video as a background video, derives an object video by detecting the area occupied by an object in the input surveillance video using an object detection model, performs embedding on the background video through a video language model to derive a background video embedding vector, and performs embedding on the object video to derive an object video embedding vector; a similarity evaluation unit that calculates the similarity between the background text embedding vector and the background video embedding vector and calculates the similarity between the object text embedding vector and the object video embedding vector; and an alarm unit that generates an alarm corresponding to the alarm text when the calculated similarity exceeds a preset threshold value.
[0012] The above similarity evaluation unit is characterized by calculating similarity through a vector dot product between the background text embedding vector and the background image embedding vector, and calculating similarity through a vector dot product between the object text embedding vector and the object image embedding vector.
[0013] The device comprises: a text processing unit that derives text embedding vectors for a plurality of candidate texts through an image language model; an image processing unit that derives image embedding vectors for a plurality of frames of a plurality of candidate images, including a plurality of suitable candidate images and a plurality of unsuitable candidate images, through an image language model; a similarity calculation unit that calculates the similarity between the image embedding vectors and the text embedding vectors; a classification unit that performs clustering on a plurality of frames of the plurality of suitable candidate images according to the similarity to classify them into a highly similar cluster, which is a group with relatively high similarity, and a low-similarity cluster, which is a group with relatively low similarity, classifies a plurality of frames of the plurality of unsuitable candidate images into the low-similarity cluster, performs similarity-based labeling to assign positive labels to frames belonging to the highly similar cluster, and assigns negative labels to frames belonging to the low-similarity cluster; and a distribution map of the similarity of frames assigned positive labels and frames assigned negative labels corresponding to each of the plurality of candidate texts, and based on the distribution map corresponding to each of the plurality of candidate texts, positive It includes an analysis unit that selects a threshold value corresponding to a target false signal rate for classification between a labeled frame and a voice-labeled frame, and selects a candidate text among the plurality of candidate texts that exhibits the lowest false signal rate relative to the target false signal rate as the alarm text.
[0014] The above text processing unit is characterized by deriving multiple candidate texts associated with the input text through a large language model when an input text is input. Effects of the invention
[0015] According to the present invention, there is no need to secure training data to build a video alarm algorithm, nor is a deep learning network training process required to learn the alarm generation algorithm. Furthermore, the present invention does not require a separate retraining process when a user wishes to add a desired alarm. Moreover, the present invention can simultaneously perform multiple video alarm analyses (e.g., falling, fighting, wearing a helmet, etc.) with only a single operation of video embedding extraction. Furthermore, the present invention can generate various types of alarms, such as alarms for non-structured objects (fire, weather, lighting) and alarms that comprehensively analyze human behavior and clothing. Additionally, the method of selecting the optimal alarm text and determining the threshold value according to the present invention can enhance the performance of text-based video alarms by proposing and selecting the optimal text that generates the desired alarm when generating text-based video alarms. Furthermore, user convenience is improved by proposing a threshold value based on the target false positive rate for the optimal alarm text. Brief explanation of the drawing
[0016] FIG. 1 is a diagram illustrating the configuration of a device for generating a text-based video alarm using a video language model according to an embodiment of the present invention. FIG. 2 is a flowchart illustrating a method for deriving alarm text according to an embodiment of the present invention. FIGS. 3 and FIGS. 4 are drawings for explaining a method for deriving alarm text according to an embodiment of the present invention. FIG. 5 is a flowchart illustrating a method for generating a text-based video alarm using a video language model according to an embodiment of the present invention. FIG. 6 is a diagram illustrating a method for generating a text-based video alarm using a video language model according to an embodiment of the present invention. FIG. 7 is an example diagram of a hardware system for implementing a device for generating a text-based video alarm using a video language model according to an embodiment of the present invention. Specific details for implementing the invention
[0017] In order to clarify the features and advantages of the means for solving the problem of the present invention, the present invention will be described in more detail with reference to specific embodiments of the present invention illustrated in the attached drawings.
[0018] However, detailed descriptions of known functions or configurations that may obscure the essence of the invention are omitted in the following description and the attached drawings. Additionally, it should be noted that identical components throughout the drawings are indicated by the same reference numerals whenever possible.
[0020] Terms and words used in the following description and drawings should not be interpreted as being limited to their ordinary or dictionary meanings, but should be interpreted in a meaning and concept consistent with the technical spirit of the invention, based on the principle that the inventor can appropriately define the concept of terms to best describe his invention. Accordingly, the embodiments described in this specification and the configurations illustrated in the drawings are merely the most preferred embodiments of the invention and do not represent all aspects of the technical spirit of the invention; therefore, it should be understood that various equivalents and modifications capable of replacing them may exist at the time of filing this application.
[0021] Furthermore, terms including ordinal numbers, such as first, second, etc., are used to describe various components and are used solely for the purpose of distinguishing one component from another, and are not used to limit said components. For example, without departing from the scope of the present invention, the second component may be named the first component, and similarly, the first component may be named the second component.
[0022] Furthermore, when it is stated that one component is “connected” or “joined” to another component, this implies that they may be connected or joined logically or physically. In other words, it should be understood that while a component may be directly connected or joined to another component, there may also be intermediate components, or the connection may be indirect.
[0023] Furthermore, the terms used in this specification are used merely to describe specific embodiments and are not intended to limit the invention. Singular expressions include plural expressions unless the context clearly indicates otherwise. Additionally, terms such as “comprising” or “having” described in this specification are intended to indicate the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0024] Additionally, terms such as “…part,” “…unit,” and “module” described in the specification refer to a unit that processes at least one function or operation, and this may be implemented in hardware, software, or a combination of hardware and software.
[0025] Additionally, “one (a or an),” “one,” “the,” and similar terms may be used in the context describing the invention (particularly in the context of the following claims) in a sense including both singular and plural, unless otherwise indicated in this specification or clearly contradicted by the context.
[0027] In addition, embodiments within the scope of the present invention include a computer-readable medium having or transmitting computer-executable instructions or data structures stored on a computer-readable medium. Such a computer-readable medium may be any available medium accessible by a general-purpose or special-purpose computer system. For example, such a computer-readable medium may include, but is not limited to, physical storage media such as RAM, ROM, EPROM, CD-ROM or other optical disk storage devices, magnetic disk storage devices or other magnetic storage devices, or any other medium accessible by a general-purpose or special-purpose computer system that can be used to store or transmit certain program code means in the form of computer-executable instructions, computer-readable instructions or data structures.
[0029] Furthermore, the present invention may be applied in a network computing environment having various types of computer system configurations, including personal computers, laptop computers, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, pagers, etc. The present invention may also be executed in a distributed system environment in which both local and remote computer systems linked via a network via a wired data link, a wireless data link, or a combination of wired and wireless data links perform tasks. In a distributed system environment, program modules may be located in local and remote memory storage devices.
[0031] First, we will describe a device for generating a text-based video alarm using a video language model according to an embodiment of the present invention. FIG. 1 is a diagram illustrating the configuration of a device for generating a text-based video alarm using a video language model according to an embodiment of the present invention.
[0032] Referring to FIG. 1, an image alarm device (10) according to an embodiment of the present invention includes a text processing unit (110), an image processing unit (120), a similarity calculation unit (130), a classification unit (140), an analysis unit (150), a text input unit (210), an image input unit (220), a similarity evaluation unit (230), and an alarm unit (240).
[0033] Here, the text processing unit (110), image processing unit (120), similarity calculation unit (130), classification unit (140), and analysis unit (150) are configured to set the optimal text for setting the alarm desired by the user as the alarm text.
[0034] Additionally, the text input unit (210), image input unit (220), similarity evaluation unit (230), and alarm unit (240) are configured to set a target alarm based on the input of the alarm text and to provide an alarm based on the set alarm.
[0035] According to an embodiment of the present invention, an artificial intelligence model comprising a Vision Language Model (VLM), an Object Detecting Model (DM), and a Large Language Model (LLM) may be used. Here, examples of the Vision Language Model (VLM) include CLIP, BLIP, etc. Examples of the Object Detecting Model (DM) include YOLO, etc. And representative examples of the Large Language Model (LLM) include ChatGPT, etc.
[0036] The specific operation of the aforementioned video alarm device (10), which includes a text processing unit (110), an image processing unit (120), a similarity calculation unit (130), a classification unit (140), an analysis unit (150), a text input unit (210), an image input unit (220), a similarity evaluation unit (230), and an alarm unit (240), will be explained in more detail below.
[0038] Next, a method for deriving an alarm text optimal for a text-based video alarm according to an embodiment of the present invention will be described. FIG. 2 is a flowchart illustrating a method for deriving an alarm text according to an embodiment of the present invention. FIGS. 3 and 4 are drawings illustrating a method for deriving an alarm text according to an embodiment of the present invention.
[0040] Referring to FIGS. 2 to 4, the user may input an input text, which is text related to an alarm, in order to derive an alarm text suitable for the alarm they want. Then, in step S110, the text processing unit (110) derives a plurality (N) of candidate texts (sentences or words) associated with the input text (background text or object text) through a large language model (LLM). For example, as shown in FIG. 3, a plurality of candidate texts (CT) can be derived for the input text “falling down” through a large language model (e.g., ChatGPT, etc.).
[0041] In step S120, the text processing unit (110) derives text embedding vectors (background text embedding vectors or object text embedding vectors) for a plurality (N) of candidate texts (sentences or words) (CT) through a visual language model (VLM).
[0042] Meanwhile, as illustrated in FIG. 4, the user can input a plurality of candidate images (CI) including a plurality of suitable candidate images (A) that contain the desired alarm content and a plurality of unsuitable candidate images (B) that do not contain the desired alarm content.
[0043] Accordingly, in step S130, the image processing unit (120) derives image embedding vectors (background image embedding vectors and object image embedding vectors) through an image language model (VLM) for a plurality of frames of a plurality of candidate images, including a plurality (A) of suitable candidate images and a plurality (B) of unsuitable candidate images.
[0044] The similarity calculation unit (130) calculates the similarity (cosine similarity) between the image embedding vector and the text embedding vector (see step S120) in step S140. This allows N similarities to be calculated from a single image frame. When calculating similarities for a suitable candidate image composed of T frames, T x N similarities can be derived.
[0045] The classification unit (140) performs clustering (e.g., K-means clustering or DBSCAN) on multiple frames of multiple (A) suitable candidate images according to the similarity calculated in step S150 to classify them into two clusters, including a highly similar cluster, which is a group with relatively high similarity, and a lowly similar cluster, which is a group with relatively low similarity; in the case of suitable candidate images, (A x T) frames for N words can be classified into two clusters based on similarity.
[0046] Next, the classification unit (140) classifies multiple frames of multiple (B) unsuitable candidate images into low-similar clusters in step S160.
[0047] Next, the classification unit (140) performs similarity-based labeling in step S170 to assign positive labels to frames belonging to the similarity cluster and negative labels to frames belonging to the low similarity cluster.
[0048] The analysis unit (150) derives a distribution of similarity between frames with positive labels and frames with negative labels at step S180. That is, the analysis unit (150) derives a distribution of similarity between frames with positive labels and frames with negative labels corresponding to each of multiple (N) candidate texts (sentences or words). For example, distributions such as (A), (B), and (C) of FIG. 4 can be derived.
[0049] In step S190, the analysis unit (150) selects a threshold value corresponding to a target false positive rate (e.g., false positive rate 1%, 0.1%, etc.) for classification between frames with positive labels and frames with negative labels based on a distribution map corresponding to each of the multiple (N) candidate texts (sentences or words).
[0050] In step S200, the analysis unit (150) selects a candidate text that shows the lowest non-display rate for the target misdisplay rate among a plurality (N) of candidate texts (sentences or words) as the alarm text, and maps and stores the previously selected threshold value of the candidate text.
[0052] Next, a method for generating a text-based video alarm using a video language model according to an embodiment of the present invention will be described. FIG. 5 is a flowchart illustrating a method for generating a text-based video alarm using a video language model according to an embodiment of the present invention. FIG. 5 is a flowchart illustrating a method for generating a text-based video alarm using a video language model according to an embodiment of the present invention. FIG. 6 is a diagram illustrating a method for generating a text-based video alarm using a video language model according to an embodiment of the present invention.
[0054] Referring to FIGS. 5 and 6, the user can input alarm text, which is text that sets an alarm of a desired form. This alarm text may be the alarm text derived earlier with reference to FIG. 2. Here, the alarm text includes background text and object text. The background text is text that describes the alarm to be detected in the background image. Examples of background text may include “fire,” “explosion,” “weather,” etc. The object text is text that describes the alarm to be detected regarding the detected object (human) image. Examples of object text may include “smoking,” “person collapsed,” “wearing a mask,” etc. Accordingly, the text input unit (210) can receive the alarm text including the background text and object text in step S310. Then, in step S320, the text input unit (210) performs embedding for the background text through a Visual Language Model (VLM) to derive a background text embedding vector, and performs embedding for the object text to derive an object text embedding vector.
[0055] Meanwhile, surveillance video captured through a video device (not shown) may be input. Accordingly, the video input unit (220) receives the surveillance video at step S330. Then, at step S340, the video input unit (220) sets the input surveillance video as a background video and can derive an object video by detecting the area occupied by an object in the surveillance video using an object detection model (DM).
[0056] Then, the image input unit (220) performs embedding for the background image through the image language model (VLM) in step S350 to derive a background image embedding vector, and performs embedding for the object image to derive an object image embedding vector.
[0057] The similarity evaluation unit (230) calculates the similarity between the text embedding vector and the image embedding vector in step S360. That is, the similarity evaluation unit (230) calculates the similarity between the background text embedding vector and the background image embedding vector, and calculates the similarity between the object text embedding vector and the object image embedding vector.
[0058] Here, the similarity can be cosine similarity. Both the text embedding vector containing the background text embedding vector and the object text embedding vector, and the image embedding vector containing the background image embedding vector and the object image embedding vector, are in the form of vectors of size 1 with dimensions of 256 to 512. Accordingly, similarity can be calculated through the vector dot product between the text embedding vector and the image embedding vector.
[0059] The alarm unit (240) determines whether the similarity exceeds a threshold value in step S370. Here, the threshold value may be a threshold value that is pre-stored and mapped to the alarm text with reference to FIG. 2.
[0060] If, as a result of the determination in step S370, the similarity does not exceed the threshold value, the aforementioned steps S330 through S370 are repeated.
[0061] On the other hand, if the similarity exceeds a threshold value as a result of the determination in step S370, the alarm unit (240) generates an alarm in step S380. According to one embodiment, if the similarity between the background text embedding vector and the background image embedding vector exceeds a threshold value, a background alarm may be generated. For example, if the alarm text is the background text "fire," a fire warning may be generated. In addition, according to another embodiment, if the similarity between the object text embedding vector and the object image embedding vector exceeds a threshold value, an object alarm may be generated. For example, if the alarm text is the object text "fallen person," an alarm indicating that a person has fallen may be generated.
[0063] According to the present invention, there is no need to secure training data to build a video alarm algorithm, and a deep learning network training process to learn the alarm generation algorithm is not required. Furthermore, the present invention does not require a separate retraining process when a user wants to add a desired alarm. In addition, the present invention can simultaneously perform multiple video alarm analyses (e.g., falling, fighting, wearing a helmet, etc.) with only a single operation of video embedding extraction. Moreover, the present invention can generate various types of alarms, such as alarms for non-standard objects (fire, weather, lighting) and alarms that comprehensively analyze human behavior and clothing.
[0064] Furthermore, the method for selecting an optimal alarm text and determining a threshold value according to the present invention can enhance the performance of a text-based video alarm by proposing and selecting the optimal text that generates the desired alarm when generating a text-based video alarm. Additionally, user convenience is improved by proposing a threshold value based on the target false positive rate for the optimal alarm text.
[0066] FIG. 7 is an example diagram of a hardware system for implementing a device for generating a text-based video alarm using a video language model according to an embodiment of the present invention.
[0067] As illustrated in FIG. 7, a hardware system (2000) according to one embodiment of the present invention may have a configuration including a processor unit (2100), a memory interface unit (2200), and a peripheral device interface unit (2300).
[0068] Each component within such a hardware system (2000) may be an individual component or integrated into one or more integrated circuits, and each of these components may be combined by a bus system (not shown).
[0069] Here, in the case of a bus system, it is an abstraction representing any one or more individual physical buses, communication lines / interfaces, and / or multi-drop or point-to-point connections connected by appropriate bridges, adapters, and / or controllers.
[0070] The processor unit (2100) performs the role of executing various software modules stored in the memory unit (2210) by communicating with the memory unit (2210) through the memory interface unit (2200) to perform various functions in the hardware system.
[0071] Here, in the memory unit (2210), each of the components including the text processing unit (110), image processing unit (120), similarity calculation unit (130), classification unit (140), analysis unit (150), text input unit (210), image input unit (220), similarity evaluation unit (230), and alarm unit (240) described above with reference to FIG. 1 can be stored in the form of a software module, and additionally, an operating system (OS) can be stored. The components including the text processing unit (110), image processing unit (120), similarity calculation unit (130), classification unit (140), analysis unit (150), text input unit (210), image input unit (220), similarity evaluation unit (230), and alarm unit (240) can be loaded into the processor unit (2100) and executed.
[0072] Each component including the text processing unit (110), image processing unit (120), similarity calculation unit (130), classification unit (140), analysis unit (150), text input unit (210), image input unit (220), similarity evaluation unit (230), and alarm unit (240) described above may be implemented in the form of a software module or a hardware module executed by a processor, or in a form in which a software module and a hardware module are combined.
[0073] In this way, software modules, hardware modules, or combinations of software and hardware modules executed by a processor can be implemented as actual hardware systems (e.g., computer systems).
[0074] In the case of an operating system (e.g., embedded operating systems such as I-OS, Android, Darwin, RTXC, LINUX, UNIX, OS X, WINDOWS, or VxWorks), it includes various procedures, instruction sets, software components, and / or drivers that control and manage general system tasks (e.g., memory management, storage device control, power management, etc.) and plays a role in facilitating communication between various hardware modules and software modules.
[0075] For reference, the memory section (2210) may include a memory hierarchy that includes, but is not limited to, a cache, main memory, and secondary memory, and such a memory hierarchy may be implemented through any combination of, for example, RAM (e.g., SRAM, DRAM, DDRAM), ROM, FLASH, magnetic and / or optical storage devices [e.g., disk drive, magnetic tape, CD (compact disk) and DVD (digital video disc), etc.].
[0076] The peripheral device interface unit (2300) performs the role of enabling communication between the processor unit (2100) and the peripheral device.
[0077] Here, the peripheral device is intended to provide different functions to the hardware system (2000), and in one embodiment of the present invention, for example, a communication unit (2310) may be included.
[0078] Here, the communication unit (2310) performs the role of providing communication functions with other devices, and for this purpose, it may include, for example, an antenna system, an RF transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a CODEC chipset, and a memory, and may include known circuits that perform this function, which are not limited thereto.
[0079] Communication protocols supported by the communication unit (2310) include, for example, Wireless LAN (WLAN), DLNA (Digital Living Network Alliance), WiBro (Wireless Broadband), WiMAX (World Interoperability for Microwave Access), GSM (Global System for Mobile communication), CDMA (Code Division Multi Access), CDMA2000 (Code Division Multi Access 2000), EV-DO (Enhanced Voice-Data Optimized or Enhanced Voice-Data Only), WCDMA (Wideband CDMA), HSDPA (High Speed Downlink Packet Access), HSUPA (High Speed Uplink Packet Access), IEEE 802.16, Long Term Evolution (LTE), LTE-A (Long Term Evolution-Advanced), 5G communication system, Wireless Mobile Broadband Service (WMBS), Bluetooth, RFID (Radio Frequency Identification), infrared Infrared Data Association (IrDA), Ultra-Wideband (UWB), ZigBee, Near Field Communication (NFC), Ultra Sound Communication (USC), Visible Light Communication (VLC), Wi-Fi, Wi-Fi Direct, etc. may be included.In addition, wired communication networks may include wired LAN (Local Area Network), wired WAN (Wide Area Network), Power Line Communication (PLC), USB communication, Ethernet, serial communication, optical / coaxial cables, etc., and any protocol capable of providing a communication environment with other devices, which is not limited thereto, may be included.
[0080] In a hardware system (2000) according to one embodiment of the present invention, each component stored in the memory unit (2210) in the form of a software module performs an interface with the communication unit (2310) through the memory interface unit (2200) and the peripheral device interface unit (2300) in the form of instructions executed by the processor unit (2100).
[0082] As explained above, although this specification contains details of a number of specific embodiments, they should not be understood as limiting the scope of any invention or claimables, but rather as descriptions of features that may be characteristic of a specific embodiment of a specific invention. In the context of individual embodiments, specific features described in this specification may be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any suitable sub-combination. Furthermore, while features may operate in a specific combination and be described as initially claimed, one or more features from the claimed combination may be excluded from the combination in some cases, and the claimed combination may be changed to a sub-combination or a variation of the sub-combination.
[0083] Likewise, although operations are depicted in the drawings in a specific order, this should not be understood as requiring that such operations be performed in that specific or sequential order depicted to obtain a desirable result, or that all depicted operations must be performed. In certain cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of the various system components of the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together into a single software product or packaged into multiple software products.
[0085] Specific embodiments of the subject matter described herein have been explained. Other embodiments fall within the scope of the following claims. For example, the operations cited in the claims may be performed in different orders and still achieve desirable results. As an example, the process illustrated in the accompanying drawings does not necessarily require the specific illustrated order or sequential order to obtain desirable results. In specific embodiments, multitasking and parallel processing may be advantageous.
[0087] The description provided herein presents the best mode of the invention and offers examples to explain the invention and to enable those skilled in the art to manufacture and use the invention. The specification thus written is not intended to limit the invention to the specific terms presented. Accordingly, although the invention has been described in detail with reference to the examples above, those skilled in the art may make modifications, changes, and variations to these examples without departing from the scope of the invention.
[0089] Therefore, the scope of the present invention should not be determined by the described embodiments but by the claims. Explanation of the symbols
[0090] 110: Text processing unit 120: Image processing unit 130: Similarity Calculation Unit 140: Classification section 150: Analysis Department 210: Text input section 220: Video input section 230: Similarity Evaluation Department 240: Alarm section
Claims
Claim 1 When an alarm text including background text and object text is input by a text input unit, a step of performing embedding on the background text through a visual language model to derive a background text embedding vector and performing embedding on the object text to derive an object text embedding vector; when a surveillance image is input by an image input unit, a step of setting the input surveillance image as a background image and using an object detection model to derive an object image in which the area occupied by an object in the input surveillance image is detected; a step in which the image input unit performs embedding on the background image through a visual language model to derive a background image embedding vector and performs embedding on the object image to derive an object image embedding vector; a step in which a similarity evaluation unit calculates the similarity between the background text embedding vector and the background image embedding vector and calculates the similarity between the object text embedding vector and the object image embedding vector; The alarm unit includes the step of generating an alarm corresponding to the alarm text when the calculated similarity between the background text embedding vector and the background image embedding vector, and the similarity between the object text embedding vector and the object image embedding vector, exceed a preset threshold value; and prior to the step of deriving the object text embedding vector, the text processing unit derives a text embedding vector for a plurality of candidate texts through an image language model; the image processing unit derives an image embedding vector for a plurality of frames of a plurality of candidate images, including a plurality of suitable candidate images and a plurality of unsuitable candidate images, through an image language model; the similarity calculation unit calculates the similarity between the image embedding vector and the text embedding vector; the classification unit performs clustering on a plurality of frames of the plurality of suitable candidate images according to the similarity and classifies them into a highly similar cluster, which is a group with relatively high similarity, and a low similarity cluster, which is a group with relatively low similarity; the classification unit [classifies] the plurality of frames of the plurality of unsuitable candidate images Step of classifying into low-similarity clusters;A method for generating a video alarm, characterized by comprising: a step in which a classification unit performs similarity-based labeling to assign a positive label to a frame belonging to a similarity cluster and assign a negative label to a frame belonging to a low similarity cluster; a step in which an analysis unit derives a distribution of similarity between a frame assigned a positive label and a frame assigned a negative label corresponding to each of the plurality of candidate texts; a step in which the analysis unit selects a threshold value corresponding to a target false voicing rate for classification between a frame assigned a positive label and a frame assigned a negative label based on the distribution corresponding to each of the plurality of candidate texts; and a step in which the analysis unit selects a candidate text among the plurality of candidate texts that exhibits the lowest non-voicing rate relative to the target false voicing rate as an alarm text. Claim 2 A method for generating an image alarm according to claim 1, wherein the step of calculating the similarity is characterized in that the similarity evaluation unit calculates the similarity through a vector dot product between the background text embedding vector and the background image embedding vector, and calculates the similarity through a vector dot product between the object text embedding vector and the object image embedding vector. Claim 3 delete Claim 4 A method for generating an image alarm according to claim 1, further comprising, prior to the step of deriving text embedding vectors for the plurality of candidate texts through an image language model, a step in which, when an input text is input, the text processing unit derives a plurality of candidate texts associated with the input text through a large language model. Claim 5 A text input unit that, when an alarm text including background text and object text is input, performs embedding on the background text through a visual language model to derive a background text embedding vector and performs embedding on the object text to derive an object text embedding vector; an image input unit that, when a surveillance image is input, sets the input surveillance image as a background image, derives an object image by detecting the area occupied by an object in the input surveillance image using an object detection model, performs embedding on the background image through a visual language model to derive a background image embedding vector, and performs embedding on the object image to derive an object image embedding vector; and a similarity evaluation unit that calculates the similarity between the background text embedding vector and the background image embedding vector and calculates the similarity between the object text embedding vector and the object image embedding vector. The alarm unit, which generates an alarm corresponding to the alarm text when the calculated similarity exceeds a preset threshold value; the text processing unit, which derives text embedding vectors for a plurality of candidate texts through an image language model; the image processing unit, which derives image embedding vectors for a plurality of frames of a plurality of candidate images, including a plurality of suitable candidate images and a plurality of unsuitable candidate images, through an image language model; the similarity calculation unit, which calculates the similarity between the image embedding vector and the text embedding vector; and the classification unit, which performs clustering on a plurality of frames of the plurality of suitable candidate images according to the similarity to classify them into a highly similar cluster, which is a group with relatively high similarity, and a low similarity cluster, which is a group with relatively low similarity, and classifies a plurality of frames of the plurality of unsuitable candidate images into the low similarity cluster, and performs similarity-based labeling to assign positive labels to frames belonging to the highly similar cluster and negative labels to frames belonging to the low similarity cluster.The device for generating a video alarm further comprises: an analysis unit that derives a distribution of similarity between a frame with a positive label and a frame with a negative label corresponding to each of the plurality of candidate texts, selects a threshold value corresponding to a target false voicing rate for classification between a frame with a positive label and a frame with a negative label based on the distribution corresponding to each of the plurality of candidate texts, and selects a candidate text among the plurality of candidate texts that exhibits the lowest non-voicing rate relative to the target false voicing rate as an alarm text. Claim 6 A device for generating an image alarm, wherein, in claim 5, the similarity evaluation unit calculates similarity through a vector dot product between the background text embedding vector and the background image embedding vector, and calculates similarity through a vector dot product between the object text embedding vector and the object image embedding vector. Claim 7 delete Claim 8 A device for generating an image alarm according to claim 5, wherein the text processing unit derives a plurality of candidate texts associated with the input text through a large language model when an input text is input.
Citation Information
Patent Citations
Method of live video event detection based on natural language queries, and an apparatus for the same
US20220138489A1