Video Surveillance Device Based On Video Sensing Method

KR1020260139043APending Publication Date: 2026-09-21XENOSYS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
KR1020260165228
Authority / Receiving Office
KR · KR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-09-01
Publication Date
2026-09-21

Smart Images

  • Figure PAT00001_ABST
    Figure PAT00001_ABST
Patent Text Reader

Abstract

The present invention relates to a VLM-based hybrid image detection method, and more specifically, to a VLM-based hybrid image detection method that maximizes the accuracy of event detection while efficiently utilizing the limited computational resources of an image detection device by inputting a policy-applied image, to which an image policy has been applied, into a VLM (Vision-Language Model) to analyze whether an event has occurred when it is determined that VLM analysis is necessary in accordance with VLM execution conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to a VLM-based hybrid image detection method, and more specifically, to a VLM-based hybrid image detection method that maximizes the accuracy of event detection while efficiently utilizing the limited computational resources of an image detection device by inputting a policy-applied image, to which an image policy has been applied, into a VLM (Vision-Language Model) to analyze whether an event has occurred when it is determined that VLM analysis is necessary in accordance with VLM execution conditions. Background Technology

[0003] In modern society, the importance of video surveillance systems is increasing day by day for purposes such as public safety, crime prevention, and facility management. Existing video surveillance systems have primarily adopted methods of simply recording video or transmitting all footage to a central server for analysis. However, as camera resolutions have increased and the number of installed channels has surged, the approach of transmitting all video data to a central server for analysis is causing serious problems such as the depletion of network bandwidth and the concentration of computational load on servers. Furthermore, there are limitations in that real-time response to emergencies is difficult due to delays occurring during data transmission.

[0004] To address these issues, technologies that enable edge terminals to analyze video independently are being introduced; however, existing technologies primarily rely on simple object detection or predefined rules. While they can detect the presence of objects or simple boundary violations, they have a problem in that they cannot accurately recognize the behavioral context of objects or complex situations.

[0005] To overcome these limitations, methods to improve analysis accuracy by applying Vision-Language Model (VLM) technology, which combines images and text to understand context, have recently been proposed. However, VLM requires a very high amount of computation, and there are practical limitations in running VLM on every frame of real-time video in image detection devices implemented on resource-limited edge terminals rather than high-performance servers, as this is practically impossible.

[0006] Therefore, there is a growing need for the development of technology that can improve analysis accuracy while efficiently managing the resources of image detection devices by selecting specific situations where object movement or change is detected using relatively simple and lightweight algorithms, and selectively driving the VLM only when precise analysis is required. Prior art literature

[0008] Korean Registered Patent KR 10-2902413 B1 “Video surveillance security control system and method equipped with an object recognition-based large-scale language model” (Dec. 16, 2025) Korean Registered Patent KR 10-1795798 B1 “Intelligent video surveillance system” (Nov. 02, 2017) Korean Registered Patent KR 10-1526499 B1 “Security network crime prevention surveillance system using object detection function and intelligent video analysis method using the same” (June 1, 2015) Korean Registered Patent KR 10-2752688 B1 “Local analysis type video surveillance system fusing thermal imaging, visible light cameras and LiDAR sensors and operation method” (January 06, 2025) US Published Patent US2021 / 0409792 A1 “DISTRIBUTED SURVEILLANCE SYSTEM WITH DISTRIBUTED VIDEO ANALYSIS” (2021.12.30) US registered patent US7136507 B2 “VIDEO SURVEILLANCE SYSTEM WITH RULE-BASED REASONING AND MULTIPLE-HYPOTHESIS SCORING” (2006.11.14) US published patent US2023 / 0360355 A1 “HYBRID VIDEO ANALYTICS FOR SMALL AND SPECIALIZED OBJECT DETECTION”(2023.11.09)US Registered Patent US11729347 B2 “VIDEO SURVEILLANCE SYSTEM, VIDEO PROCESSING APPARATUS, VIDEO PROCESSING METHOD, AND VIDEO PROCESSING PROGRAM”(2023.08.15) The problem to be solved

[0009] The present invention aims to provide a VLM-based hybrid image detection method, and more specifically, a VLM-based hybrid image detection method that maximizes the accuracy of event detection while efficiently utilizing the limited computational resources of an image detection device by inputting a policy-applied image, to which an image policy has been applied, into a VLM (Vision-Language Model) to analyze whether an event has occurred when it is determined that VLM analysis is necessary in accordance with VLM execution conditions. means of solving the problem

[0011] To solve the above-mentioned problem, a video detection method performed in a VLM (Vision-Language Model)-based hybrid video surveillance device that aims for resource optimization comprises: a tracking information generation step of detecting an object in a frame to be analyzed of video information captured by a camera, and generating and managing tracking information including one or more of the ID, class, location, velocity, and movement path of the detected object; a VLM execution condition application step of determining whether the tracking information of the frame to be analyzed satisfies one or more preset VLM execution conditions, and if it satisfies, applying an image policy corresponding to the VLM execution condition to the video image of the frame to be analyzed to generate a policy-applied image; and a VLM analysis step of inputting a prompt into the VLM that includes the policy-applied image generated for the frame to be analyzed, video images for a plurality of preset frames before and after the frame to be analyzed, and a plurality of preset queries, to derive an analysis score including semantic association between the video image of the frame to be analyzed or the policy-applied image and each of the plurality of queries. The present invention provides a video detection method comprising: a final analysis step of analyzing whether an event to be detected in each of the multiple queries occurs in the video image or policy application image of the analysis target frame based on an analysis score derived for each of the multiple queries.

[0012] In one embodiment of the present invention, the VLM execution condition includes a periodic condition that determines that VLM analysis for the frame to be analyzed should be performed according to a preset time period; and the image policy may include a policy that generates the original image of the frame to be analyzed as the policy-applied image when the frame to be analyzed meets the periodic condition.

[0013] In one embodiment of the present invention, the tracking information includes an object state indicator comprising one or more of an object detection status, bounding box size, position coordinates, and movement speed, and the VLM execution condition may include a user-configured condition that determines that VLM analysis should be performed on the frame to be analyzed when the amount of change calculated by comparing the object state indicator of the frame to be analyzed and the previous frame exceeds a preset threshold.

[0014] In one embodiment of the present invention, the VLM execution condition includes a user-configured condition, and the user-configured condition may include one or more detailed conditions among: a new object detection condition that determines whether a new object that was not detected in the previous frame is detected in the frame to be analyzed; a region entry condition that determines whether the position coordinates of the object enter into a preset region of interest; and a direction violation condition that determines whether the direction of movement according to the change in the position coordinates of the object differs from the allowed direction of movement set in the virtual direction of movement setting area where the object exists by more than a preset angle.

[0015] In one embodiment of the present invention, the image policy includes one or more detailed policies to be applied to the image of the frame to be analyzed when the tracking information of the frame to be analyzed matches the detailed condition for each of the one or more detailed conditions included in the user-configured conditions, and the VLM execution condition application step can determine whether the tracking information of the frame to be analyzed satisfies each of the one or more detailed conditions, and if satisfied, apply the detailed policy for the detailed condition to the image of the frame to be analyzed to generate a policy-applied image.

[0016] In one embodiment of the present invention, the detailed policy included in the image policy may include one or more of: a policy for generating a policy-applied image by cropping an area including a bounding box of a newly detected object when a new object that was not detected in a previous frame is detected in an analysis target frame; a policy for generating a policy-applied image by cropping an area including a pre-set area of ​​interest when the position coordinates of an object enter into a pre-set area of ​​interest; and a policy for generating a policy-applied image by cropping an area including a bounding box of an object and a movement direction setting area when the movement direction according to the change in the position coordinates of the object differs from the allowed movement direction set in a virtual movement direction setting area where the object exists by an angle greater than or equal to a pre-set angle.

[0017] In one embodiment of the present invention, the final analysis step may include: a step of setting a sliding frame section configured to include the analysis target frame and a predetermined number of adjacent frames that are temporally continuous with the analysis target frame; a step of deriving a section analysis score for each of a plurality of queries of the analysis target frame based on an analysis score derived for each of a plurality of frames included in the sliding frame section; and a step of determining that an event to be detected through a query corresponding to the section analysis score has occurred in the analysis target frame when the section analysis score exceeds a predetermined standard for the corresponding section analysis score.

[0018] In one embodiment of the present invention, the query may include: a positive query describing a situation in which an event to be detected has occurred; and a negative query describing an exceptional situation in which the visual characteristics are similar to the event but the event has not actually occurred.

[0019] In one embodiment of the present invention, the analysis score may be calculated as higher by adding as the semantic association between the image or policy application image of the frame to be analyzed and the positive query is higher, and as lower by subtracting as the semantic association between the image or policy application image of the frame to be analyzed and the negative query is higher.

[0020] In one embodiment of the present invention, the VLM analysis step may input into the VLM by adding a phrase to the prompt requesting that, among the one or more VLM execution conditions, the event or query related to the VLM execution condition to which the tracking information of the analysis target frame matches be analyzed in particular, thereby causing the VLM to evaluate the semantic relevance of the query related to the VLM execution condition as higher than other queries, and to calculate the analysis score for the query related to the VLM execution condition with weights applied.

[0021] To solve the above-mentioned problem, a Vision-Language Model (VLM)-based video surveillance device for resource optimization comprises: a tracking information generation unit that detects an object in an analysis target frame of video information captured by a camera, and generates and manages tracking information including one or more of the ID, class, location, velocity, and movement path of the detected object; a VLM execution condition application unit that determines whether the tracking information of the analysis target frame satisfies one or more preset VLM execution conditions, and if it satisfies them, applies an image policy corresponding to the VLM execution condition to the video image of the analysis target frame to generate a policy-applied image; and a VLM analysis unit that inputs a prompt into the VLM including the policy-applied image generated for the analysis target frame, video images for a plurality of preset frames before and after the analysis target frame, and a plurality of preset queries, and derives an analysis score including semantic association between the video image of the analysis target frame or the policy-applied image and each of the plurality of queries. The present invention provides an image detection device comprising: an analysis unit that analyzes whether an event to be detected in each of a plurality of queries occurs in an image of an analysis target frame or a policy application image based on an analysis score derived for each of a plurality of queries.

[0022] To solve the above problem, a computer-readable recording medium that performs an image detection method, comprising one or more processors and one or more memories, wherein the computer-readable recording medium stores instructions for performing the following steps, and the following steps include: a tracking information generation step of detecting an object in an analysis target frame of image information captured by a camera, and generating and managing tracking information including one or more of the ID, class, location, velocity, and movement path of the detected object; and a VLM execution condition application step of determining whether the tracking information of the analysis target frame satisfies one or more preset VLM execution conditions, and if it satisfies, applying an image policy corresponding to the VLM execution condition to the image of the analysis target frame to generate a policy-applied image. A computer-readable recording medium is provided, comprising: a VLM analysis step of inputting a prompt into a VLM including a policy application image generated for the frame to be analyzed, video images for a plurality of preset frames before and after the frame to be analyzed, and a plurality of preset queries, to derive an analysis score including semantic associations between the video image or policy application image of the frame to be analyzed and each of the plurality of queries; and a final analysis step of analyzing whether an event to be detected in each of the plurality of queries occurs in the video image or policy application image of the frame to be analyzed based on the analysis score derived for each of the plurality of queries. Effects of the invention

[0024] According to one embodiment of the present invention, by using tracking information for an image of a frame to be analyzed to determine whether VLM analysis is necessary for the image, and by performing a high-computation VLM analysis step only when the VLM execution conditions are met, it is possible to achieve the effect of optimizing resource usage of the image detection device while ensuring real-time event detection.

[0025] According to one embodiment of the present invention, by setting periodic conditions as VLM execution conditions and performing VLM analysis on the frame to be analyzed according to a preset time period even in situations where there are no special events, it is possible to periodically monitor image information.

[0026] According to one embodiment of the present invention, by applying a user-defined condition in which an object state indicator (amount of change) included in the tracking information exceeds a threshold, and performing VLM analysis on the frame to be analyzed for specific situations satisfying specific user-defined conditions, such as the discovery of an object with rapid movement or a large change in state, a precise analysis can be selectively performed.

[0027] According to one embodiment of the present invention, by setting specific detailed conditions such as new object detection conditions, area entry conditions, and direction violation conditions as user-configured conditions, the user can save computational resources and reduce false alarms by selectively performing the VLM analysis step only for situations where a specific event to be detected by the user is predicted to occur.

[0028] According to one embodiment of the present invention, by applying an image policy that crops an image centered on an object or region matching the VLM execution condition and generating a policy-applied image, intensive analysis of the region related to the event to be detected within the entire image can be performed.

[0029] According to one embodiment of the present invention, by finally determining whether an event occurs in the frame to be analyzed based on the segment analysis score derived from the sliding frame segment that is variably set over time, temporary noise or misrecognition that may occur during single-frame analysis can be excluded and the reliability of the event occurrence determination can be ensured.

[0030] According to one embodiment of the present invention, by configuring the query into positive and negative queries, the VLM can detect events that are visually similar but have different meanings while clearly distinguishing them.

[0031] According to one embodiment of the present invention, by dynamically configuring a prompt to focus analysis on events related to VLM execution conditions that match the tracking information of the frame to be analyzed, it is possible to perform a more precise analysis on a specific event for which a focus analysis is requested among a plurality of events that the VLM can detect. Brief explanation of the drawing

[0033] FIG. 1 illustrates the components of a system implementing an image detection method according to one embodiment of the present invention. FIG. 2 illustrates embodiments of a system implementing an image detection method according to one embodiment of the present invention. FIG. 3 illustrates tracking information for an analysis target frame according to one embodiment of the present invention. FIG. 4 illustrates a process in which tracking information of an analysis target frame is determined to meet VLM execution conditions according to an embodiment of the present invention. FIG. 5 illustrates a process in which tracking information of an analysis target frame according to one embodiment of the present invention is determined to correspond to a plurality of VLM execution conditions. FIG. 6 illustrates an image policy for VLM execution conditions according to an embodiment of the present invention, and a policy-applied image generated according to the image policy. FIG. 7 illustrates examples of VLM execution conditions and image policies according to an embodiment of the present invention. FIG. 8 illustrates a VLM analysis step according to one embodiment of the present invention. FIG. 9 illustrates a VLM analysis step according to another embodiment of the present invention. FIG. 10 illustrates the section analysis score for a sliding frame section according to one embodiment of the present invention. FIG. 11 illustrates the form of a query according to one embodiment of the present invention. FIG. 12 schematically illustrates the internal configuration of a computing device according to one embodiment of the present invention. Specific details for implementing the invention

[0034] Hereinafter, various embodiments and / or aspects are disclosed with reference to the drawings. For illustrative purposes, numerous specific details are disclosed in the following description to aid in a general understanding of one or more aspects. However, it will also be recognized by those skilled in the art that these aspects may be practiced without such specific details. The following description and the accompanying drawings describe specific exemplary aspects of one or more aspects in detail. However, these aspects are exemplary, and some of the various methods in the principles of the various aspects may be used, and the description is intended to include all such aspects and their equivalents.

[0036] In addition, various aspects and features will be presented by a system that may include multiple devices, components and / or modules, etc. It should also be understood and recognized that various systems may include additional devices, components and / or modules, etc., and / or may not include all of the devices, components, modules, etc. discussed in relation to the drawings.

[0037] Terms such as “embodiment,” “example,” “aspect,” “example,” etc. as used herein may not be interpreted as implying that any aspect or design described is superior or more advantageous than other aspects or designs. Terms used below, such as “part,” “component,” “module,” “system,” “interface,” etc., generally refer to computer-related entities and may, for example, refer to hardware, a combination of hardware and software, or software.

[0038] Additionally, the terms “comprising” and / or “comprising” should be understood to mean that the relevant feature and / or component is present, but not to exclude the presence or addition of one or more other features, components and / or groups thereof.

[0039] Additionally, terms including ordinal numbers, such as first, second, etc., may be used to describe various components, but said components are not limited by said terms. Such terms are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of the present invention, the first component may be named the second component, and similarly, the second component may be named the first component. The term "and / or" includes a combination of a plurality of related described items or any of a plurality of related described items.

[0040] Furthermore, in the embodiments of the present invention, all terms used herein, including technical or scientific terms, unless otherwise defined, have the same meaning as generally understood by those skilled in the art to which the present invention pertains. Terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in the embodiments of the present invention.

[0042] FIG. 1 illustrates the components of a system implementing an image detection method according to one embodiment of the present invention.

[0044] As illustrated in FIG. 1, a Vision-Language Model (VLM)-based image detection method for resource optimization comprises: a tracking information generation step for detecting an object in an analysis target frame of image information captured by a camera, and generating and managing tracking information including one or more of the ID, class, location, velocity, and movement path of the detected object; a VLM execution condition application step for determining whether the tracking information of the analysis target frame satisfies one or more preset VLM execution conditions, and if it satisfies, applying an image policy corresponding to the VLM execution condition to the image of the analysis target frame to generate a policy-applied image; and a VLM analysis step for inputting a prompt into the VLM that includes the policy-applied image generated for the analysis target frame, image images for a plurality of preset frames before and after the analysis target frame, and a plurality of preset queries, to derive an analysis score including semantic association between the image of the analysis target frame or the policy-applied image and each of the plurality of queries. It may include a final analysis step of analyzing whether an event to be detected in each of the multiple queries occurs in the image of the analysis target frame or the policy application image based on the analysis score derived for each of the multiple queries.

[0046] Specifically, the present invention is a technology for performing artificial intelligence-based image analysis independently within an image detection device (1) rather than a high-performance server, and in particular, without performing a VLM analysis step that requires a high computational load for all frames of image information captured by a camera (3), it is possible to examine whether VLM analysis is necessary for the image of the frame to be analyzed based on tracking information generated for the frame to be analyzed first.

[0047] In addition, according to the results of a review on whether VLM analysis is necessary, the VLM analysis step can be selectively performed only on specific target frames that meet the VLM execution conditions.

[0048] Through this, high-level semantic image analysis can be performed without data processing bottlenecks even in image detection devices (1) with limited hardware resources.

[0050] A system implementing the present invention may be composed of a camera (3) for capturing images, an image detection device (1) for performing image analysis, and a server system (2) for managing the system and setting queries, etc.

[0051] Specifically, the camera (3) can photograph the area under surveillance and transmit the generated image information to the image detection device (1) in real time via a wired or wireless network.

[0052] The server system (2) can generate or manage a query defined in natural language text for a specific event (e.g., fire, collapse, fight, etc.) that the user wants to detect, and transmit this to the video detection device (1) to set the purpose and criteria for the monitoring that the video detection device (1) must perform. Alternatively, the server system (2) can receive the analysis result derived from the video detection device (1), that is, the analysis result regarding whether the event the user wants to detect through the query has occurred, from the video detection device (1) and provide it to the user.

[0054] The image detection device (1) may be a computing device comprising one or more processors and one or more memories as a component implementing the present invention. Specifically, the image detection device (1) may perform an independent image detection process based on an image received from a camera (3) and a query received from a server system (2), and may include a tracking information generation unit (10), a VLM execution condition application unit (11), a VLM analysis unit (12), and a final analysis unit (13).

[0055] The tracking information generation unit (10) can detect and track an object in an input image and generate tracking information that indicates the movement path and state change of the object.

[0056] The VLM execution condition application unit (11) can determine in real time whether the generated tracking information matches the pre-set VLM execution conditions and decide whether to perform VLM analysis. In addition, if the VLM execution condition application unit (11) determines that VLM analysis is necessary, it can generate a policy application image using an image policy for the VLM execution conditions that the tracking information matches.

[0057] The VLM analysis unit (12) can derive an analysis score by inputting a policy-applied image and a query generated according to the image policy into the VLM, and the query may be information describing an event that the user wants to detect in the image information captured by the camera (3).

[0058] The final analysis unit (13) can determine whether an actual event occurred in the frame to be analyzed by comprehensively considering the analysis scores derived for each of the pre-set number of adjacent frames that are temporally consecutive with the frame to be analyzed.

[0060] In one embodiment of the present invention, the image detection device (1) may be connected to an external or internal VLM. Specifically, the VLM (Vision-Language Model) may correspond to a learned multimodal artificial intelligence model capable of simultaneously understanding and processing image and text (Language) information as a vision-language model.

[0062] According to one embodiment of the present invention, by using tracking information for an image of a frame to be analyzed to determine whether VLM analysis is necessary for the image, and by performing a high-computation VLM analysis step only when the VLM execution conditions are met, the effect of optimizing resource usage of the image detection device (1) while ensuring real-time event detection can be achieved.

[0064] FIG. 2 illustrates embodiments of a system implementing an image detection method according to one embodiment of the present invention.

[0066] As illustrated in FIG. 2, the image detection device (1) of the present invention can be physically implemented as an edge terminal of various forms, and can be flexibly applied depending on the type of camera (3) and the configuration of the existing image surveillance system. Specifically, the image detection device (1) may be configured as separate hardware independent of the image shooting function, or it may be implemented by being integrated into the camera (3) in the form of a software module or an embedded chipset.

[0067] FIG. 2(a) illustrates an embodiment in which a camera (3) and an image detection device (1) are implemented as a single unit. In this case, the camera (3) is an intelligent AI camera capable of performing not only a simple image shooting function but also an image analysis function according to the present invention, and may include an image detection device (1) itself. That is, image analysis is performed simultaneously with image shooting inside the camera (3), and the analysis results can be transmitted to an external server system (2), etc.

[0068] FIG. 2(b) illustrates an embodiment in which the camera (3) is an IP camera and the image detection device (1) is implemented in the form of a decoder. Specifically, the IP camera transmits captured images through a network, and in order to monitor or store them, a decoder that decodes the image signal may be required. In this embodiment, the decoder can be implemented to perform the role of the image detection device (1) of the present invention. That is, the decoder including the image detection device (1) can perform image analysis as described in the present invention in addition to the image decoding function.

[0069] FIG. 2(c) illustrates an embodiment in which the camera (3) is an HD-SDI (High Definition Serial Digital Interface) camera and the image detection device (1) is implemented in the form of a video server. Specifically, the HD-SDI camera transmits high-quality uncompressed video signals through a coaxial cable and may require a video server to utilize this in a digital network environment. In one embodiment of the present invention, the video server includes an image detection device (1) and can perform the image analysis described in the present invention in addition to the function of converting analog signals into digital signals.

[0071] Meanwhile, the server system (2) is connected to the video detection device (1) via a wired or wireless network and can perform the role of controlling the entire system and managing settings. Specifically, the server system (2) can receive a query from the user defining a specific event that the user wants to detect and transmit it to the video detection device (1), and the video detection device (1) can store the query internally and utilize it in the VLM analysis stage.

[0072] Meanwhile, in one embodiment of the present invention, the server system (2) may be implemented by being integrated into the image detection device (1). For example, the function of providing a user interface for receiving a query from a user among the functions of the server system (2) may be implemented in a form where the image detection device (1) provides such a function. In that case, the user can input a query directly connected to the image detection device (1).

[0074] As such, the present invention can be implemented by configuring the image detection device (1) as an integral unit with the camera (3), or by software-integrating the function of the image detection device (1) into an edge terminal such as a decoder or video server in an environment with an existing IP camera or analog camera.

[0076] FIG. 3 illustrates tracking information for an analysis target frame according to one embodiment of the present invention.

[0078] As illustrated in FIG. 3, the tracking information may include an object state indicator comprising one or more of the following: whether an object is detected, bounding box size, position coordinates, and movement speed.

[0080] For each frame of image information received from the camera (3) of the image detection device (1), object detection and tracking can be performed in real time to generate tracking information. For example, when f3 among the frames (f1 to f9) of the image information is assumed to be the frame to be analyzed, the image detection device (1) can identify each object and generate tracking information including detailed information about the object when there are multiple objects (two people, one firework) in the image of the frame to be analyzed f3.

[0081] Specifically, for each object detected in the image, the tracking information may include one or more object state indicators among an identifier (ID) for uniquely identifying each object, a class (Class) indicating the type of object, position coordinates (x, y) indicating the current location of the object, a velocity ($v$) indicating the speed of the object's movement, and a movement path indicating the direction or trajectory of the object's movement.

[0083] Meanwhile, the image detection device (1) according to the present invention can determine whether VLM analysis is required for the corresponding frame by first determining whether the tracking information generated for each frame in this manner meets the predefined VLM execution conditions. The image detection device (1) can prevent unnecessary resource waste by selecting whether the frame to be analyzed requires precise analysis based on the object state indicator of the tracking information, and by performing subsequent VLM analysis only in such cases.

[0085] FIG. 4 illustrates a process in which tracking information of an analysis target frame is determined to meet VLM execution conditions according to an embodiment of the present invention.

[0087] As illustrated in FIG. 4, the image detection device (1) can determine whether the tracking information of the frame to be analyzed meets one or more preset VLM execution conditions. A VLM execution condition may correspond to a criterion for the image detection device (1) to select whether VLM analysis is necessary for the image of the frame to be analyzed in order to efficiently use limited resources. Specifically, the VLM execution condition may include periodic conditions for performing VLM analysis at preset time intervals and user-set conditions for detecting events that the user intends to detect, and this will be described later.

[0089] FIG. 4(a) corresponds to a case where tracking information generated for a frame to be analyzed is compared with a plurality of preset conditions (VLM execution conditions 1, 2, and 3), and the tracking information does not match 'VLM execution condition 1' and 'VLM execution condition 3' but matches 'VLM execution condition 2'. The image detection device (1) can determine that VLM analysis is required for the image of the frame to be analyzed when the tracking information of the frame to be analyzed matches any one of the one or more VLM execution conditions.

[0090] For example, if 'VLM execution condition 2' is a condition for determining whether a new object is newly detected in the image, and a new object (flame) is detected in the image of the frame to be analyzed, it can be determined that VLM analysis is required for the image of the frame to be analyzed.

[0092] FIG. 4(b) illustrates a case where the tracking information generated for the frame to be analyzed is compared with a plurality of preset conditions (VLM execution conditions 1, 2, and 3), and the result does not meet all of 'VLM execution condition 1', 'VLM execution condition 2', and 'VLM execution condition 3'.

[0093] In this way, if the tracking information of the frame to be analyzed does not meet any of the one or more pre-set VLM execution conditions, the image detection device (1) may determine that VLM analysis is not required for the image of the frame to be analyzed.

[0094] Specifically, if no specific event or periodic monitoring point corresponding to the user is detected in the image of the frame to be analyzed, the image detection device (1) can terminate the analysis procedure for the frame to be analyzed without performing a high-computation VLM analysis step.

[0095] In this way, the image detection device (1) does not unconditionally perform VLM analysis on all frames of image information received from the camera (3), but can perform VLM analysis only on some meaningful frames selected through tracking information. In other words, the image detection device (1) can reduce unnecessary computation by applying a filter called a VLM execution condition to the tracking information generated for each frame of image information, selectively performing high-computation VLM analysis only on frames that meet the VLM execution condition, and not performing high-computation VLM analysis on frames that do not meet the condition.

[0097] FIG. 5 illustrates a process in which tracking information of an analysis target frame according to one embodiment of the present invention is determined to correspond to a plurality of VLM execution conditions.

[0099] As illustrated in FIG. 5, the VLM execution condition includes a periodic condition that determines that VLM analysis of the frame to be analyzed should be performed according to a preset time period; and the image policy may include a policy that generates the original image of the frame to be analyzed as the policy-applied image when the frame to be analyzed meets the periodic condition.

[0101] In one embodiment of the present invention, the VLM execution conditions may include periodic conditions and user-configured conditions. Specifically, the periodic condition may be a condition that forces VLM analysis to be performed according to a preset time period even if there are no special abnormal signs in the tracking information. That is, by performing VLM analysis according to the periodic condition according to a preset time period regardless of whether the user-configured condition is met, it is possible to check whether potential risk factors that were not detected by the algorithm for the user-configured condition exist in the image.

[0102] User-configured conditions are conditions set by the user to reflect the characteristics of specific events to be detected, and may be conditions that trigger VLM analysis when object state indicators of tracking information (e.g., position, velocity, acceleration, change in size, etc.) exceed a threshold or exhibit a specific pattern. Specifically, user-configured conditions may include multiple detailed conditions to detect different events. For example, 'Detailed Condition 1' can be set as a new object detection condition, 'Detailed Condition 2' as an area entry condition, and 'Detailed Condition 3' as a direction violation condition, allowing for monitoring of different situations.

[0104] As illustrated in FIG. 5, the tracking information for frame f1 satisfies the periodic condition, so the image of frame f1 can be determined as a subject for VLM analysis regardless of whether other user-configured conditions are satisfied. The tracking information for frame f2 does not satisfies the periodic condition, but since it satisfies 'Detailed Condition 1' among the user-configured conditions, the image of frame f2 can be determined as a subject for VLM analysis. The tracking information for frame f3 does not satisfies the periodic condition, but since it satisfies 'Detailed Condition 1' and 'Detailed Condition 3' among the user-configured conditions, the image of frame f3 can be determined as a subject for VLM analysis. Since the tracking information for frame f4 does not satisfies the periodic condition and all of Detailed Conditions 1, 2, and 3, the image of frame f4 can be excluded from VLM analysis.

[0106] In this way, the image detection device (1) can examine in parallel whether the tracking information for each frame corresponds to the periodic condition and the multiple user-set conditions, and if there is at least one condition that corresponds to it, the image of the frame can be determined as a target for VLM analysis.

[0108] FIG. 6 illustrates an image policy for VLM execution conditions according to an embodiment of the present invention, and a policy-applied image generated according to the image policy.

[0110] As illustrated in FIG. 6, the VLM execution condition may include a user-configured condition that determines that VLM analysis should be performed on the frame to be analyzed when the amount of change calculated by comparing the object state indicator of the frame to be analyzed and the previous frame exceeds a preset threshold.

[0111] In addition, the above image policy includes one or more detailed policies to be applied to the image of the frame to be analyzed when the tracking information of the frame to be analyzed matches the detailed condition for each of the one or more detailed conditions included in the user-configured conditions, and the VLM execution condition application step can determine whether the tracking information of the frame to be analyzed satisfies each of the one or more detailed conditions, and if satisfied, apply the detailed policy for the detailed condition to the image of the frame to be analyzed to generate a policy-applied image.

[0113] As illustrated in FIG. 6(a), the image detection device (1) may store image policies corresponding to one or more preset VLM execution conditions, and in one embodiment, the image policy may be set by a user and received from the server system (2).

[0114] An image policy refers to a set of rules defining how to process original video images during the VLM analysis stage in order to increase the recognition rate of the VLM and input the optimal image that meets the analysis objectives.

[0115] Specifically, regarding periodic conditions, an image policy may be set to generate the entire original video image of the analysis target frame as a policy-applied image without separate image processing. As described above, periodic conditions are conditions for performing VLM analysis to monitor the overall situation of the entire video without being limited to specific objects or areas. By periodically performing VLM analysis on the entire video image, the video detection device (1) can comprehensively analyze abnormal signs or changes in the overall atmosphere that the algorithm according to the VLM execution conditions might miss.

[0117] Meanwhile, different detailed policies may be matched and stored for each detailed condition of the user-configured conditions. For example, 'Detailed Policy 1' may be set for 'Detailed Condition 1', 'Detailed Policy 2' for 'Detailed Condition 2', and 'Detailed Policy 3' for 'Detailed Condition 3'. A detailed policy may correspond to an image policy set for each detailed condition of the user-configured conditions, and in one embodiment, it may be set by the user and received from the server system (2).

[0118] As mentioned above, for each detailed condition, the area of ​​interest (ROI) or object to be focused on may be set differently depending on the characteristics of the event to be detected. For example, if the condition is to detect a fire in a specific area, intensive analysis of the image centered on that area is required, and if the condition is to detect an intrusion by a specific object, intensive analysis of the image magnified around that object may be required.

[0119] The image detection device (1) can improve the accuracy and efficiency of analysis by applying an optimized image policy for each detailed condition set by the user to generate a policy-applied image, thereby limiting the range of information that the VLM must analyze for event analysis to be detected through the detailed condition and providing only key visual information.

[0120] For example, if it is a detailed condition for detecting a fire in a specific area, an image policy may be stored as a detailed policy for that condition that crops an image of the area to generate a policy-applied image.

[0122] Figure 6(b) illustrates a process of determining whether the tracking information of the frame to be analyzed meets the VLM execution conditions, and if it does, applying an image policy for the corresponding VLM execution conditions to generate a policy-applied image.

[0123] The image detection device (1) can determine whether the tracking information of the frame to be analyzed meets each of the detailed conditions (detailed condition 1, detailed condition 2, detailed condition 3) of the pre-set VLM execution conditions, namely periodic conditions and user-set conditions.

[0124] As described above, when the tracking information of the frame to be analyzed does not meet the periodic condition, detailed condition 1, and detailed condition 3, but meets detailed condition 2, the image detection device (1) can call the 'detailed policy 2' that is pre-matched and stored to the matched 'detailed condition 2', and apply the 'detailed policy 2' to the original image of the frame to be analyzed to generate a policy-applied image.

[0125] For example, if 'Detailed Policy 2' is an image policy that crops a specific area containing an object presumed to be a flame, the policy-applied image generated according to Detailed Policy 2 may correspond to an image cropped from the area surrounding the flame object, rather than the entire original image.

[0126] In this way, the image detection device (1) generates a policy-applied image, thereby allowing the VLM to perform analysis intensively only on the flame object of interest without performing analysis on unnecessary background information.

[0127] That is, according to one embodiment of the present invention, by applying an image policy that crops an image centered on an object or region that meets the VLM execution conditions and generating a policy-applied image, intensive analysis can be performed on the region related to the event to be detected among the entire image.

[0129] FIG. 7 illustrates examples of VLM execution conditions and image policies according to an embodiment of the present invention.

[0131] As illustrated in FIG. 7, the VLM execution condition includes user-configured conditions, and the user-configured conditions may include one or more detailed conditions among: a new object detection condition that determines whether a new object that was not detected in the previous frame is detected in the frame to be analyzed; a region entry condition that determines whether the position coordinates of the object enter into a preset region of interest; and a direction violation condition that determines whether the direction of movement according to the change in the position coordinates of the object differs from the allowed direction of movement set in the virtual direction of movement setting area where the object exists by more than a preset angle.

[0132] In addition, the detailed policies included in the above image policy may include one or more of the following: a policy for generating a policy-applied image by cropping an area containing the bounding box of a newly detected object when a new object that was not detected in the previous frame is detected in the frame to be analyzed; a policy for generating a policy-applied image by cropping an area containing the area of ​​interest when the position coordinates of an object enter a pre-set area of ​​interest; and a policy for generating a policy-applied image by cropping an area containing the bounding box of an object and the movement direction setting area when the movement direction resulting from the change in the position coordinates of the object differs from the allowed movement direction set in the virtual movement direction setting area where the object exists by an angle greater than or equal to a pre-set angle.

[0134] FIG. 7(a) illustrates a case where a new object detection condition is set as a detailed condition among user-configured conditions, and a policy-applied image generated according to the image policy for the new object detection condition. The new object detection condition may be a condition for determining whether a new object that did not exist in the previous frame has been detected in the frame to be analyzed. For example, the new object detection condition may be a condition for detecting cases where a 'flame' object that does not normally exist newly appears, or where a 'person' appears in a security zone.

[0135] The image detection device (1) can apply an image policy that creates a policy-applied image by cropping the image around the new object when the tracking information of the frame to be analyzed meets the new object detection condition.

[0136] In this way, by generating policy-applied images centered on new objects and inputting them into the VLM, the size of the images that the VLM needs to analyze is reduced and background noise is removed, thereby improving computation speed and enabling precise analysis of the corresponding objects.

[0138] FIG. 7(b) illustrates a case where an area entry condition is set as a detailed condition among user-configured conditions, and a policy-applied image generated according to the image policy for the area entry condition. The area entry condition may be a condition for determining whether an object has entered the area of ​​interest within the analysis target frame when the user has pre-set an area of ​​interest to specifically monitor. For example, it can be used to detect when a person or vehicle approaches a security area with controlled access or a specific area where hazardous materials are stored.

[0139] The image detection device (1) can apply an image policy to create a policy-applied image by cropping the image centered on the area of ​​interest when the location coordinates of a specific object in the tracking information meet the area entry condition, that is, when the object enters or is located within the area of ​​interest. For example, in FIG. 7 (b), as a 'person' object enters the area of ​​interest within the entire image, only the part of the area of ​​interest can be cut out to create a policy-applied image.

[0140] In this way, by generating policy application images centered on regions of interest and inputting them into the VLM, unnecessary analysis of the entire video image can be minimized, and intensive analysis of the regions of interest requiring monitoring can be performed.

[0142] Figure 7(c) illustrates a case where a direction violation condition is set as a detailed condition among user-configured conditions, and a policy-applied image generated according to the image policy for the direction violation condition. A direction violation condition may be a condition that detects whether the actual direction of movement of an object in a specific area within a video image differs from the allowed direction of movement when the allowed direction of movement for an object to move in a specific area within the video image is set in advance.

[0143] For example, as illustrated in Fig. 7(c), the user can set an allowed movement direction for each of the upward escalator area (A1) and the downward escalator area (A2) within the image. The image detection device (1) can analyze the movement path of an object in the tracking information in real time and compare the allowed movement direction of the area where each object is located with the actual movement direction.

[0144] For example, if an object moving upward in the downward escalator area (A2) is detected and the object's direction of movement differs from the set allowed direction of movement by more than a preset angle, the image detection device (1) can crop the area including the bounding box of the object and the direction of movement setting area (A2) to which the object belongs to generate a policy application image.

[0145] In this way, by cropping the object moving in the wrong direction and the area where it is located together to generate a policy-applied image, the VLM can go beyond simple detection of movement direction violations and make more accurate situational judgments by comprehensively considering interactions with the surrounding environment (e.g., the direction of travel of the escalator or other objects present on the escalator).

[0146] For example, in a situation where a downward escalator is temporarily stopped due to a malfunction, people may be allowed to walk in the opposite direction (upward) under on-site control; in such cases, if judgment is made based solely on the direction of movement of the object, a false positive may occur in which all pedestrians are mistaken for dangerous reverse movement.

[0147] On the other hand, in the present invention, by cropping the object moving in the wrong direction and the area where the object is located together to create a policy-applied image and inputting it into the VLM, if multiple people are moving in the opposite direction simultaneously, the VLM recognizes this as a special situation such as an escalator malfunction or emergency evacuation rather than an individual accident of moving in the wrong direction, thereby preventing false positives.

[0149] FIG. 8 illustrates a VLM analysis step according to one embodiment of the present invention.

[0151] As illustrated in FIG. 8, the image detection device (1) can derive an analysis score by inputting a prompt containing a generated policy-applied image and a predefined plurality of queries into the VLM.

[0152] A query can consist of natural language questions or sentences related to a specific event that the user intends to detect. For example, 'Query 1' can be defined in the form "Has a fire occurred?", 'Query 2' in the form "Is there a person in the restricted area?", and 'Query 3' in the form "Is there a person moving in the wrong direction on the escalator?".

[0153] VLM can analyze the semantic association between the input policy application image and multiple queries to output an analysis score for each query. Specifically, the analysis score can be expressed as a probability value between 0 and 1 or as a normalized score, and a higher score may indicate a higher probability that the situation described by the query matches the policy application image.

[0154] For example, when VLM determines a policy-applied image as a fire situation, the analysis score for 'fire monitoring' is calculated as high at 0.8, while other items may be calculated as low at 0.1 and 0.2.

[0156] As described above, in one embodiment of the present invention, tracking information is primarily determined to determine whether it meets the VLM execution conditions, and only when it meets, high-computation VLM analysis is selectively performed to efficiently use the limited resources of the image detection device (1). Additionally, when performing VLM analysis, instead of using the original image as is, a policy-applied image processed to include only core objects or areas related to the event according to an image policy is used, thereby allowing intensive analysis to be performed only on objects related to the event without interference from unnecessary backgrounds.

[0158] In one embodiment of the present invention, the prompt may include not only a policy application image for the frame to be analyzed, but also video images for a plurality of preset frames (e.g., 5 previous frames, 5 subsequent frames) before and after the frame to be analyzed. Through this, the VLM can understand and analyze dynamic situations or temporal contexts that are difficult to grasp with only a still image of the frame to be analyzed.

[0160] FIG. 9 illustrates a VLM analysis step according to another embodiment of the present invention.

[0162] As illustrated in FIG. 9, the VLM analysis step can be input into the VLM by adding a phrase to the prompt requesting that, among the one or more VLM execution conditions, the event or query related to the VLM execution condition to which the tracking information of the analysis target frame matches be analyzed in particular, thereby causing the VLM to evaluate the semantic relevance of the query related to the VLM execution condition as higher than other queries, and to calculate the analysis score for the query related to the VLM execution condition with weights applied.

[0164] As described above, the image detection device (1) determines whether the tracking information corresponds to each of the plurality of VLM execution conditions, and in one embodiment, the image detection device (1) may include in the prompt a phrase requesting analysis of a query or event related to a specific VLM execution condition to which the tracking information corresponds, such as "analyze fire monitoring in detail," in order to increase the accuracy of analysis of a query related to a specific VLM execution condition to which the tracking information corresponds among the plurality of VLM execution conditions.

[0165] For example, as illustrated in FIG. 9, when tracking information meets a VLM execution condition for monitoring a fire situation (e.g., a new object detection condition for monitoring the new creation of a flame object), the image detection device (1) can input it into the VLM by adding a phrase to the prompt requesting intensive analysis of fire monitoring.

[0166] Accordingly, VLM can evaluate semantic relevance more highly when performing an analysis on a query related to the VLM execution condition (Query 1 requesting an analysis of fire monitoring) compared to other queries (Query 2 requesting an analysis of entry into a restricted area, and Query 3 requesting an analysis of escalator reverse movement). As a result, when calculating the analysis score for fire monitoring, a weight is applied, which can result in a high score of 0.99.

[0168] In this way, the present invention can exert the effect of re-verifying and reinforcing the detection results in the VLM execution condition application step during the VLM analysis step by inducing the analysis of events related to the abnormal signs when performing VLM analysis secondarily on the abnormal signs detected primarily through tracking information.

[0169] In other words, when VLM analyzes the semantic association between the policy application image and each of the multiple queries, it can focus on analyzing specific events detected through tracking information to derive a higher or more accurate analysis score for those events.

[0170] As such, according to one embodiment of the present invention, by adding a phrase to the prompt to focus on analyzing events corresponding to VLM execution conditions that match tracking information, the VLM can draw attention to the event, assign weight to the analysis score, and minimize the possibility of non-detection.

[0172] FIG. 10 illustrates the section analysis score for a sliding frame section according to one embodiment of the present invention.

[0174] As illustrated in FIG. 10, the final analysis step may include: a step of setting a sliding frame section configured to include the analysis target frame and a predetermined number of adjacent frames that are temporally continuous with the analysis target frame; a step of deriving a section analysis score for each of a plurality of queries of the analysis target frame based on an analysis score derived for each of a plurality of frames included in the sliding frame section; and a step of determining that an event to be detected through a query corresponding to the section analysis score has occurred in the analysis target frame when the section analysis score exceeds a predetermined standard for the corresponding section analysis score.

[0176] As illustrated in FIG. 10 (a), the image detection device (1) can set a sliding frame section to include a predetermined number of adjacent frames that are temporally continuous with respect to the frame to be analyzed, and perform a final analysis on whether an event to be detected in the frame to be analyzed has occurred based on a section analysis score obtained by synthesizing the analysis scores derived for the sliding frame section.

[0177] Specifically, for frames for which VLM analysis has been performed in accordance with the VLM execution conditions for each frame over time, an analysis score for each query may be derived. For example, in FIG. 10, Q1, Q2, and Q3 may represent the analysis scores calculated by VLM for each of the different queries included in the prompt. The image detection device (1) combines the analysis scores derived for each of the multiple frames included in the sliding frame section to obtain a final section analysis score (Q1) for the frame to be analyzed (f3). avg , Q2. avg , Q3. avg ) can be derived.

[0178] In one embodiment of the present invention, the interval analysis score may be calculated as the average value, median value, or weighted average value of the analysis scores for each query within a sliding frame interval. For example, the average value of the Q1 scores derived from the frames (f1, f3, f4, f5) within the sliding frame interval including the analysis target frame (f3) may be the interval analysis score for Q1 (Q1. avg It can be calculated as ).

[0179] In this way, the image detection device (1) does not make a final determination of whether an event has occurred in the frame to be analyzed by relying on the analysis result of a single frame to be analyzed, that is, the analysis score (Q1, Q2, Q3) derived for the frame to be analyzed, but rather synthesizes the analysis scores for multiple frames included in the sliding frame section, which is a preset time range, to derive a section analysis score (Q1. avg , Q2. avg , Q3. avg Based on ), it is possible to make a final determination of whether an event has occurred in the frame to be analyzed.

[0180] For example, the video detection device (1) can finally determine that a fire has occurred in the frame to be analyzed when the section analysis score related to fire monitoring exceeds a preset standard.

[0181] Consequently, according to one embodiment of the present invention, even if the analysis score for a specific event is abnormally high or low due to camera noise or temporary obscuration in a specific frame, such temporary errors can be mitigated by considering the analysis scores of adjacent frames together to determine whether the event has occurred. That is, the image detection device (1) determines that an actual event has occurred only when a consistently high score is maintained, thereby preventing momentary false positives and improving the reliability of the analysis results.

[0183] Meanwhile, as shown in Fig. 10 (b), as the frame to be analyzed changes over time, the sliding frame section also moves and the section analysis score can be updated.

[0184] Specifically, the image detection device (1) can reset the sliding frame section to the section from f4 to f8 based on f6 when a new frame is received as time elapses and the frame to be analyzed corresponds to f6.

[0185] In this way, the sliding frame section is not a fixed section but can be continuously changed based on the frame to be analyzed, and the image detection device (1) can update the section analysis score (Q1.avg, Q2.avg, Q3.avg) for the frame to be analyzed f6 using the analysis score (Q1, Q2, Q3) of the frames included in the newly set sliding frame section (f4~f8) for the frame to be analyzed f6, and perform a final analysis of whether an event occurred in the frame to be analyzed f6 based thereon.

[0187] According to one embodiment of the present invention, by finally determining whether an event occurs in the frame to be analyzed based on the segment analysis score derived from the sliding frame segment that is variably set over time, temporary noise or misrecognition that may occur during single-frame analysis can be excluded and the reliability of the event occurrence determination can be ensured.

[0189] FIG. 11 illustrates the form of a query according to one embodiment of the present invention.

[0191] As illustrated in FIG. 11, the query may include a positive query describing a situation in which an event to be detected has occurred; and a negative query describing an exception situation in which the event has visual characteristics similar to the event but has not actually occurred.

[0192] In addition, the above analysis score may be calculated higher by adding as the semantic association between the video image or policy application image of the above analysis target frame and the above positive query is higher, and may be calculated lower by subtracting as the semantic association between the video image or policy application image of the above analysis target frame and the above negative query is higher.

[0193] In addition, the tracking information generation step may include a tracking information compression step that excludes from the tracking information information about an object that is assigned the same class as the object detected when an exception situation described in the negative query occurs, among a plurality of objects detected in the image.

[0195] FIG. 11 illustrates the form of a query according to one embodiment of the present invention.

[0196] As illustrated in FIG. 11, the query may include a positive query describing a situation in which an event to be detected has occurred; and a negative query describing an exception situation in which the event has visual characteristics similar to the event but has not actually occurred. Additionally, the analysis score may be calculated higher by adding as the semantic correlation between the video image or policy application image of the analysis target frame and the positive query increases, and may be calculated lower by subtracting as the semantic correlation between the video image or policy application image of the analysis target frame and the negative query increases. Furthermore, the tracking information generation step may include a tracking information compression step that excludes from the tracking information information regarding an object assigned the same class as the object detected when the exception situation described in the negative query occurs, among a plurality of objects detected in the video image.

[0198] As illustrated in FIG. 11 (a), the query may be configured to include positive and negative queries. Specifically, the query may include not only instructions to simply detect a specific event, but also instructions to explicitly exclude exceptions that are likely to be confused with the event.

[0199] For example, in a query requesting the VLM to detect a fire monitoring situation, the positive query may be text-based information defining or describing an actual hazardous situation, such as "detection of flames rising from a fire," and the negative query may be text-based information defining or describing a situation that could be mistaken for a fire, such as "non-detection of steam from a humidifier or red brakes from a car."

[0200] As such, when a VLM receives a query containing positive and negative queries regarding fire monitoring situations, it can calculate a low analysis score related to the fire situation by considering the association with the negative query, even if there are objects with visual characteristics similar to flames (such as steam) when analyzing the input image.

[0201] In other words, the analysis score can be calculated by adding when an event or object described in a positive query occurs or is detected, and subtracting when an event or object described in a negative query occurs or is detected.

[0203] As illustrated in FIG. 11 (b), the image detection device (1) can optimize tracking information based on a negative query. As described above, the tracking information may include class information for each detected object, and the image detection device (1) can check the class information of the objects detected in the image and compare whether the class of the object is the same as the class of the object defined or described in the negative query.

[0204] For example, when an exception situation defined or described by a negative query occurs, if an object (04) having the same class as the object (water vapor) is detected, the image detection device (1) can perform a process of deleting or excluding data about the object from the tracking information table.

[0205] Conversely, the image detection device (1) can maintain information about objects (03) that have the same class as the object class (flame) detected when a situation defined or described by a positive query occurs, or objects (01, O2) that are not related to a negative query, in the tracking information.

[0206] In this way, in one embodiment of the present invention, by preemptively excluding information about an object related to a negative query from the tracking information, it is possible to prevent the VLM execution condition from being triggered by an object related to a negative query during the process of determining whether the tracking information meets the VLM execution condition.

[0207] For example, when tracking information includes an object with the class 'steam' as a new object, the tracking information may be determined to meet the new object detection condition. Consequently, unnecessary VLM analysis may be performed by an object with the class 'steam' that is unrelated to the fire situation, potentially wasting computational resources.

[0208] In order to prevent such problems, the present invention excludes objects of class 'steam' from the tracking information before determining whether the tracking information meets the new object detection condition, thereby preventing VLM analysis from being performed because the tracking information meets the new object detection condition due to objects of class 'steam'. That is, according to one embodiment of the present invention, VLM analysis may not be performed due to objects of class 'steam' that are unrelated to the fire situation.

[0210] FIG. 12 schematically illustrates the internal configuration of a computing device according to one embodiment of the present invention.

[0212] The image detection device and server system illustrated in FIG. 1 described above may include the components of the computing device (11000) illustrated in FIG. 12.

[0213] As illustrated in FIG. 12, the computing device (11000) may include at least one processor (11100), memory (11200), peripheral interface (11300), input / output subsystem (I / O subsystem) (11400), power circuit (11500), and communication circuit (11600). In this case, the computing device (11000) may correspond to the image sensing device and server system illustrated in FIG. 1.

[0214] The memory (11200) may include, for example, high-speed random access memory, a magnetic disk, SRAM, DRAM, ROM, flash memory, or non-volatile memory. The memory (11200) may include software modules, instruction sets, or various other data required for the operation of the computing device (11000).

[0215] At this time, access to memory (11200) from other components, such as the processor (11100) or peripheral device interface (11300), can be controlled by the processor (11100).

[0216] The peripheral device interface (11300) can connect input and / or output peripheral devices of the computing device (11000) to the processor (11100) and memory (11200). The processor (11100) can perform various functions for the computing device (11000) and process data by executing software modules or instruction sets stored in the memory (11200).

[0217] The input / output subsystem can connect various input / output peripherals to the peripheral interface (11300). For example, the input / output subsystem may include a controller for connecting peripherals such as a monitor, keyboard, mouse, printer, or, if necessary, a touchscreen or sensor to the peripheral interface (11300). According to another aspect, input / output peripherals may be connected to the peripheral interface (11300) without passing through the input / output subsystem.

[0218] The power circuit (11500) can supply power to all or part of the components of the terminal. For example, the power circuit (11500) may include one or more power sources such as a power management system, a battery or alternating current (AC), a charging system, a power failure detection circuit, a power converter or inverter, a power status indicator, or any other components for power generation, management, and distribution.

[0219] The communication circuit (11600) can enable communication with another computing device using at least one external port.

[0220] Alternatively, as described above, the communication circuit (11600) may enable communication with other computing devices by including an RF circuit and transmitting and receiving an RF signal, also known as an electromagnetic signal.

[0221] The embodiment of FIG. 12 is merely an example of a computing device (11000), and the computing device (11000) may have some components shown in FIG. 12 omitted, additional components not shown in FIG. 12 added, or a configuration or arrangement that combines two or more components. For example, a computing device for a communication terminal in a mobile environment may include a touchscreen or sensors in addition to the components shown in FIG. 12, and the communication circuit (11600) may include a circuit for RF communication of various communication methods (WiFi, 3G, LTE, Bluetooth, NFC, Zigbee, etc.). The components that can be included in the computing device (11000) may be implemented as hardware, software, or a combination of both hardware and software, including one or more integrated circuits specialized for signal processing or applications.

[0222] Methods according to embodiments of the present invention may be implemented in the form of program instructions that can be executed through various computing devices and recorded on a computer-readable medium. In particular, the program according to the present embodiment may be configured as a PC-based program or an application dedicated to a mobile terminal. An application to which the present invention is applied may be installed on a computing device (11000) through a file provided by a file distribution system. For example, the file distribution system may include a file transmission unit (not shown) that transmits the file in response to a request from the computing device (11000).

[0224] The device described above may be implemented as a hardware component, a software component, and / or a combination of a hardware component and a software component. For example, the device and components described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and one or more software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. Additionally, other processing configurations, such as parallel processors, are also possible.

[0225] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or command the processing unit independently or collectively. Software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computing devices and stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media.

[0226] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either alone or in combination. The program instructions recorded on the medium may be those specifically designed and configured for the embodiment, or they may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc. The hardware devices described above may be configured to operate as one or more software modules to perform the operation of the embodiment, and vice versa.

[0228] Although the embodiments have been described above with reference to limited examples and drawings, those skilled in the art can make various modifications and variations from the description above. For example, suitable results may be achieved even if the described techniques are performed in a different order than described, and / or if the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents. Therefore, other implementations, other embodiments, and equivalents to the claims below are also within the scope of the claims.

Claims

Claim 1 A Vision-Language Model (VLM)-based image detection method comprising: a tracking information generation step for detecting an object in a frame of image information to be analyzed and generating and managing tracking information; a VLM execution condition application step for determining whether the tracking information of the frame to be analyzed satisfies a VLM execution condition and, if it satisfies it, generating a policy application image corresponding to the VLM execution condition; a VLM analysis step for inputting a prompt including a policy application image, an image, and a plurality of predefined queries into the VLM to derive an analysis score including semantic associations between the image or policy application image of the frame to be analyzed and each of the plurality of queries; and a final analysis step for analyzing whether an event to be detected occurs in each of the plurality of queries in the image or policy application image of the frame to be analyzed; wherein the queries include a positive query describing a situation in which the event to be detected has occurred; and a negative query describing an exception situation in which the event has similar visual characteristics but has not actually occurred. Claim 2 A method for detecting images according to claim 1, wherein the VLM execution condition includes a user-configured condition, and the user-configured condition includes one or more detailed conditions among: a new object detection condition for determining whether a new object that was not detected in the previous frame is detected in the frame to be analyzed; a region entry condition for determining whether the position coordinates of the object enter into a preset region of interest; and a direction violation condition for determining whether the direction of movement according to the change in the position coordinates of the object differs from the allowed direction of movement set in a virtual direction of movement setting area where the object exists by more than a preset angle. Claim 3 A method for detecting images according to claim 1, wherein the VLM execution condition includes a periodic condition that determines that VLM analysis for the frame to be analyzed should be performed according to a preset time period, and the image policy includes a policy that generates the original image of the frame to be analyzed as the policy-applied image when the frame to be analyzed meets the periodic condition.