Method and device for detecting object using vision language model

KR103002833B1Active Publication Date: 2026-08-12SEO CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2026-08-12

Smart Images

  • Figure 112025098329013-PAT00003_ABST
    Figure 112025098329013-PAT00003_ABST
Patent Text Reader

Abstract

An object detection method using a vision language model executed by a processor is disclosed. The object detection method using the vision language model comprises the steps of receiving a prompt from a user, encoding the prompt to output a plurality of text vectors, encoding a single video frame generated by a camera to output a plurality of visual vectors, measuring the similarity between the outputted plurality of text vectors and the outputted plurality of visual vectors, and providing the user with information of an object corresponding to the highest similarity in the single video frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to an object detection method and apparatus using a vision language model, and more specifically, to an object detection method and apparatus using a vision language model capable of detecting an object in an image that corresponds to a prompt input by a user. Background Technology

[0002] Object detection systems are utilized for various purposes, such as monitoring traffic flow, security, and safety. Existing object detection systems adopt a method of recognizing objects based on predefined classes; however, this approach has limitations in that it is difficult to respond quickly and flexibly to diverse user requirements and dynamic environmental changes. Therefore, a new method is required that can effectively reflect users' intuitive needs and respond more flexibly. Prior art literature

[0003] Korean Registered Patent Publication No. 10-2737656 (2024.11.28.) The problem to be solved

[0004] The technical problem that the present invention aims to solve is to provide an object detection method and device capable of interacting with a user. means of solving the problem

[0005] An object detection method using a vision language model performed by a processor according to an embodiment of the present invention comprises the steps of: receiving a prompt from a user; encoding the prompt and outputting it as a plurality of text vectors; encoding a single video frame generated by a camera and outputting it as a plurality of visual vectors; measuring the similarity between the outputted plurality of text vectors and the outputted plurality of visual vectors; and providing the user with information of an object corresponding to the highest similarity in the single video frame.

[0006] The object detection method using the above-described vision language model further includes the step of determining whether the prompt is clear based on the similarity between the outputted plurality of text vectors and the outputted plurality of visual vectors, and the step of determining whether to provide information about the object to the user using a single video frame or using multiple video frames when it is determined that the prompt is not clear.

[0007] The step of determining whether the prompt is clear based on the similarity between the outputted plurality of text vectors and the outputted plurality of visual vectors includes the step of determining that the prompt is clear when the highest similarity is greater than a first threshold value, and the step of determining that the prompt is not clear when the highest similarity is less than the first threshold value.

[0008] When it is determined that the above prompt is unclear, the step of determining whether to provide information about the object to the user using a single video frame or to provide information about the object to the user using multiple video frames includes: a step of determining to provide information about the object to the user using a single video frame when the highest similarity is greater than a second threshold value; and a step of determining to provide information about the object to the user using multiple video frames when the highest similarity is less than the second threshold value.

[0009] The step of determining to provide information of the object to the user using a plurality of video frames when the highest similarity is less than the second threshold value includes the step of encoding the plurality of video frames generated by the camera to output a plurality of different visual vectors, the step of measuring the similarity between the output plurality of text vectors and the plurality of different visual vectors, and the step of providing information of the object corresponding to the highest similarity in the plurality of video frames to the user. The plurality of different visual vectors are related to actions.

[0010] An apparatus according to an embodiment of the present invention includes a processor that executes object detection commands using a vision language model, and a memory that stores object detection commands using the vision language model.

[0011] Object detection commands using the above vision language model are implemented to receive a prompt from a user, encode the prompt to output multiple text vectors, encode a single video frame generated by a camera to output multiple visual vectors, measure the similarity between the output multiple text vectors and the output multiple visual vectors, and provide the user with information about the object corresponding to the highest similarity in the single video frame.

[0012] Object detection commands using the above vision language model determine whether the prompt is clear based on the similarity between the output multiple text vectors and the output multiple visual vectors, and when it is determined that the prompt is not clear, determine whether to provide information about the object to the user using a single video frame or to provide information about the object to the user using multiple video frames.

[0013] Commands for determining whether the prompt is clear based on the similarity between the output multiple text vectors and the output multiple visual vectors are implemented such that when the highest similarity is greater than a first threshold value, the prompt is determined to be clear, and when the highest similarity is less than the first threshold value, the prompt is determined to be unclear.

[0014] When it is determined that the above prompt is unclear, commands determining whether to provide information about the object to the user using a single video frame or to provide information about the object to the user using multiple video frames are implemented such that when the highest similarity is greater than a second threshold value, the information about the object is provided to the user using a single video frame, and when the highest similarity is less than the second threshold value, the information about the object is provided to the user using multiple video frames.

[0015] Commands determining to provide information of the object to the user using multiple video frames when the highest similarity is less than the second threshold value are implemented to encode the multiple video frames generated by the camera to output multiple different visual vectors, measure the similarity between the output multiple text vectors and the multiple different visual vectors, and provide information of the object corresponding to the highest similarity in the multiple video frames to the user, and the multiple different visual vectors are related to the operation. Effects of the invention

[0016] The object detection method and device using a vision language model according to an embodiment of the present invention provide the user with information about an object corresponding to a prompt entered by the user, thereby enabling the user to focus on monitoring a desired object. Brief explanation of the drawing

[0017] Detailed descriptions of each drawing are provided to help to more fully understand the drawings cited in the detailed description of the present invention. Figure 1 shows a block diagram of an object detection system using a vision language model according to an embodiment of the present invention. Figure 2 shows a block diagram of a vision language model according to an embodiment of the present invention. FIG. 3 shows a diagram for explaining the operation of a vision language model according to an embodiment of the present invention. FIG. 4 shows another block diagram of a vision language model according to an embodiment of the present invention. FIG. 5 shows a flowchart for explaining an object detection method using a vision language model according to an embodiment of the present invention. FIG. 6 shows a flowchart for explaining a method for determining whether a prompt is clear based on the similarity between an output text vector and a plurality of output visual vectors according to an embodiment of the present invention. Specific details for implementing the invention

[0018] Specific structural or functional descriptions regarding embodiments according to the concept of the present invention disclosed herein are provided merely for the purpose of explaining embodiments according to the concept of the present invention, and embodiments according to the concept of the present invention may be implemented in various forms and are not limited to the embodiments described herein.

[0019] Embodiments according to the concept of the present invention may be subject to various modifications and may take various forms; therefore, embodiments are illustrated in the drawings and described in detail in this specification. However, this is not intended to limit the embodiments according to the concept of the present invention to specific disclosed forms, and includes all modifications, equivalents, or substitutions that fall within the spirit and scope of the present invention.

[0020] Terms such as "first" or "second" may be used to describe various components, but said components should not be limited by said terms. For the sole purpose of distinguishing one component from another, for example, without departing from the scope of rights according to the concept of the present invention, the first component may be named the second component, and similarly, the second component may be named the first component.

[0021] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. Conversely, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between. Other expressions describing the relationship between components, such as "between" and "exactly between," or "adjacent to" and "directly adjacent to," should be interpreted in the same way.

[0022] The terms used herein are used merely to describe specific embodiments and are not intended to limit the invention. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this specification, terms such as “comprising” or “having” are intended to indicate the existence of the described features, numbers, steps, actions, components, parts, or combinations thereof, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0023] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as generally understood by those skilled in the art to which the present invention pertains. Terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in this specification.

[0024] Hereinafter, the present invention will be described in detail by explaining preferred embodiments of the present invention with reference to the attached drawings.

[0025] Figure 1 shows a block diagram of an object detection system using a vision language model according to an embodiment of the present invention.

[0026] An object detection system (100) using a vision language model refers to a system capable of detecting objects in video frames captured by a camera (30) using a vision language model. The object detection system (100) using a vision language model includes a server (10) and a camera (30). The server (10) and the camera (30) can communicate with each other through a network (101).

[0027] The server (10) refers to an electronic device implemented as a laptop, smartphone, tablet PC, or desktop. The server (10) includes a processor (11) that executes object detection commands using a vision language model and a memory (13) that stores said object detection commands. The server (10) can be connected to a display (20). An administrator (3) using the server (10) can view multiple video frames received from a camera (30) through the display (20).

[0028] A camera (30) is installed on a road or building, etc., to photograph the surroundings (200) and generate multiple video frames. The multiple video frames generated by the camera (30) can be transmitted to a server (10) via a network (101).

[0029] FIG. 2 shows a block diagram of a vision language model according to an embodiment of the present invention. FIG. 3 shows a diagram for explaining the operation of a vision language model according to an embodiment of the present invention.

[0030] Referring to FIGS. 1 to 3, the processor (11) executes commands related to the VLM (Vision Language Model, 300). The VLM (300) refers to an artificial intelligence system that includes computer vision and natural language processing for simultaneously understanding visual data and text data. Depending on the embodiment, the VLM (300) may be referred to as a VLM engine. For example, the VLM (300) may be implemented as VisualBERT, Visual ChatGPT, Flamingo, or CLIP.

[0031] The VLM (300) includes a video encoder (310), a text encoder (320), and a decoder (330). It should be understood that the operation of the VLM (300) is performed by the processor (11). That is, it should be understood that the operation of the video encoder (310), the text encoder (320), and the decoder (330) is performed by the processor (11).

[0032] A video encoder (310) is implemented as a neural network to extract and encode features associated with a single video frame (210) generated by a camera (30). For example, the video encoder (310) includes one or more fully connected layers or partially connected layers to extract image patches from the video frame (210) and to encode features of the video frame (210). The layers include multiple nodes. For example, the video encoder (310) may be implemented as a ResNet or a Vision Transformer (ViT), etc. The video encoder (310) encodes a single video frame (210) generated by the camera (30) and outputs multiple visual vectors (311).

[0033] A user (3) of the server (10) can input a prompt using an input means such as a keyboard, touchpad, or voice. For example, the user (3) can input a prompt (260) saying "Find a suspicious person." The term query may be used instead of the above prompt.

[0034] The text encoder (320) includes a tokenizer, an embedding layer, and a transformer encoder stack.

[0035] The tokenizer converts the input prompt (260) into tokens, that is, a token sequence. Each token sequence has an ID, and each ID corresponds to subwords, words, and characters.

[0036] The embedding layer converts each ID into a vector.

[0037] The Transformer encoder stack includes multiple layers. Each layer includes a multi-head self-attention, a feed-forward neural network (FFN), residual connections, and a layer norm.

[0038] Multi-head self-attention determines whether each token is related to other tokens.

[0039] Feed-forward neural networks (FFNs) are used to further learn features. Feed-forward neural networks (FFNs) include hidden layers and activation functions.

[0040] Residual connections and LayerNorm are used to make neural network training easier and more stable.

[0041] Multi-head self-attention, feed-forward neural networks (FFN), or residual connections and LayerNorm are widely known concepts, so a detailed explanation thereof is omitted.

[0042] The text encoder (320) outputs the prompt (260) entered by the user (3) as a plurality of text vectors (321).

[0043] The decoder (330) measures the similarity between the output multiple text vectors (321) and the output multiple visual vectors (311). The similarity is represented by a matrix (331) in FIG. 3. The similarity can be calculated as a dot product or cosine similarity between the output multiple text vectors (321) and the output multiple visual vectors (311). The similarity has a range from -1 to 1. A similarity closer to 1 indicates a stronger similarity, while a similarity closer to -1 indicates a weaker similarity. For example, in the matrix (331) of FIG. 3, 'I1T1' represents the similarity between one of the multiple text vectors (321) (I1) and one of the multiple visual vectors (311) (T1).

[0044] The decoder (330) provides the user (3) with information about the object corresponding to the highest similarity in a single video frame (210). The information about the object can be expressed as text (340). For example, when the prompt (260) says "Find the suspicious person," the decoder (330) can output text (340) such as "The suspicious person in the video frame (210) is likely a person running."

[0045] FIG. 4 shows another block diagram of a vision language model according to an embodiment of the present invention.

[0046] Referring to FIGS. 1 to 4, according to an embodiment, the processor (11) may further determine whether the highest similarity among the similarities measured between a plurality of output text vectors (321) and a plurality of output visual vectors (311) is greater than a first threshold value (e.g., 0.5). This is to determine whether the prompt (260) is clear.

[0047] When the highest similarity among the similarities measured between the output multiple text vectors (321) and the output multiple visual vectors (311) is determined to be greater than the first threshold value (e.g., 0.5), the processor (11) determines that the prompt (260) is clear.

[0048] When the prompt (260) is determined to be clear, the processor (11) provides the user (3) with information about the object corresponding to the highest similarity in one video frame (210) as text (340). That is, the text (340) can be displayed on the display (20).

[0049] When the highest similarity among the similarities measured between the output multiple text vectors (321) and the output multiple visual vectors (311) is less than the first threshold value (e.g., 0.5), the processor (11) determines that the prompt (260) is unclear.

[0050] When it is determined that the prompt (260) is unclear, the processor (11) decides whether to provide information about the object to the user (3) using one video frame (210) or to provide information about the object to the user (3) using multiple video frames (220).

[0051] In order to determine whether to provide information about the object to the user (3) using a single video frame (210) or to provide information about the object to the user (3) using multiple video frames (220), the processor (11) determines whether the highest similarity among the similarities measured between the output multiple text vectors (321) and the output multiple visual vectors (311) is greater than a second threshold value (e.g., 0.3) (S710). The first threshold value (e.g., 0.5) is greater than the second threshold value (e.g., 0.3).

[0052] When the highest similarity is greater than the second threshold value (e.g., 0.3), the processor (11) decides to provide information about the object to the user (3) using one video frame (210). For example, the prompt (260) may be "Find a big tree."

[0053] Assuming that a single video frame (210) contains multiple trees and the height of the included trees is greater than a certain height (e.g., 5m), the prompt (260) for "find a big tree" is determined to be clear. This is because when the video encoder (310) is trained, trees greater than 5m are trained as big trees. Assuming that a single video frame (210) contains multiple trees and the height of the included trees is greater than a certain height (e.g., 5m), the highest similarity among the similarities measured between the output multiple text vectors (321) and the output multiple visual vectors (311) is greater than the first threshold value (e.g., 0.5).

[0054] According to an embodiment, when multiple trees are included in a single video frame (210) and the height of the trees is assumed to be less than a certain height (e.g., 2m), the prompt (260) for "find a big tree" is determined to be unclear. This is because when the video encoder (310) is trained, trees less than 2m are not trained as big trees. When multiple trees are included in a single video frame (210) and the height of the trees is assumed to be less than a certain height (e.g., 2m), the highest similarity among the similarities measured between the output multiple text vectors (321) and the output multiple visual vectors (311) is between the first threshold value (e.g., 0.5) and the second threshold value (e.g., 0.3).

[0055] When the highest similarity is between the first threshold value (e.g., 0.5) and the second threshold value (e.g., 0.3), the processor (11) decides to provide information about the object to the user (3) using one video frame (210). The processor (11) uses one video frame (210) to output text (340) to the user (3) indicating what the big tree is in one video frame (210). For example, the text (340) may be "The big tree is the tree in the upper left of one video frame (220)." However, in this case, the big tree included in the text (340) refers to a tree that is relatively larger than the remaining trees included in one video frame (210), and may not be an objectively large tree.

[0056] According to another embodiment, assuming that multiple people are included in a single video frame (210), the prompt (260) for "find the suspicious person" is determined to be unclear. This is because the video encoder (310) cannot be trained on suspicious behavior with only a single video frame (210). Assuming that multiple people are included in a single video frame (210), the highest similarity among the similarities measured between the output multiple text vectors (321) and the output multiple visual vectors (311) is less than the second threshold value (e.g., 0.3).

[0057] When the highest similarity is less than the second threshold value (e.g., 0.3), the processor (11) decides to provide information about the object to the user (3) using a plurality of video frames (220).

[0058] The VLM (300) further includes a second video encoder (315). The second video encoder (315) is used when the highest similarity is less than a second threshold value (e.g., 0.3). The video encoder (310) may be referred to as the first video encoder, and the video encoder (315) may be referred to as the second video encoder.

[0059] The first video encoder (310) can be implemented with an image encoder and a pooling layer. The image encoder treats each of the multiple video frames (220) as a single image, encodes each of the multiple video frames (220), and outputs multiple visual vectors. The pooling layer pools the output multiple visual vectors. By implementing the first video encoder (310) with an image encoder and a pooling layer, multiple visual vectors can be output quickly.

[0060] The second video encoder (315) can be implemented as a VideoCILP or TimeSformer, etc. Unlike the first video encoder (310), the second video encoder (315) can be trained to recognize the actions of an object (e.g., walking, fighting, stealing, suspicious actions, etc.) in multiple video frames (220). The second video encoder (315) encodes multiple video frames (220) and outputs multiple visual vectors related to the actions of the object. The suspicious actions may refer to actions such as looking around.

[0061] The decoder (330) measures the similarity between the outputted multiple text vectors (321) and the multiple visual vectors (311) output from the first video encoder (310) and the second video encoder (315), and provides information about the object to the user (3) as text (340). For example, the text (340) may be "The suspicious person in the multiple video frames (220) is the person running across the crosswalk."

[0062] Specifically, the decoder (330) calculates first cosine similarities between multiple text vectors (321) output from the text encoder (320) and first visual vectors output from the first video encoder (310), and calculates second cosine similarities between multiple text vectors (321) output from the text encoder (320) and second visual vectors output from the second video encoder (315). The decoder (330) calculates the first cosine similarities and the second cosine similarities to calculate a score. This is calculated as shown in the following mathematical formula.

[0063] [Mathematical Formula 1]

[0064] Score = αCos1+(1-α)CoS2

[0065] The above Score represents the score, the above α represents the weight, the above Cos1 represents the first cosine similarity calculated between a plurality of text vectors (321) output from the text encoder (320) and a first visual vector output from the first video encoder (310), and the above CoS2 represents the second cosine similarity calculated between a plurality of text vectors (321) output from the text encoder (320) and a second visual vector output from the second video encoder (315).

[0066] The above weight has a value between 0 and 1. The weight has a smaller value as the highest similarity, which is smaller than the second threshold value, is smaller. For example, when the highest similarity, which is smaller than the second threshold value, has a range greater than 0.1 and less than 0.3, the weight is 0.2, and when the highest similarity, which is smaller than the second threshold value, has a range greater than 0.3 and less than 0.5, the weight may be 0.4. That is, by making the second cosine similarity calculated between the multiple text vectors (321) output from the text encoder (320) output from the second video encoder (315) and the second visual vectors output from the second video encoder (315) more reliable, the ambiguity of the prompt (260) is set to be related to the action. The fact that the highest similarity is smaller than the second threshold value may not always mean that it is related to the action. That is, the above weight can be set to make the first cosine similarity calculated between the plurality of text vectors (321) output from the text encoder (320) and the first visual vectors output from the first video encoder (310) more reliable.

[0067] The processor (11) provides the user (3) with information about the object corresponding to the highest score in a plurality of video frames (220).

[0068] FIG. 5 shows a flowchart for explaining an object detection method using a vision language model according to an embodiment of the present invention.

[0069] Referring to FIGS. 1 to 5, the processor (11) receives a prompt (260) from the user (3) (S100). For example, the prompt (260) may be "find a suspicious person."

[0070] The processor (11) encodes the prompt (260) and outputs it as a plurality of text vectors (321) (S200).

[0071] The processor (11) encodes a single video frame (210) generated by the camera (30) and outputs a plurality of visual vectors (311) (S300).

[0072] The processor (11) measures the similarity between the output multiple text vectors (321) and the output multiple visual vectors (311) (S400).

[0073] The processor (11) determines whether the prompt (260) is clear based on the similarity between the output multiple text vectors (321) and the output multiple visual vectors (311) (S500).

[0074] FIG. 6 shows a flowchart for explaining a method for determining whether a query is clear based on the similarity between an output text vector and a plurality of output visual vectors according to an embodiment of the present invention.

[0075] Referring to FIGS. 1 to 6, the processor (11) determines whether the highest similarity among the similarities measured between the outputted plurality of text vectors (321) and the outputted plurality of visual vectors (311) is greater than a first threshold value (e.g., 0.5) (S510).

[0076] When the highest similarity among the similarities measured between the output multiple text vectors (321) and the output multiple visual vectors (311) is determined to be greater than a first threshold value (e.g., 0.5), the processor (11) determines that the prompt (260) is clear (S520).

[0077] When the prompt (260) is determined to be clear, the processor (11) provides the user (3) with information about the object corresponding to the highest similarity in one video frame (210) (S600).

[0078] When the highest similarity among the similarities measured between the output multiple text vectors (321) and the output multiple visual vectors (311) is less than the first threshold value (e.g., 0.5), the processor (11) determines that the prompt (260) is unclear.

[0079] When it is determined that the prompt (260) is unclear, the processor (11) decides whether to provide information about the object to the user (3) using one video frame (210) or to provide information about the object to the user (3) using multiple video frames (220) (S700).

[0080] In order to determine whether to provide information about the object to the user (3) using a single video frame (210) or to provide information about the object to the user (3) using multiple video frames (220), the processor (11) determines whether the highest similarity among the similarities measured between the output multiple text vectors (321) and the output multiple visual vectors (311) is greater than a second threshold value (e.g., 0.3) (S710). The first threshold value (e.g., 0.5) is greater than the second threshold value (e.g., 0.3).

[0081] When the highest similarity is greater than the second threshold value (e.g., 0.3), the processor (11) decides to provide information about the object to the user (3) using one video frame (210) (S720).

[0082] When the highest similarity is less than the second threshold value (e.g., 0.3), the processor (11) decides to provide information about the object to the user (3) using a plurality of video frames (220) (S730).

[0083] The present invention has been described with reference to an exemplary embodiment illustrated in the drawings, but this is merely illustrative, and those skilled in the art will understand that various modifications and equivalent alternative embodiments are possible therefrom. Accordingly, the true technical scope of protection of the present invention should be determined by the technical spirit of the appended claims. Explanation of the symbols

[0084] 100: System; 3: Administrator; 10: Server; 11: Processor; 13: Memory; 20: Display; 30: Camera; 101: Network; 260: Prompt; 300: Vision Language Model; 310, 315: Video encoder; 320: Text encoder; 330: Decoder; 340: Text output;

Claims

Claim 1 A method for object detection using a vision language model executed by a processor, comprising: receiving a prompt from a user; encoding the prompt and outputting it as a plurality of text vectors; encoding a single video frame generated by a camera and outputting it as a plurality of visual vectors; measuring the similarity between the outputted plurality of text vectors and the outputted plurality of visual vectors; determining that the prompt is clear when the highest similarity among the measured similarities is greater than a first threshold value; determining that the prompt is not clear when the highest similarity is less than the first threshold value; determining to provide information about the object to the user using the single video frame when the prompt is determined not to be clear and the highest similarity is greater than a second threshold value; and determining to provide information about the object to the user using a plurality of video frames when the prompt is determined not to be clear and the highest similarity is less than the second threshold value. Claim 2 delete Claim 3 delete Claim 4 delete Claim 5 In claim 1, the step of determining to provide information of the object to the user using a plurality of video frames when the highest similarity is smaller than the second threshold value comprises: encoding the plurality of video frames generated by the camera to output a plurality of different visual vectors; measuring the similarity between the outputted plurality of text vectors and the plurality of different visual vectors; and providing information of the object corresponding to the highest similarity in the plurality of video frames to the user, wherein the plurality of different visual vectors are an object detection method using a vision language model related to motion. Claim 6 A device comprising: a processor that executes object detection commands using a vision language model; and a memory that stores object detection commands using the vision language model, wherein the object detection commands using the vision language model receive a prompt from a user, encode the prompt and output it as a plurality of text vectors, encode a single video frame generated by a camera and output it as a plurality of visual vectors, measure the similarity between the outputted plurality of text vectors and the outputted plurality of visual vectors, and when the highest similarity among the measured similarities is greater than a first threshold value, determine that the prompt is clear; when the highest similarity is less than the first threshold value, determine that the prompt is not clear; when the prompt is determined not to be clear and the highest similarity is greater than a second threshold value, determine to provide information about the object to the user using the single video frame; and when the prompt is determined not to be clear and the highest similarity is less than the second threshold value, provide information about the object to the user using a plurality of video frames.

Citation Information

Patent Citations

  • Instance level scene recognition with a vision language model

    KR102737656B1

  • Method and computer system for inference using a vision-language model based on cached information associated with input prompt

    KR102798633B1

  • Method and apparatus for searching video section based artificial intelligence

    KR1020240057254A

  • Mining unlabeled images with vision and language models for improving object detection

    US20230281858A1