Fast video search systems and methods with vision language models
The system addresses the impracticality of existing video search systems by employing fast region proposal algorithms and embedding similarity search to enable efficient video search using linguistic and visual descriptors, achieving rapid and accurate object detection without supervised models.
Patent Information
- Application Number
- PCT/US2024/054026
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-31
- Filing Date
- 2024-10-31
- Publication Date
- 2025-05-08
AI Technical Summary
Existing video search systems require supervised detector models and large amounts of labeled data, making them impractical for applications with thousands of objects or activities, or where user needs change frequently.
A system that uses fast region proposal algorithms and embedding similarity search to enable fast video search without supervised models, by processing linguistic and visual query descriptors and storing embeddings for efficient nearest neighbor searches.
Enables rapid localization and detection of objects and activities within unstructured video data using natural language queries, improving search accuracy with minimal computational overhead and no need for extensive labeled data.
Smart Images

Figure US2024054026_08052025_PF_FP_ABST
Abstract
Description
PATENT APPLICATIONFORFAST VIDEO SEARCH SYSTEMS AND METHODS WITH VISION LANGUAGE MODELSAPPLICANT:Percipient.ai, Inc.INVENTORS:Matthew Guay,San Francisco, CAVasudev Parameswaran, Fremont, CAJasvinder Singh, Newark, CARustu Seyhun Sariyildiz, Union City, CAAlison Higuera, San Jose, CAMike Higuera, San Jose, CARichard M. Lansky, Boulder, COSPECIFICATION RELATED APPLICATIONS
[0001] The present application is a conversion of U. S. Patent Application S.N. 63 / 546,749, filed 10 / 31 / 2023, and further is a continuation-in-part of U.S. patent Application S.N. 18 / 811 ,630, filed 8 / 21 / 2024, which in turn is a 371 conversion of PCT application PCT / US2023 / 020634, filed 5 / 1 / 2023, which in turn is a conversion of U.S. Patent Application S.N. 63 / 337,595, filed 5 / 2 / 2022, and claims the benefit of each of the foregoing, all of which are incorporated herein by reference.FIELD OF THE INVENTION
[0002] This invention relates generally to search methods for vision language models, and more particularly relates to systems and methods for querying video and imagery data in a fast manner without the need for a supervised detector model by providing linguistic and / or visual examples of the object sought.BACKGROUND
[0003] Object and activity detection in visual data typically requires finding and localizing objects and activities from within an unstructured set of images, such as either a video sequence or a quantity of still frame images. Various techniques have been developed for automated video analysis, but such object and activity detection is typically performed in a supervised setting where a deep neural network is trained to detect the object and / or activities of interest by feeding to the neural network a large quantity of labeled data, frequently images. Such a supervised training process can be practical when the set of objects and / or activities of interest is known beforehand and is relatively small.
[0004] However, scaling such a conventional process is difficult if not impossible in applications involving thousands of objects or activities of interest, or where the users’ needs for object / activity detection keep changing. In such situations, conventional supervised training processes become substantially unworkable because each new object / activity of interest requires that the process of data collection, manual labeling, and training must be performed all over again.
[0005] Recent research in the computer vision community has sought to break out of this onerous paradigm by open-vocabulary computer vision methods, which relate the natural language semantics of object class labels with image content to enable computer vision applications with no or low training overhead. Recent methods for segmentation and object detection offer precise localization, but suffer from latency and accuracy challenges. Meanwhile, recent methods for open-vocabulary image classification are designed to work on low-res images without dense content, making them impractical for use as-is for object or activity detection in video imagery.
[0006] The shortcomings of the existing techniques have created a long-felt need for a system and method capable of enabling fast responses to a query seeking to find and localize objects or activities without the need for a supervised detector model.SUMMARY OF THE INVENTION
[0007] In an embodiment, the present invention is a system that allows a user to query video and imagery data in a fast manner, by providing, in any combination, query descriptors in the form of linguistic and / or visual examples of what the user wishes to find, without the need for a supervised detector model. In an embodiment, the linguistic descriptors can be anything in an open vocabulary. In a first aspect of the invention, the image data is processed by fast region proposal algorithms that trade spatial precision for speed to achieve fast, resource-efficient determinations of which portions of a video frame or other image need to be processed.
[0008] The results of the region proposal algorithms are then processed by a variant of an embedding similarity search algorithm that ranks stored image content according to a weighted sum of descriptor embedding similarities. The weighted sum comprises embeddings of query descriptors provided by the user combined with object embeddings extracted from search results. As the process iterates with new search results and / or new descriptors, search accuracy rapidly improves. In some embodiments, a zero-shot or low-shot learning approach is used, that is, the models perform inference tasks for which no, or few, labeled examples were provided during training.
[0009] In an embodiment, the invention ingests video or other image data by first making a region of interest determination (ROID) to determine which parts of the image include data. Natural languages descriptions of the visual content may also be provided in at least some instances. Each region of interest is then processed by a Vision Language Model (VLM). The VLM processing of a given region of interest produces an embedding representing in at least some instances a combination of the visual content of the region together with the associated natural language descriptions of that visual content.
[0010] The resulting embeddings are stored in any convenient manner but insome embodiments are stored as a vector database that can then be used to conduct fast nearest neighbor searches.
[0011] The data store that results from the embedding can then be queried for live monitoring and alerting or can be queried at a later time. In either case the query can comprise imagery, text / natural language descriptors, or a combination of both. The query is passed through the same VLM as the ingested video or other imagery and natural language descriptors, resulting in an embedding as described above. In some instances, multiple examples of the query’s object or action exist and those multiple examples can be combined in some instances. The resulting query embeddings are then compared with embeddings in the data store, for example using the “nearest neighbor” approach to yield one or more tentative matches.
[0012] The matches are provided to a user or other downstream process, for review and confirmation or rejection. The downstream process can be automated in some embodiments. The confirmed matches are fed back to a saved search store and used in at least some instances to improve the results of the similarity search function. The matches selected as the final output can then be provided as training data for the development of a model, for example the processes described in copending patent application S.N. 18 / 811 ,630, filed 8 / 21 / 2024.
[0013] It is one object of the present invention to provide a system and method for rapidly finding and localizing objects and activities from within an unstructured set of images, where in at least some instances one or more text or natural language descriptors are associated with an image.
[0014] It is a further object of the present invention to provide a system and method for localizing objects or activities within an image through the use of query embeddings where the query comprises at least in part natural language descriptors.
[0015] It is a still further object of the present invention to provide a system and method for identifying image data including localization of an object or activity responsive to a query based in part of natural language descriptors.
[0016] A still further object of the present invention is to provide a system and method for classifying and detecting objects or activities within an image using nearest neighbor comparisons of vector-based embeddings.
[0017] Yet a further object of the present invention is to provide a system and method for rapidly responding to a query where ingested image data is represented in a vector embeddings store wherein the embeddings were developed using a vision language model, and a natural language query is represented as a vector embedding developed using a vision language model such that the response comprises images wherein the object of the query is identified and localized.
[0018] A still further object of the invention is to provide a system and method for searching stored embeddings in accordance with a visual concept where the visual concept can comprise both positive and negative descriptors and positive and negative search confirmations.
[0019] Yet a further object of the invention is to provide a system and method wherein each query descriptor, whether imagery or natural language, comprises a query element, and an embedding is developed for each query element by the vision language model used to ingest the data to be searched.
[0020] Another object of the invention is to provide a system and method in which the query forms a collection of query elements and the embeddings developed from the collection elements are used to rank search results.
[0021] Still another object of the invention is to provide a system and method that allows a user to query video and imagery data without the need for a supervised detector model by providing, in any combination, query descriptors in the form of linguistic and / or visual examples of what the user wishes to find.
[0022] These and other objects of the invention can be better appreciated from the following Detailed Description of the Invention, taken in combination with the appended Figures described below.The Figures
[0023] Figure 1 shows in process flow format an embodiment of the present invention as used for querying imagery and recorded video.
[0024] Figure 2 shows in process flow format an embodiment of the present invention as used for live video alerting.
[0025] Figure 3 shows in process flow format an embodiment of the gridding of regions of interest within an image such as a frame of video.
[0026] Figure 4 shows an embodiment of overlapping grids in accordance with the process of Figure 3.
[0027] Figure 5 shows an embodiment of a processor-based system suitable for executing the processes described herein.Detailed Description of the Invention
[0028] Referring first to Figures 1 and 2, exemplary alternative embodiments of the present invention can be appreciated, where the embodiment illustrated in Figure 1 shows a process for querying stored data, for example at a later time. The embodiment illustrated in Figure 2 shows a process for ingesting video or other imagery and processing it for live video alerting. It can be appreciated that both show a plurality of software modules which taken together, in each case comprise alternative systems and, taken together with the components shown in Figure 5, comprise a computer system where stored instructions can cause the system to perform the tasks defined by the software modules. It will further be appreciated that many of the process elements are the same in each embodiment, and thus like reference numerals are assigned to like elements for clarity and ease of reference.
[0029] At a high level, the invention comprises an input processing portion 100 which, in the embodiment of Figure 1 , comprises ingesting video or other imagery at 105 and providing that ingested data into a Region of Interest Module or Server 110. The video or other imagery can be any form of digital image files such as JPEG, PNG, DNG, PSD, CR2, NEF, or similar) and / or video files (MP4, MKV, AVCHD, MOV, 3gp, or similar). The Region of Interest Module is discussed in greater detail with reference to Figure 3 hereinafter, and functions to determine which parts of the imagery, whether video or still, need to be processed.
[0030] Once the regions of interest are identified, each such region is processed by a Vision Language Model, orVLM, 115. In the context of the present invention, a VLM is a machine learning model which accepts as input both visual data (images or videos) and language data (text), and is capable of performing inference on them jointly or simultaneously. In an embodiment, this stage of the system uses VLMs such as a neural network that performs Contrastive Language- Image Pretraining or CLIP. CLIP produces embeddings of images and text in ashared embedding space, allowing for cross-comparison between data modalities. As used herein, text inference describes the process of producing an embedding from an input text datum, and image inference describes the process of producing an embedding from an input image datum. Other VLMs that could be used in place of CLIP are FLAVA, ViLT, or Open CLIP.
[0031] The VLM 115 produces an embedding that represents the visual content of the region being processed, and may also represent natural language descriptions of that visual content. The VLM 115 develops embeddings that have a high degree of invariance to the many different appearances of the same visual content, but also can have numerous different terms for describing the visual content with natural language descriptors.
[0032] The embeddings resulting from processing of the image data in the VLM module is then augmented with metadata, shown at 120, where the metadata comprises for example, the video or image name, frame number of the video, the coordinates of the region, etc. The augmented embeddings may be stored in many different ways such as parquet formatted files, databases, etc. In an embodiment, the embeddings are stored as vector databases, shown in Figure 1 at 125, because the embeddings so stored enable fast and efficient nearest neighbor searches. Vector databases allow a trade-off between accuracy and speed, where a modest reduction in accuracy permits the retrieval of nearest neighbor embeddings dramatically faster than with many other approaches. When the embeddings store 125 is populated with a quantity of embeddings representing a suitably broad range of objects and activities, typically in a wide variety of settings and times of image capture, the ingested video or other imagery data represented in the embeddings store 125 can be searched by the system in response to a query.
[0033] For the embodiment of Figure 1 , such queries are developed in a Multimodal Interactive Search portion, indicated at 140. For the embodiment of Figure 1 , it is assumed that the query will be made later in time than when the embeddings were added to the embeddings store 125. The subject of a video search query can be interactively represented as a Visual Concept, indicated at 145. The Visual Concept is seeded initially by one or more positive descriptors and zero or more negative descriptors. By iteratively traversing search results, a user may add additional positive or negative visual or linguistic descriptors and positiveor negative search confirmations to attach new embeddings to the Visual Concept. The additions result in reweighting of search result scoring as described hereinafter in connection with Embedding Matching and Retrieval. This process produces a similarity function tuned to surface examples of the desired subject while downranking false positives that may be particular to the datasets of a given use case.
[0034] The query can, for example, comprise only natural language descriptors with no imagery data, i.e. , no images that serve as example of the object or activity sought by the query, indicated at 150 in Figure 1. The query can also comprise one or more example images as visual descriptors, indicated at 155, and can further comprise a combination of natural language descriptors and pictures or images as examples. Still further, the user’s query can comprise a collection of natural language representations / descriptors along with a collection of pictures or other images as examples. For example, a query may comprise just the term “ladder” as a natural language descriptor, or may comprise the phrase “red SUV traveling on a roadway”, and may also comprise multiple such terms. Still further, the query may comprise one or more such terms in combination with one or more example images. In an embodiment, each of the terms and each of the images forms an element of the query. Further, the visual and linguistic elements can be either positive or negative, where a descriptor is positive if it describes aspects of what the subject is, and the descriptor is negative if it describes aspects of what the subject is not. For example, if a query is searching for a red fire hydrant, “red” or “fire hydrant” are positive descriptors, while “blue” or “newsstand” are negative descriptors.
[0035] The Visual Concept can also comprise the results of previous searches, as indicated at 160, and comprises both positive and negative search confirmations. A search confirmation is feedback from the user about a search result. A search confirmation is positive if the result is an example of the desired subject, or negative if it is not an example of the desired subject. The confirmation can be either an embedding confirmation or an object confirmation. An embedding confirmation associates the confirmation’s search result with the embedding of that result’s Region of Interest (ROI) determination. An object confirmation produces a new embedding from a new ROI closely bounding an object that intersects with, butmay not be perfectly contained by, the confirmation’s search result ROI. The new ROI is obtained from an interactive segmentation III overlaid on the frame containing the search result. The search results can, in an embodiment be stored in the embeddings store 125, or can be stored separately.
[0036] In an embodiment, the Visual Concept makes use of a plurality of “importance weights”, indicated as {Wi}i=i6, where each weight is a real number and not all weights are 0, where, for example, the weights are associated with (1) positive text descriptors (Wi), (2) negative text descriptors (W2), (3) positive image descriptors (W3), (4) negative image descriptors (W4), (5) positive search confirmations (Ws), and (6) negative search confirmations (We). In such an example, additional Visual Concept elements can be added to category “i” and can be assigned weight Wi where all of the weights are used to produce the similarity function. It will be appreciated by the skilled person that if additional elements are used, there may be more than six weights, e.g. n weights. Alternatively, some of the descriptors may not be used, in which case their respective weights can be 0 or other null indicator or there can be fewer than six weights.
[0037] In an embodiment, each element of the query, whether natural language descriptor or example image, is processed through the same VLM 115 as used to develop the embeddings of the ingested imagery. In some cases, the embeddings developed for different query elements will be combined into a single embedding for computational efficiency. For example, where the query comprises many similar examples of the visual concept, resulting in too many embeddings, say N of them, it becomes computationally expensive to query all of them. In such cases, a smaller set of representative embeddings can be created, for example m where m « N Such representative embeddings can be developed using, as just some example methods, averaging, the median, or clustering of the original embeddings and then selecting the centroids of each cluster as the representatives. This can provide similar accuracy as the full set of embeddings but at considerably reduced computational cost.
[0038] Following development of embeddings for the query elements, in an embodiment those embeddings are passed to a Similarity Function Construction step 165. The Similarity Function Construction module 165 performs a similarity metric function which aggregates the embeddings developed for the query elementswith the embeddings associated with each search confirmation. In an embodiment, the aggregation is a weighted averaging of the two groups of embeddings - that is, the query embeddings and search confirmation embeddings. The similarity metric function is applied on the distances between the embeddings from the data store and the set of query and confirmation embeddings, where the output of the similarity metric function is a similarity score. In an embodiment, the metric is a weighted sum of similarities among Visual Concept embeddings, such that the weight of a positive descriptor or confirmation from category “I” is (VWS), and the weight of a negative descriptor or confirmation from category “j” is (-Wj / S), where S is the sum of the absolute values of the weights of each descriptor and embedding associated with the Visual Concept. For example, assume there are just two groups of embeddings and they have weights w1 and w2. Further, assume that there are two query embeddings q1 and q2, one from each group. Then the similarity metric function can be written as f(x) = w1 * (d(x,q1 ) + w2 * d(x,q2) where x is any candidate embedding. Thus, for a given embedding s from the embeddings store, f(s) is computed and the result represents the similarity of that given embedding s to the object that is being searched for. Negative weights can also be used in some instances, to ensure that s is as far away as possible from the negative examples.
[0039] Next, an Embedding Matching and Retrieval step, indicated at 170, is performed using the similarity function from the Similarity Function Construction step 165 and searches the embeddings store 125 for embeddings which maximize the similarity function. The vectors of the embeddings can be compared based on a nearest neighbor calculation, or other convenient approach, for example a linear scan of all stored vectors, or alternatively an approximation indexing technique such as HNSW or IVFFIat. The search results 175 from the Embedding Matching and Retrieval step comprise the embeddings from the embeddings store that yielded the highest similarity scores as determined by the constructed Similarity Metric function. Depending upon the embodiment, a specific number of results, e.g.can be preset such that the search results will be the k embeddings from the embeddings store that get the top scores on the when evaluated by the Similarity Function. In addition to the embeddings themselves, the search results can return associated metadata, for example video and frame indices and the associated search result ROI bounding box. In an embodiment, the results can be displayed in anyconvenient manner, for example in order of decreasing confidence where the metadata is presented in the user interface as an overlay positioned on corresponding video frame images.
[0040] The search results are then provided to a user where the object or activity discovered by the search process is displayed for the user to provide feedback by confirming a particular result as a good response to the query or to reject as not a good response, as indicated at 180. As discussed above, the confirmation can be either an embedding confirmation or an object confirmation. When the user or automated process conclude the search is complete, the resulting images or their representative embeddings can be used for training of a model, such as the processes for developing teacher-student models as described in U.S. patent Application S.N. 18 / 811 ,630, filed 8 / 21 / 2024, incorporated herein by reference.
[0041] In an embodiment, the search result images displayed in Search Results 175 also can have a result confirmation overlay that allows the user the option of indicating whether a search result should be added to the active Visual Concept. In some embodiments, to produce an object confirmation a user can click on an object that is the subject of a search result, and the click coordinate is fed into an interactive segmentation algorithm such as SAM or RITM. This produces a candidate object mask and a bounding box around this mask. The bounding box can be manually edited, if need be, to create a Region of Interest in the search result frame that tightly bounds the desired object. This object ROI is then passed through the VLM to generate an embedding that may better represent the object than the initial search result embedding. Confirmations may be added as additional examples of the user’s query intent, and used in a subsequent search in order to retrieve more accurate results. The nearest neighbor search process can also suppress any embeddings that are close to the embeddings corresponding to rejections.
[0042] The alternative embodiment shown in Figure 2 is particularly applicable to systems and processes intended to provide alerts from analysis of live video feeds. As a result, the embodiment of Figure 2 follows a slightly different workflow from the system and process shown in Figure 1. Generally, embeddings processed from a live video feed are matched against a collection of descriptors to score feed content for similarity to a desired visual concept as described above. Aswith the embodiment of Figure 1 , the visual concept can be iteratively refined. Content with scores exceeding a specified threshold triggers an alert for downstream handling of matching content.
[0043] More specifically, the first portion of the system of Figure 2 performs Live Video Stream Processing, shown at 210, where the incoming data is a live video stream indicated at 215. The live video stream can be in any suitable format, for example RTMP, HLS, MPEG-DASH or any other convenient format. The frames comprising the live video stream are fed to the Region of Interest module or server 110 which processes each frame as described above. The resulting regions of interest are then processed in a VLM 115, again as described above, with the output being embeddings representative of the visual content of the region or regions of interest, together with linguistic descriptors. The output of the VLM is then augmented with metadata associated with the content of the region of interest, as indicated at Embeddings + Metadata 120. The output of the Embeddings + Metadata step is provided, first, to the Embeddings Store 125, but also to an Embedding Matching step that is part of the alerting sequence.
[0044] The second portion of Figure 2 is a process for Multimodal interactive alert forming 215 and comprising the same Visual Concept Representation 140 shown in Figure 1 , where the concept is formed with the same type of text and image descriptors 150 and 155, respectively, as for Figure 1 , and also the same facility for Saved Search Result Embeddings 160. The image and linguistic descriptions are again processed by a VLM 115, as with Figure 1 , and the output is then augmented with the embeddings of the saved search results at an All Descriptors Embeddings step 225.
[0045] Embedding matching 230 is performed substantially as with Figure 1 , except that with live monitoring, the matching or comparison is between the embeddings and metadata of the live video feed and the augmented All Descriptors Embeddings 225 without retrieval from the embeddings store 125. As with Figure 1 , this comparison is, in at least some embodiments, a “nearest neighbor” search and can be done in many ways such as calculating the Euclidean distance between them and declaring a match if the distance is below a threshold. Historical data, such as previously reported matches that have been confirmed by a user, can also be retrieved. The search results 175 of the Embedding Matching step 230 are thenprovided to the user as with the process of Figure 1 , where the user can confirm or reject each result 180. Optionally, for future searches each result can be added to the search query or can be used to suppress a negative result. In addition, the search results are provided to a Threshold Alert function 235. In an embodiment, a Threshold Alert is generated when score threshold T, trigger count parameter C, and trigger duration window of W seconds meet predetermined parameters. As with the embodiment of Figure 1 , any final output of the system of Figure 2 can be used to train a downstream model, such as that disclosed in U.S. patent Application S.N. 18 / 81 1 ,630, filed 8 / 21 / 2024.
[0046] Referring next to Figure 3, the Region of Interest Service or module can be better appreciated. As used herein, the Region of Interest Server, Service or Module is an algorithm which receives as input: (1 ) an image or video frame, (2) video frame metadata such as timestamp, resolution, etc., and (3) zero or more other computer vision inference algorithms, to produce as output an Nx4 list of coordinates for N bounding boxes in the input frame for which vision-language model inference should be run. The areas indicated by the bounding boxes are regions of interest, or ROIs. This configuration provides low-latency, model-free, scene-adapted tiling sequences that provide sufficient localization for video search while enabling real-time operation.
[0047] Figure 3 illustrates a Region of Interest Determination [ROID] process for identifying Regions of Interest in a frame of video or other imagery. Many, if not most, visual monitoring and search needs can be met without a highly precise bounding box around the objects / activities of interest. It is often sufficient to bring the user’s attention to the frame where their object / activity of interest can be seen, together with a coarse bounding box within the frame. With this in mind, the ROID module partitions the imagery into relatively coarse regions where a VLM can be run. Figure 3 shows several ways in which this can be accomplished for live video. Shown in Figure 3 are four of the many options for identifying regions of interest in an image 300. The image can be provided to any of uniform gridding 305, scene geometry respecting gridding 310, optional flow-based regions 315, and segmentation derived regions 320. Each of these four approaches is explained in greater detail below, and in each case the output comprises the regions of interest 325.
[0048] Uniformly sized regions: The simplest region of interest determination creates a set of regions slightly overlapping with each other to ensure that objects near region boundaries are not missed. The following is one possible example of the way in which a grid of regions can be constructed parameterized by the image width, image height, number of horizontal regions, number of vertical regions, and the overlap fraction (0 for no overlap): def uniform_gridding(width, height, num_horizontal_regions, num_vertical_regions, overlap_fraction): regions = [] side_x = math.ceil(float(width) I (num_horizontal_regions - (num_horizontal_regions - 1 ) * overlap_fraction)) side_y = math.ceil(float(height) I (num_vertical_regions - (num_vertical_regions - 1 ) * overlap_fraction)) top = 0 for i in range(0, num_vertical_regions): left = 0 bottom = min(height - 1 , math.ceil(top + side_y - 1 )) forj in range(0, num_horizontal_regions): right - min(width - 1 , math.ceil(left + side_x - 1 )) regions. append((left, top, right, bottom)) left = math.ceil(right - side_x * overlap_fraction + 1 ) top = math.ceil(bottom - side_y * overlap_fraction + 1 ) return regions
[0049] The exemplary embodiment of Figure 4 illustrates a frame 400 gridded into twelve equal-sized regions 410 generated by running the above gridding for a 640 x 480 image with four horizontal and three vertical regions, and an overlap fraction of 0.15.
[0050] Scene Geometry Respecting Regions: Visual scene monitoring has many applications including use in security, safety, forensic investigation, etc. Thevast majority of cameras used for visual scene monitoring are static cameras. The cameras are typically installed higher up in the environment and do not change their viewpoint In such cases, one can select regions that vary in size over the image, instead of being uniformly sized as in the previous example. In an embodiment, the size of a region can, for example, be a fixed fraction of the size of an average person at that location. During processing, each region can be scaled to a fixed size making the embeddings invariant to scale.
[0051] Optical Flow Based Regions: An approach for selecting regions can be based on optical flow, a foundational computer vision task. There are a number of methods that allow calculation of the flow vector (u(x,y,t), v(x,y,t)) at each pixel (so called “dense optical flow”), where u(x,y,t) is the horizontal component and v(x,y,t) the vertical component of the movement of a scene element at (x,y) and time t. The resulting flow field can be segmented into coherent regions showing independently moving objects. The motion-coherent regions then become the regions of interest for running the VLM.
[0052] Segmentation Based Regions: An approach to select regions can also be based on semantic segmentation, also a foundational computer vision task. Semantic segmentation seeks to decompose an image into regions belonging to an object or a major part of the object. These regions become the regions of interest for running the VLM. A modern method to do this is described in Kirillov, Alexander, et al. "Segment Anything", arXiv preprint arXiv:2304.02643 (2023) and the associated website https: / / segment-anything.com / , a widely and generally applicable method for semantic segmentation to ‘cut out’ an object in an image where the Segment Anything Model is a promptable segmentation system with zero-shot generalization to unfamiliar objects and images without the need for additional training. This method provides three levels of segmentations - whole, part, and subpart of each object. Any or all of the segmentations can be used as a region of interest for the VLM.
[0053] Referring next to Figure 5, system hardware capable of executing the processes described herein can be appreciated. Image input devices 500, for example cameras, video cameras, video streams or other image sources provide digital image data and to processor and associated memory 505. In some instances the image input devices can also provide linguistic descriptors. User interface 510,which can include a display, keyboard, mouse, or other user input / output devices, permits the user to enter linguistic descriptors and queries, and also permits user review of search results and entry of confirm / reject instructions as well as other instructions. Data such as embeddings store 125 can be stored on local data store 515, or can be stored remotely on remote data store 520, accessible through the internet or other network 525.
[0054] In some embodiments described herein, plural instances may implement components, operations, or structures described as a single instance and vice versa. Likewise, individual operations of one or more embodiments may be illustrated and described collectively, one or more of the individual operations may be performed concurrently, and the operations may be performed in an order different than that illustrated. Structures and functionality presented as separate components in example configurations may be implemented as a combined structure or single component. Similarly, structures and functionalities presented as separate components may be implemented as a single component. These and other variations, modifications, additions, and improvements fall within the scope of the subject matter herein.
[0055] Embodiments described herein as including components, modules, or mechanisms may comprise either software modules (e.g., code embodied on a machine-readable medium or in a transmission signal) or hardware modules. A hardware module comprises a tangible unit configured or arranged to perform the requisite operations. In example embodiments, one or more computer systems (e.g., a standalone, client or server computer system, co-located or remote from one another) or one or more hardware modules of a computer system (e.g., a processor or a group of processors) may be configured either by software (e.g., an application or application portion) or as a hardware module that operates to perform certain operations as described herein.
[0056] In various embodiments, a hardware module may comprise dedicated circuitry or logic that is permanently configured (e.g., as a specialpurpose processor, such as a field programmable gate array (FPGA) or an application-specific integrated circuit (ASIC)) to perform certain operations. A hardware module may also comprise programmable logic or circuitry (e.g., as encompassed within one or more general-purpose processors or otherprogrammable processors) that is temporarily configured by software to perform certain operations. It will be appreciated that the implementation of a hardware module in a particular configuration may be driven by cost and time considerations.
[0057]
[0061] Embodiments in which one or more hardware modules are temporarily configured (e g., programmed), each of the hardware modules need not be configured or instantiated at any one instance in time. For example, where the hardware modules comprise a general-purpose processor configured using software, the general-purpose processor may be configured as respective different hardware modules at different times. Software may accordingly configure a processor, for example, to constitute a particular hardware module at one instance of time and to constitute a different hardware module at a different instance of time.
[0058] Embodiments in which one or more hardware modules are temporarily configured (e.g., programmed), each of the hardware modules need not be configured or instantiated at any one instance in time. For example, where the hardware modules comprise a general-purpose processor configured using software, the general-purpose processor may be configured as respective different hardware modules at different times. Software may accordingly configure a processor, for example, to constitute a particular hardware module at one instance of time and to constitute a different hardware module at a different instance of time.
[0059] The one or more processors may also operate to support performance of the relevant operations in a “cloud computing” environment or as a “software as a service” (SaaS). For example, at least some of the operations may be performed by a group of computers (as examples of machines including processors), these operations being accessible via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., application program interfaces (APIs).) The performance of certain of the operations may be distributed among the one or more processors, not only residing within a single machine, but deployed across a number of machines. In some example embodiments, the one or more processors or processor-implemented modules may be located in a single geographic location (e.g., within a home environment, an office environment, or a server farm). In other example embodiments, the one or more processors or processor-implemented modules may be distributed across a number of geographic locations.
[0060] Some portions of this specification are presented in terms of algorithms or symbolic representations of operations on data stored as bits or binary digital signals within a machine memory (e.g., a computer memory). These algorithms or symbolic representations are examples of techniques used by those of ordinary skill in the data processing arts to convey the substance of their work to others skilled in the art. As used herein, an “algorithm” is a self-consistent sequence of operations or similar processing leading to a desired result. In this context, algorithms and operations involve manipulation of physical quantities. Typically, but not necessarily, such quantities may take the form of electrical, magnetic, or optical signals capable of being stored, accessed, transferred, combined, compared, or otherwise manipulated by a machine. It is convenient at times, principally for reasons of common usage, to refer to such signals using words such as “data,” “content,” “bits,” “values,” “elements,” “symbols,” “characters,” “terms,” “numbers,” “numerals,” or the like. These words, however, are to be understood merely as convenient labels associated with appropriate physical quantities.
[0061] Unless specifically stated otherwise, terms such as “processing,” “computing,” “calculating,” “determining,” “presenting,” “displaying,” or the like may refer to actions or processes of a machine (e.g., a computer) that manipulates or transforms data represented as physical (e.g., electronic, magnetic, or optical) quantities within one or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other machine components that receive, store, transmit, or display information.
[0062] As used herein any reference to “one embodiment” or “an embodiment” means that a particular element, feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. The phrase “in an embodiment” used in various places in the specification do not necessarily all refer to the same embodiment.
[0063] As used herein, the terms “comprises,” “comprising,” “includes,” “including,” “has,” “having” or any other variation thereof, are intended to cover a non-exclusive inclusion. For example, a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to only those elements but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Further, unless expressly stated to the contrary, “or”refers to an inclusive or and not to an exclusive or. For example, a condition A or B is satisfied by any one of the following: A is true (or present) and B is false (or not present), A is false (or not present) and B is true (or present), and both A and B are true (or present).
[0064] From the foregoing, those skilled in the art will recognize that new and novel devices, systems and methods for rapidly identifying and localizing objects and activities, including multiple such objects and activities, using both imagery and natural language queries, have been disclosed, together with techniques, systems and methods for alerting a user to such detections when monitoring a live video stream. Given the teachings herein, those skilled in the art will recognize numerous alternatives and equivalents that do not vary from the invention, and therefore the present invention is not to be limited by the foregoing description, but only by the appended claims.
Claims
We claim:
1. A method for querying video and image data without the need for a supervised detector model comprising the steps of receiving a query wherein query descriptors comprise, in any combination, one or both of linguistic and visual examples of an object or action of interest, processing each of the query descriptors in a vision language model to generate an embedding representative of the query descriptors, providing embeddings representative of selected regions of image data processed by the vision language model, developing a plurality of similarity values, each representative of a weighted sum of a combination of the embeddings of the query descriptors with the embedding representative of a selected one of the regions of image data, and updating the query in response to one or more of the plurality of similarity values.
2. The method of claim 1 further comprising the step of providing the similarity values to at least one of group comprising a user and an automated process and accepting or rejecting a given region of image data based at least in part on the similarity values.
3. The method of claim 1 wherein the query descriptors comprise both linguistic examples and visual examples of an object or action of interest.
4. The method of claim 1 wherein the query is intended to detect an action.
5. The method of claim 1 wherein the query is intended to detect an object.
6. The method of claim 5 wherein the object is at least a portion of a person.
7. A system for querying video and image data without the need for a supervised detector model comprising in a processor interoperable with a data store, receiving a query wherein query descriptors comprise, in any combination, one or both of linguistic and visual examples of an object or action of interest, processing in the processor each of the query descriptors in a vision language model to generate an embedding representative of the query descriptors, receiving from the data store embeddings representative of selected regions of image data processed by the vision language model, developing in the processor a plurality of similarity values, each representative of a weighted sum of a combination of the embedding of the query descriptors with the embedding representative of a selected one of the regions of image data wherein the similarity value indicates how well a selected one of the regions of image data matches the query descriptors, and in the processor, updating the query in response to one or more of the plurality of similarity values.
8. The system of claim 7 wherein the embeddings are stored as a vector database.
9. The system of claim 7 wherein the similarity values are determined by a nearest neighbor search.
10. The method of claim 1 wherein the query descriptors comprise both positive and negative descriptors.11 . The method of claim 1 wherein the selected regions of image data are chosen from live imagery or video streams.
12. The method of claim 1 wherein the selected regions of image data are chosen from previously stored imagery or video streams.
13. The method of claim 11 further comprising the step of providing a threshold alert to initiate downstream handling in response to a similarity value exceeding a predetermined threshold.
14. The method of claim 11 wherein the embeddings representative of selected regions of image data processed by the vision language model are augmented by metadata associated with content of the selected regions of image data.
15. The method of claim 1 wherein the embedding representative of the query descriptors is a plurality of embeddings.
16. The method of claim 15 wherein embeddings developed for different query elements are combined into fewer representative embeddings.
17. The method of claim 1 wherein embeddings developed for different query elements are combined into a representative embedding.
18. The method of claim 16 wherein each representative embedding is developed using at least one of a group comprising averaging, median, or clustering of the original embeddings and then selecting the centroids of each cluster.
19. The system of claim 7 wherein the processor comprises a plurality of processors and the data store comprises a plurality of data stores.
20. A non-transitory computer readable storage medium comprising stored instructions, the instructions when executed causing at least one processor and data storage in communication therewith to: receive a query comprising a plurality of query descriptors wherein the descriptors comprise linguistic and visual examples of an object or action of interest process each of the query descriptors in a vision language model which generates an embedding representative of the query descriptors,retrieve embeddings representative of selected regions of image data processed by the vision language model, develop a plurality of similarity values, each representative of a weighted sum of a combination of the embeddings of the query descriptors with the embedding representative of a selected one of the regions of image data, and update the query in response to one or more of the plurality of similarity values.
Citation Information
Patent Citations
Efficient and fine-grained video retrieval
US20200302294A1
Modality adaptive information retrieval
US20220230061A1
Systems and methods for open vocabulary object detection
US20230154213A1