A method and a system for selecting video frames

The method optimizes video frame selection for large language models by using a frame controller process with sub-processes to analyze and select relevant frames in real-time, addressing the resource-intensive challenge of processing streaming video content.

WO2026024617A1PCT designated stage Publication Date: 2026-01-29ENDLESS TECH LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/038464
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-07
Filing Date
2025-07-21
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Feeding streaming video content to large language models (LLM) is expensive and resource-intensive, requiring large storage and processing power, making it impractical.

Method used

A method and system for selecting relevant video frames using a frame controller process with sub-processes that analyze video streams in real-time, optimizing frame selection based on user-defined criteria and latency constraints, and communicating selected frames to a destination within a predetermined latency period.

Benefits of technology

Reduces the computational load and storage requirements by selectively processing only relevant frames, enhancing the efficiency and cost-effectiveness of video processing by large language models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025038464_29012026_PF_FP_ABST
    Figure US2025038464_29012026_PF_FP_ABST
Patent Text Reader

Abstract

A computer-implemented method for configuring a system, the method comprising the actions of: obtaining a large language model, providing the large language model with a system-prompt comprising configuration options of the system, obtaining from a user a prompt describing the purpose of the system, obtaining, from the large language model, an output, the output comprising configuration data for the system, the configuration data complying with the user-prompt, and communicating the configuration data to the system.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] A Method and a System for Selecting Video Frames

[0002] FIELD

[0003] The method and apparatus disclosed herein are related to the field of artificial intelligence (Al) and data optimization. More particularly but not exclusively to processing streaming content by Al and / or Machine Learning (ML) systems. More particularly but not exclusively to processing video by Al. And more particularly but not exclusively to processing video by a large language model (LLM)

[0004] BACKGROUND

[0005] Artificial intelligence (Al) and particularly large language models (LLM) are expensive, requiring large storage systems and much processing power. Feeding streaming content such as audio and particularly video is therefore very expensing, both requiring even larger storage systems and more processing power. Feeding streaming content such as audio and particularly video to train a large language model (LLM), or to be analyzed by an LLM is therefore practically prohibitive. There is therefore a need for a method and a system that may overcome these deficiencies.

[0006] SUMMARY OF THE INVENTION

[0007] According to one exemplary embodiment, there is provided a computer- implemented method for configuring a system, the method including: obtaining a large language model, providing the large language model with a system-prompt including configuration options of the system, obtaining from a user a prompt describing the purpose of the system, obtaining, from the large language model, an output, the output including configuration data for the system, the configuration data complying with the user-prompt, communicating the configuration data to the system.

[0008] According to another exemplary embodiment, the system also includes at least two processing modules, the system-prompt includes at least one requirement for each of the at least two processing modules, the user-prompt includes a characteristic, and the configuration data includes at least one of the modules, which requirement complies with the characteristic. According to yet another exemplary embodiment, the system-prompt includes at least one requirement in the form of a time constraint, the user-prompt includes a characteristic, and the configuration data includes at least one time constraint that complies with the characteristic.

[0009] According to still another exemplary embodiment, the computer-implemented method includes: receiving a video streaming data including a first plurality of video frames, selecting at least one frame from the first plurality of video frames, and communicating the at least one selected video frame to a destination, where the communication of the selected video frames is performed within a predetermined latency period after the start of receiving the video streaming data.

[0010] Further according to another exemplary embodiment, the method additionally includes receiving a video streaming data including receiving a group of frames, and where the predetermined latency period is measured from the first frame of the group of frames.

[0011] Further according to another exemplary embodiment, the method for selecting frames from a video stream includes receiving a video streaming data including a first plurality of video frames, selecting at least one frame from the first plurality of video frames, and communicating the selected video frame(s) to a destination. Wherein the selecting of the frame(s) includes: generating visual embeddings for each frame, comparing vectors to determine frames of a group, and selecting at least one frame representative of the group.

[0012] Yet further according to another exemplary embodiment, the computer- implemented method for selecting frames from a video stream includes: receiving a video streaming data including a first plurality of video frames, selecting at least one frame from the first plurality of video frames, and communicating the at least one selected video frame to a destination, where the action of selecting at least one frame includes: detecting at least one object within each frame, comparing objects detected in different frames to determine frames of a group, and selecting at least one frame representative of the group. Still further according to another exemplary embodiment, the computer- implemented method for selecting frames from a video stream includes: receiving a definition of at least one undesired object, receiving a video streaming data including a first plurality of video frames, selecting at least one frame from the first plurality of video frames, and communicating the at least one selected video frame to a destination, where the action of selecting at least one frame includes: detecting at least one undesired object within a frame, and deselecting at least one frame including said undesired object.

[0013] Even further according to another exemplary embodiment, the computer- implemented method for selecting frames from a video stream includes: receiving a video streaming data including a first plurality of video frames, selecting at least one frame from the first plurality of video frames, and communicating the at least one selected video frame to a destination, where the action of selecting at least one frame includes: detecting at least one moving object within at least two frames, measunng the distance passed by the object between the at least two frames, and if the distance is larger than a predetermined value selecting at least one frame including the moving object.

[0014] Additionally, according to another exemplary embodiment, the computer- implemented method for selecting frames from a video stream includes: receiving a video streaming data including a first plurality of video frames, selecting at least one frame from the first plurality of video frames, and communicating the at least one selected video frame to a destination, wherein the selecting at least one frame includes: detecting a change of a camera field of view between at least two frames, measuring the change between the at least two frames, and if the change is larger than a predetermined value, selecting at least one frame including the change.

[0015] Yet according to another exemplary embodiment, any of the methods above additionally includes repeating the action of selecting at least one frame until a latency period is about to elapse.

[0016] Still according to another exemplary' embodiment there is provided a computer- implemented method for selecting frames from a video stream, the method including: receiving a video streaming data comprising a first plurality' of video frames, performing at least one of: selecting at least one frame from the first plurality of video frames to form a reduced frame set, and removing at least one frame from the first plurality of video frames to form a reduced frame set, and communicating the reduced frame set to a Large Language Model.

[0017] Further according to another exemplary embodiment there is provided a computer-implemented method for selecting frames from a video stream additionally including at least one of: performing the at least one of selecting at least one frame and removing at least one frame in real-time, performing the at least one of selecting at least one frame and removing at least one frame within a latency period, and repeating the performing the at least one of selecting at least one frame and removing at least one frame until a latency period is about to elapse.

[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the relevant art. The materials, methods, and examples provided herein are illustrative only and not intended to be limiting. Except to the extent necessary' or inherent in the processes themselves, no particular order to steps or stages of methods and processes described in this disclosure, including the figures, is intended or implied. In many cases the order of process steps may vary without changing the purpose or effect of the methods described.

[0019] BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Various embodiments are described herein, by way of example only, with reference to the accompanying drawings. With specific reference now to the drawings in detail, it is stressed that the particulars shown are by way of example and for purposes of illustrative discussion of the preferred embodiments only, and are presented in order to provide what is believed to be the most useful and readily understood description of the principles and conceptual aspects of the embodiment. In this regard, no attempt is made to show structural details of the embodiments in more detail than is necessary for a fundamental understanding of the subject matter, the description taken with the drawings making apparent to those skilled in the art how the several forms and structures may be embodied in practice. In the drawings:

[0021] Fig. 1 is a simplified block diagram of a frame selection system;

[0022] Fig. 2 is a simplified flow chart of one type of a perceived visual selector, which is one of a plurality of sub-processes operated by a frame controller process of the frame selection system;

[0023] Fig. 3 is a simplified flow chart of another type of the perceived visual selector;

[0024] Fig. 4 is a simplified flow chart of object detector selector, which may be one of the sub-processes operated by the frame controller; and

[0025] Fig. 5 is a simplified flow chart of large language model frame selector, which may be one of the sub-processes operated by the frame controller.

[0026] DESCRIPTION OF THE EMBODIMENTS

[0027] The present embodiments comprise a method, one or more devices, and one or more software programs for selecting video frames to be further processed by an artificial intelligence (Al) system or software, and particularly (but not exclusively), for processing by a large language model (LLM).

[0028] Current systems for video processing and machine learning often prioritize human-perceived quality in data compression. However, machine learning models have distinct requirements for data that can impact their performance and energy consumption during both learning and inference. There is a need for a method to efficiently select particular video frames for further machine learning.

[0029] The present embodiments provide a system and method for optimizing machine learning and inference processes by selecting frames a from video stream. The system uses task-specific data selection. Among these tasks, it uses a large language model (LLM) to identify relevant frames.

[0030] The principles and operation of the system, a method, and / or a computer program for selecting video frames according to the several exemplary embodiments may be better understood with reference to the following drawings and accompanying description.

[0031] Before explaining at least one embodiment in detail, it is to be understood that the embodiments are not limited in its application to the details of construction and the arrangement of the components set forth in the following description or illustrated in the drawings. Other embodiments may be practiced or carried out in various ways. Also, it is to be understood that the phraseology and terminology employed herein is for the purpose of description and should not be regarded as limiting.

[0032] In this document, an element of a drawing that is not described within the scope of the drawing and is labeled with a numeral that has been described in a previous drawing has the same use and description as in the previous drawings. Similarly, an element that is identified in the text by a numeral that does not appear in the drawing described by the text, has the same use and description as in the previous drawings where it was described.

[0033] The drawings in this document may not be of any scale. Different Figures may use different scales and different scales can be used even within the same drawing, for example different scales for different views of the same object or different scales for the two adjacent objects.

[0034] The phrases ‘at least one’, ‘one or more’ and ‘and / or’, etc. are open-ended expressions that are both conjunctive and disjunctive in operation. For example, each of the expressions ‘at least one of A, B and C’, ‘at least one of A, B, or C’, ‘one or more of A, B, and C’, ‘one or more of A, B, or C’, and ‘A, B, and / or C may mean ‘A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B and C together’.

[0035] The terms ‘a’ or ‘an entity’ may refer to one or more of that entity. As such, the terms ‘a’ (or ‘an’), ‘one or more1and ‘at least one1can be used interchangeably herein. It is also to be noted that the terms ‘comprising’, ‘including’, and ‘having’ can be used interchangeably. Reference throughout this specification to “one embodiment,” “an embodiment,” or similar language means that a particular feature, structure, or characteristic that is descnbed in connection with the embodiment is included in at least one embodiment of the present disclosure. Thus, appearances of the phrases “in one embodiment,” “in an embodiment, and similar language throughout this specification may, but do not necessarily, all refer to the same embodiment.

[0036] The term plurality, as used herein, is defined as two or more than two. The term another, as used herein, is defined as at least a second or more. The term coupled, as used herein, is defined as connected, although not necessarily directly, and not necessarily mechanically.

[0037] In this document, the term ‘computing device’ may refer to any type of computing machine, including but not limited to, a computer, a portable computer, a laptop computer, a tablet computer, a mobile communication device, a network server, a cloud computer, etc., as well as any combination thereof. Such computing device or computing machine may include any type or combination of devices, including, but not limited to, a processor or a processing device, a memory device, a storage device, a user interface device, and / or a communication device.

[0038] The terms ‘execute’, ‘perform’, ‘compute’, ‘calculate’, ‘process’, etc. may refer to a processor of a computational device executing a software program code embodied on a non-transitory computer readable medium to achieve a result such as described after any of the terms ‘execute’, ‘perform’, ‘compute’, ‘calculate’, ‘process’, etc.

[0039] The term ‘client computing device’, or ‘client device’, ‘user device’ may refer to any type of computing device that is directly used, or operated, by a user. Such a device may include a user interface that may be used by a user directly, including means for user input and / or user output. Such a device may be communicatively coupled to another computing devices such as a network server via a communication network.

[0040] Means for user input may include a keyboard, a pointing devices such as a mouse, a microphone, a camera, a touch-sensitive plate, or display, means for user gesture control, means for haptic user control, etc. Means for user output may include a display, and / or any other means for providing visual information, a speaker, or an earphone, and / or any other means for providing audible information, means for providing tactile and / or haptic information, etc.

[0041] The term ‘mobile communication device’ may refer to devices such as a tablet, a mobile telephone, a smartphone, etc.

[0042] The term ‘network server’ or ‘server’ may refer to any type of ‘computing device’ that is communicatively coupled to a communication network and may include a cloud computer, etc.

[0043] The term ‘communication network’ or ‘network’ may refer to any type or technology for digital communication including, but not limited to, the Internet, WAN, LAN, MAN, PSDN, etc. Any of the abovementioned technologies may be wired or wireless, for example, Wireless WAN such as WiMAX, WLAN (Wi-Fi), WPAN (Bluetooth), etc. Wireless networking technology may also include PLMN, and / or any type of cellular network.

[0044] The term ‘communication network’ or ‘network’ may refer to any combination of communication technologies, and to any combination of physical networks. The term ‘communication network’ or ‘network’ may refer to any number of interconnected communication networks that may be operated by one or many network operators.

[0045] The term ‘application’ may refer to a software program running on, or executed by, one or more processors of a computing devices, and particularly by a mobile computing device such as a mobile telephone, a tablet, a smartphone, etc., as well as any other mobile or portable computing facility. The term ‘mobile application’ may refer to an application executed by a mobile computing device.

[0046] The term ‘communication’ may refer to the use of any communication network, or means of communication, by a user (person, human) to communicate content to another user. Such communication may be direct like in a telephone call, or indirect (or store and forward), such as in messaging. Messaging can be half-duplex, for example, when the message is completed, stored, forwarded to the recipient, and then consumed by the recipient in whole before responding to the sender. Messaging can be full-duplex, for example, when the message may be forwarded to the recipient before it is completed and the recipient may respond to the sender before the message ends.

[0047] The term ‘streaming content’, or ‘streaming’, may refer to data provided as a stream of data elements being sent and / or received at a predetermined repetition. For example, video may be sent and / or received at the frequency of 30 frames per second (fps), where each frame may include the same number of pixels, and each pixel may include the same number of bytes. The number of fps here (temporal resolution) is arbitrary as well as the number of pixels in a frame (spatial resolution) and number of bits in a pixel (color resolution).

[0048] The term ‘system prompt’ may refer to any prompt that is fed to a large language model (LLM) prior to a user prompt. The term ‘system prompt’ may also be known as a “model prompt”, and “technical prompt”.

[0049] Reference is now made to Fig. 1, which is a simplified block diagram of a frame selection system 10, according to one exemplary embodiment.

[0050] As seen in Fig. 1, frame selection system 10 may include a frame controller process 11. Frame controller process 11 may receive a stream 12 of video frames, which source may be a camera 13, or a storage 14, etc. Frame controller 11 may then select one or more frames 15 from the video stream 12 and provide it to one or more of a destination large language model (LLM) 16, a storage 17, a display, a communication unit, etc. (the display and the communication unit are not shown in Fig. 1).

[0051] It is appreciated that the frame selection process, for example, from the camera 13 to the destination LLM 16, is a real-time process at least in the sense that it is contained in a predefined latency. In this regard the term ‘latency’ may refer to a time period by which a selected frame 15 should be provided to the destination LLM 16 (for example) after that frame is received from the source (e.g., camera 13) within video stream 12. In this regard, the latency may be provided as a measure of time (e.g., milliseconds), or as a number of frames, or a combination thereof. The term ‘group latency’ may refer to a time period by which a selected frame 15 should be provided to the destination LLM 16 after the first frame of the group is received from the source. In this document the term latency may also refer to group latency. To select the required frames within the latency period, frame controller 11 may use one or more sub-processes 18 to process video frames 19 of the input video stream 12. Each subprocess 18 may evaluate a different aspect of the video frame 19 it may receive. Fig. 1 shows four subprocesses 18 but any number of subprocess 18 is contemplated. Any combination of subprocesses 18 may be processed in parallel, for example using a parallel processing architecture.

[0052] Therefore, the term ‘series feed’ may refer to one subprocess feeding another subprocess, while the term ‘parallel feed’ may refer to the frame controller 11 feeding two or more subprocesses with substantially the same frames. The term ‘parallel processing’ may refer to two or more subprocesses executing their respective processes at the same time whether these subprocesses are fed in series or in parallel.

[0053] The LLM configurator 23 may therefore determine how frame selection system 10 may perform regarding ‘series feed’, ‘parallel feed’ and / or ‘parallel processing’. The LLM configurator 23 may then instruct frame controller 11 via selector configuration data 26, and frame controller 11 may then manage the data-flow and processing manner between the relevant subprocesses 18.

[0054] It is appreciated that any combination of subprocesses 18 may be processed in series feed in the sense that the output of a first subprocess 18 is the input of a second subprocess 18. Any combination of subprocesses 18 may be processed in series feed and in parallel processing so that, for example, a first subprocess 18 is processing a first frame which is then fed (in series) to a second subprocess 18, while in the same time (in parallel) first subprocess 18 is processing a second frame.

[0055] Frame controller 11 determines which sub-process 18 to operate and to what extent, and communicates video frame 19 to the selected sub-process 18 according to the associated task performed by the sub-process 18. The number of frames in the group is therefore determined by the frame controller 11 according to the determined latency and the processing time associated with the particular task. Frame controller 11 may also communicate control data 20 to any of sub-processes 18 to set the means, and / or goals, and / or targets of each of sub-processes 18. Each sub-process 18 may therefore return to the frame controller 11 one or more selected frames 21 with their respective meta-data that may include a reason for the selection of the particular selected frame. It is appreciated that different sub-processes 18 may return different selected frames 21, or the same selected frame 21, or any combination thereof. The frame controller 11 may then determine which frames to communicate to the destination LLM 16, and / or the storage 17, and / or the display and / or the communication unit. The frame controller 11 may also determine the meta data associated with each frame of the selected frames 15.

[0056] Frame selection system 10 may be operated by a user, via user interface 22 and via a configurator 23. In the example of Fig. 1, the configurator 23 is a large language model (LLM Configurator) and the user’s input is a prompt 24.

[0057] It is appreciated that frame selection system 10 may include several different large language models. The LLM configurator 23 is used to configure frame selection system 10, an LLM frame selector 25, which is one of the sub-processes 18 may select frames 19, and the destination LLM 16 may receive the selected frames 15.

[0058] The user prompt 24, may include one or more characteristics of the system and the way it is expected to operate. For example, user prompt 24 may include a sentence such as “This is a camera system that tracks employee movements within the factory. It does not care about their activities but only their movement.” As another example, user prompt 24 may include “This system is analyzing a policeman body cam. It needs to map people in an incident, any dangerous or potentially dangerous tools or weapons, and rapid movements by people on the scene.”

[0059] The LLM configurator 23 may then configure the frame controller 11 by sending it selector configuration data 26. For example, The LLM configurator 23 may determine, for example, the latency and / or the sub-process(es) 18 to be operated by the frame controller 11.

[0060] In this respect, the LLM configurator 23 may have an input received from a user interface in the form of plain text (prompt 24), and an output (Selector Configuration 26) which is directed at the frame controller 11 in the form of a structured language (e g., XML, JSON, YAML, TOML, etc ). It is appreciated that the LLM configurator 23 may receive another input, prior to receiving the user prompt 24, in the form of a system prompt 27. The system prompt 27 may include any number of requirements, or restrictions, on the system configuration and modes of operation. The system prompt 27 may therefore enable LLM configurator 23 to translate the user-prompt 24 into the appropriate selector configuration data 26.

[0061] An example of a system prompt may look like this:

[0062] 1. Provided are system modules that can be used to allow for correct frame selection given a series of continuous stream of frames.

[0063] 1. 1 LLM Frame selector

[0064] 1.1.1 Explanation for how the frame selector works (The exact algorithm by name or formula)

[0065] 1.1.2 Explanation on frame selector latency and limitations in general (for example, can analyze N frames per second).

[0066] 1.2 Perceived visual selector

[0067] 1.2.1 Explanation for how the perceived visual selector works. For example, the algorithm involves utilization of a CLIP Model allowing to find text relating to objects / portions of visual data. The selector can provide the frames where relevant information provided by the system user points to the object / visual portion (such as a scenery including dangerous objects on beach sand).

[0068] 1.2.2 Explanation on selector latency and limitations in general such as can analyze N frames per second.

[0069] 1.3 Obj ect detection selector

[0070] 1.3.1 Explanation for how the selector works.

[0071] 1.3.2 Explanation on selector latency and limitations in general such as can analyze N frames per second.

[0072] 1.4 Area coverage selector

[0073] 1.4.1 Explanation for how the selector works (The exact algorithm by name or formula)

[0074] 1.4.2 Explanation on selector latency and limitations in general such as can analyze N frames per second. 2. Please provide the best combination of selectors to use in order to select the correct frames from the given stream with a maximum latency of 500 milliseconds to the stream clock (where stream clock is defined as the series of timestamps provided with streamed frames.

[0075] 2. 1 Please provide the order of selectors in the following manner - <TotalSelectors>NumberOfSelectors< / TotalSelectors>, <SelectorOrder>Area Coverage selector-> Perceived Visual Selector -> LLM Frame Selector< / SelectorOrder> denotes the delimiter and the order the stream flows through.

[0076] 3. Given a user prompt requesting an analysis of a stream with parameters provided by said user, please use the information herein in order to configure the system to the user's needs.

[0077] It is appreciated that the frame controller module 11 may receive the output of the LLM configurator 23, namely the selector configuration data 26, may interpret it and may implement it. The frame controller module 11 may therefore be responsible to implement the order of the processing by the selected modules. In this respect, each subprocess module 18, when processing frames 19, may return to the frame controller module 11, the respectively selected frames 21, and the frame controller module 11 may then forw ard the selected frames 21 (as frames 19) to the next subprocess module 18 for further processing.

[0078] It is appreciated that the frame controller module 11 may also be responsible to maintain the required latency. If the accumulative time of processing according to the order of the subprocess module 18 is larger than the required latency then the frame controller module 11 may return an error message to the LLM configurator 23, which may issue a response 28 to the user interface 22.

[0079] It is appreciated that the metadata of each frame may include, among other data, a time stamp. Such timestamp may be generated by the camera (or stored in a video file for later analysis). Each metadata may also include a cut-off lime, which may be the timestamp value plus the latency value, namely the maximum allowed delay for the particular frame. The cut-off time may be set by the frame selector 26. The metadata may also include a clarity score (and / or blurriness score) that may also be set by the frame selector 26.

[0080] It is appreciated that the metadata of each frame may also include a score as determined by any sub-process 18. For example, the metadata may include a PVS-score as determined by the perceived visual selector (PVS) 29.

[0081] Reference is now made to Fig. 2, which is a simplified flow chart of perceived visual selector 29, which may be one of the sub-processes 18 operated by the frame controller 11, according to one exemplary embodiment.

[0082] As an option, the flow chart of Fig. 2 may be viewed in the context of the previous Figures. Of course, however, Fig. 2 may be viewed in the context of any desired environment. Further, the aforementioned definitions may equally apply to the description below.

[0083] The purpose of perceived visual selector (PVS) 29 is to scan the stream of frames, detect one or more groups of similar frames, determine one or more frames that are best representative of each group, and communicate the selected frames to the frame controller. There may be several ways to achieve this purpose and the LLM configurator 23 has the capability to control PVS 29 via control data 20 to operate in a particular way. Fig. 2 shows an exemplary process of PVS 29 denoted 29A

[0084] As shown in Fig. 2, PVS 29A may start by receiving (30) control data 20 from frame controller 11 (that may have received control data 20 from LLM configurator 23 (frame controller 11 and LLM configurator 23 are not shown in Fig. 2). This control data 20 may set the way PVS 29 operates (namely, PVS 29A). PVS 29A may then receive (31) a raw group of video frames 19 from frame controller 11. This raw group may include, for example, 30 frames representing 1 second of video stream.

[0085] The latency period (as set by LLM configurator 23) may be, for example, 2 seconds. One second is consumed by the 30 frames, leaving 1 second for processing by PVS 29A. The goal here (of PVS 29A ) may be, for example, to communicate up to 5 frames to destination LLM 16 and thus, for example, saving about 1 minute of processing the 30 raw frames by destination LLM 16. PVS 29A may therefore generate visual embeddings (32) for each of the raw group of 30 frames 19. For example: fl = <5.1122334455, 6.55667788, 7.44556677>

[0086] .f3 = <1.1122334455, 2.55667788, 3.44556677> f5 = <5.1122334455, 6.55667788, 7.44556677>

[0087] ,f7 <1.1122334455, 2.55667788, 3.44556677>

[0088] .fN = altogether 30 vectors.

[0089] The terms ‘visual embedding’, ‘vector embedding’, or simply ‘embedding’ may refer to generating a data set, such as an array of numbers, for example in the form of a vector, representing visual content of an image, such as a video frame.

[0090] PVS 29A may then analyze the frames (33), or their respective vectors, for example by using cosine similarity formula that measures the cosine value between the vectors of various frames to create respective groups (34). For example:

[0091] Groupl = {fl, f5, ... },

[0092] Group2 = {fl, f7, ... },

[0093] Etc.

[0094] For example, a group of frames may include a series of frames (e.g., a video clip) which content is similar according to a required threshold that may be set by LLM configurator 23. If, for example, the scenery is changing quickly enough so that one second of video does include two similar frames then no group may be created (35). PVS 29A would then proceed to inspect images (36) and return all frames to frame controller 11 (37).

[0095] If a group is identified, it is generated (defined, collected) as a group of frames for further procesing. If (38) the assigned latency period is over, or if the data is over, namely the raw group of 30 frames are all analyzed, PVS 29A may proceed to inspect images (36) and return (37) all frames to frame controller 11. The process of inspect images (36) may select the best quality frame for each group. For example, {gl, fl}, {g2, f7}, etc. The group and its best frame(s) may now be registered (39). If time (latency) permits, and there are more video frames to process (40) PVS 29A may generate (34) more groups. If latency has elapsed or there are no more video frames to process, PVS 29A may check (41) if the control data 20 requires to determine one best group (42). Otherwise, all the best frames of all groups are returned (37) to frame controller 11. The criterion for selecting the best group may depend on the type of model used by the PVS. For example, the PVS may be trained to recognize changes in location within a confined space. The criterion may then prefer images from the confined space. Alternatively, the PVS model is trained to recognize objects, and the criterion prefers a group with a dog.

[0096] Reference is now made to Fig. 3, which is another simplified flow chart of perceived visual selector 29, which may be one of the sub-processes 18 operated by the frame controller 11, according to one exemplary embodiment.

[0097] As an option, the flow chart of Fig. 3 may be viewed in the context of the previous Figures. Of course, however, Fig. 3 may be viewed in the context of any desired environment. Further, the aforementioned definitions may equally apply to the description below.

[0098] Another purpose of the perceived visual selector (PVS) 29 is to detect frames containing one or more target objects and score the object image (PVS-score) for each object in each frame. Fig. 3 shows an exemplary process of PVS 29 denoted 29B. The perceived visual selector (PVS) 29B may provide the PVS-score of each image to the frame controller process 11. In this respect, PVS 29B may create a group of video frames that include the same visual information and select the best (most) representative frame.

[0099] Frame controller process 11 may then perform one of the following two actions, or both:

[0100] 1. Select the best frames of a group of frames according to the PVS-score and similar scores provided by the other sub-processes 18, and provide the selected frames as selected frames 15 of Fig. 1.

[0101] 2. Forw ard the frame, or group of frames, or the best frames of a group of frames (e.g., frames 19 of Fig. 1), to another sub-process 18 for further processing and scoring. To maintain real-time processing, the processing time of each frame by the perceived visual selector (PVS) 29 (or any other of sub-processes 18) is limited by the rate of frames received in its input. For example, at the rate of 25 frames per second in the input of visual selector (PVS) 29, the frame processing time of visual selector (PVS) 29 is limited to 40 millisecond per frame.

[0102] The latency is determined for the entire process of selecting frames including all sub-processes 18 that are engaged in the processing of frame selection as determined by the selector configuration data 26 of Fig. 1.

[0103] Perceived visual selector (PVS) 29B may therefore start by receiving (43) a frame 19 at a time. Perceived visual selector (PVS) 29B may then analyze (44) each frame 19 to determine the goals as set by the frame controller 11 in the received (30) control data 20. A typical goal may be to identify a particular target object in the frame and determine a PVS-score as to the ‘clarity’ of the determined target object.

[0104] For example, the PVS 29B may use an Al (artificial intelligence) classification model that is capable of recognizing the required target object(s) (e.g., an Image classifier neural network). If a target object is identified (45), PVS 29 may output the identity of the object, for example, as meta-data of the particular frame (e.g., mark object type 46).

[0105] For example, the PVS 29B may provide as an output a vector representing the identified object in ‘object-space’. For example, the PVS 29B may additionally provide as an output a PVS-score based on the probability score that the identified object is the target object. As shown in Fig. 3, PVS 29 may repeat the analysis until no more objects are detected.

[0106] Within the allowed latency window, the PVS 29B may look for significant changes and group similar frames. The frames may then go through image processing inspection to select the frame having best quality (e.g., score, clarity, probability) from each group. The last selected frame, regardless of the algorithm that made it selected, is being compared to the first group. This may prevent unnecessary frames being selected in each new window. To keep optimal results and respect the requested latency, the PVS 29B may perform selection only if a group has been already defined within a window (two changes), or when a group that is still ongoing is already X latency apart from where the new group started.

[0107] To maintain the desired latency while processing a group of frames 19 the PVS 29B may break a group into two or more sub-groups of frames from which the frame controller process 11 may select the best frames. Therefore, if the current frame is the last frame (47) in the group (or subgroup) or if latency is about to expire (48), then PVS 29 may close the current group (49) and open a new group (50).

[0108] Additionally, or optionally, PVS 29B may be required to limit the number of selected frames, for example, to N frames. In such case (51) the PVS 29B may select or indicate (e.g., in the metadata) the best N frames (52) as frames being 21 returned (53) to the frame controller process 11.

[0109] It is appreciated that any number of types of any subprocess is contemplated, and the any number of types of the same subprocess can be executed in the same time, whether is series feed, parallel feed, and / or parallel processing.

[0110] Reference is now made to Fig. 4, which is a simplified flow chart of object detector selector 54, which may be one of the sub-processes 18 operated by the frame controller 11, according to one exemplary embodiment.

[0111] As an option, the flow chart of Fig. 4 may be viewed in the context of the previous Figures. Of course, however, Fig. 4 may be viewed in the context of any desired environment. Further, the aforementioned definitions may equally apply to the description below.

[0112] The purpose of the object detector selector 54 is to recognize (identify) objects within each frame of the input video frames 19 and report the recognized object and the associated frame to the frame controller 11. Object detector selector 54 may also report the location of each object in a particular frame as well as the motion of the object between frames. Object detector selector 54 may be configured, for example, by LLM configurator 23 (via frame controller 11) with different types of parameters, such as: A. Ignore objects. For example, a list of objects that may be completely ignored for the purpose of frame selection.

[0113] B. Specified accuracy for specific objects. For example, a specific list of detected objects the system can specify its inclusion rules, the requested accuracy for movement and potentially a minimal forced frame rate for specific detected objects, unless no movement was detected.

[0114] C. Default handling of unlisted objects such as: ignore, report appear or disappear, track (e.g., motion between frames) or accuracy.

[0115] The process of object detector selector 54 may start by receiving (55) control data 20, and then receiving (56) video frames 19. If (57) control data 20 included ‘ignore objects’ data (option A above) and a listed object is detected (58) the associated frame may be removed (59).

[0116] If ignore is not enabled or if such listed object is not detected object detector selector 54 may proceed to option B above to detect an object (60) and determine accuracy, location, motion, etc. The frame (e.g., the meta data of the frame) may then be marked (61) with the required data.

[0117] The process of object detector selector 54 may then proceed to option C above (if set by control data 20) to process unlisted objects (62). Then, if the data is over, or if the latency period elapsed (63), the list of selected frames is created (64) and (all marked frames are) returned (65) to frame controller 11.

[0118] Reference is now made to Fig. 5, which is a simplified flow chart of Large language model (LLM) frame selector 25, which may be one of the sub-processes 18 operated by the frame controller 11, according to one exemplary embodiment.

[0119] As an option, the flow chart of Fig. 5 may be viewed in the context of the previous Figures. Of course, however, Fig. 5 may be viewed in the context of any desired environment. Further, the aforementioned definitions may equally apply to the description below. LLM frame selector 25 may be another example of a subprocess 18 of Fig. 1. As shown in Fig. 5, LLM frame selector 25 may start by receiving (66) control data 20, and then receiving (67) video frames 19. Control data 20 may be generated by LLM configurator 23 and provided to LLM frame selector 25 via Frame controller 11. Control data 20 for LLM configurator 23 may be a set of system prompts, followed by user prompt, followed by a group of video frames 19.

[0120] LLM configurator 23 may choose to use this subprocess 18 to select meaningful frames in more complex situations. An example of a simplified prompt that can be generated by the LLM configurator 23 is:

[0121] “The set of photos and their timestamps represent a real-time video. Please select the minimum number of photos so that every object appears at least once in its fullest form. Additionally, include all frames where a weapon or a tool appears clearly, and if such a tool does not appear clearly include the frame that is of the best quality. “

[0122] The following may be added if there were photos selected on the previous batch of photos:

[0123] “The following photos are the last three photos and their timestamps were selected from the previous batch of photos so you can consider that in the need to include the first photos.”

[0124] The LLM frame selector 25 may be fed by a window^ of N frames (photos) 19, depending on the capacity of the LLM frame selector 25 and the required latency. The latency window includes the time required to acquire N photos plus the processing time by the LLM frame selector 25. The LLM frame selector 25 may therefore execute frame selection analysis (68) and then return (69) the (selected) frames to the frame controller 11.

[0125] Covered area selector may be another example of a subprocess 18 of Fig. 1. The covered area selector is used for moving cameras only and uses available sensors such as gyro, accelerometer, camera information, or image analysis, to determine camera motion. The covered area selector may trigger frame inclusion (namely, report the frame-to-frame controller 11) based on the amount of change of the area coverage, regardless of detected objects. For example, by way of a change of the field of view (e.g., size of imaged area), or by angular change of the camera pointing. For example, the LLM configurator 23 (of Fig. 1, via frame controller 11) may set a threshold value such as 20% change of the covered area to mark the particular frame and return it to the frame controller 11.

[0126] It is appreciated that frame selection system 10 conveys images from a source of images (e.g., camera 13 or storage 14) to a destination Large Language Model 16, while selecting or removing frames from the input stream. Therefore, the number of images (photos, frames) provided to the Large Language Model 16 is much smaller than the number of images in the input stream.

[0127] The computer-implemented method for reducing the number of frames of a video stream may therefore include the actions of:

[0128] Receiving video streaming data comprising a first plurality of video frames.

[0129] Performing any of: selecting at least one frame from the first plurality of video frames to form a reduced frame set, and removing at least one frame from the first plurality of video frames to form a reduced frame set, and

[0130] Communicating the reduced frame set to the Large Language Model 16.

[0131] Further according to another exemplary embodiment there is provided a computer-implemented method for selecting frames from a video stream additionally including at least one of: performing the at least one of selecting at least one frame and removing at least one frame in real-time, performing the at least one of selecting at least one frame and removing at least one frame within a latency period, and repeating the performing the at least one of selecting at least one frame and removing at least one frame until a latency period is about to elapse.

[0132] It is appreciated that certain features, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable subcombination. Although descriptions have been provided above in conjunction with specific embodiments thereof, it is evident that many alternatives, modifications, and variations will be apparent to those skilled in the art. Accordingly, it is intended to embrace all such alternatives, modifications and variations that fall within the spirit and broad scope of the appended claims. All publications, patents and patent applications mentioned in this specification are herein incorporated in their entirety by reference into the specification, to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated herein by reference. In addition, citation, or identification of any reference in this application shall not be construed as an admission that such reference is available as prior art.

Claims

CLAIMS:What is claimed is:

1. A computer-implemented method for configuring a system, the method comprising: obtaining a large language model; providing the large language model with a system-prompt comprising configuration options of the system; obtaining from a user a prompt describing the purpose of the system; obtaining, from the large language model, an output, the output comprising configuration data for the system, the configuration data complying with the userprompt; communicating the configuration data to the system.

2. The method according to claim 1, additionally comprising: wherein the system comprises at least two processing modules; wherein the system-prompt comprises at least one requirement for each of the at least two processing modules; wherein the user-prompt comprises a characteristic; and wherein the configuration data comprises at least one of the at least two modules which requirement complies with the characteristic.

3. The method according to claim 1, additionally comprising: wherein the system-prompt comprises at least one requirement in the form of a time constraint; wherein the user-prompt comprises a characteristic; and wherein the configuration data comprises at least one time constraint that complies with the characteristic.

4. A computer-implemented method for selecting frames from a video stream, the method comprising:receiving a video streaming data comprising a first plurality of video frames; selecting at least one frame from the first plurality of video frames; and communicating the at least one selected video frame to a destination; wherein the communicating the selected video frames is performed within a predetermined latency since said receiving the video streaming data.

5. The method according to claim 4, additionally comprising: wherein the receiving a video streaming data comprises receiving a group of frames, and wherein the predetermined latency is measured from a first frame of the group of frames.

6. A computer-implemented method for selecting frames from a video stream, the method comprising: receiving a video streaming data comprising a first plurality of video frames; selecting at least one frame from the first plurality of video frames; and communicating the at least one selected video frame to a destination; wherein the selecting at least one frame comprises: generating visual embeddings for each frame; comparing vectors to determine frames of a group; and selecting at least one frame representative of the group.

7. The method according to claim 6, additionally comprising: repeating the selecting at least one frame until a latency period is about to elapse.

8. A computer-implemented method for selecting frames from a video stream, the method comprising: receiving a video streaming data comprising a first plurality of video frames; selecting at least one frame from the first plurality of video frames; and communicating the at least one selected video frame to a destination; wherein the selecting at least one frame comprises: detecting at least one object within each frame;comparing objects detected in different frames to determine frames of a group; and selecting at least one frame representative of the group.

9. The method according to claim 8, additionally comprising: repeating the selecting at least one frame until a latency period is about to elapse.

10. A computer-implemented method for selecting frames from a video stream, the method comprising: receiving a definition of at least one undesired object; receiving a video streaming data comprising a first plurality of video frames; selecting at least one frame from the first plurality of video frames; and communicating the at least one selected video frame to a destination; wherein the selecting at least one frame comprises: detecting at least one undesired object within a frame; and deselecting at least one frame including said undesired object.

11. The method according to claim 10, additionally comprising: repeating the deselecting at least one frame until a latency period is about to elapse.

12. A computer-implemented method for selecting frames from a video stream, the method comprising: receiving a video streaming data comprising a first plurality of video frames; selecting at least one frame from the first plurality of video frames; and communicating the at least one selected video frame to a destination; wherein the selecting at least one frame comprises: detecting at least one moving object within at least two frames; measuring the distance passed by the object between the at least two frames; and if the distance is larger than a predetermined value selecting at least one frame including said moving object.

13. The method according to claim 12, additionally comprising: repeating the selecting at least one frame until a latency period is about to elapse.

14. A computer-implemented method for selecting frames from a video stream, the method comprising: receiving a video streaming data comprising a first plurality of video frames; selecting at least one frame from the first plurality of video frames; and communicating the at least one selected video frame to a destination; wherein the selecting at least one frame comprises: detecting a change of a camera field of view between at least two frames; measuring the change between the at least two frames; and if the change is larger than a predetermined value, selecting at least one frame including said change.

15. The method according to claim 12, additionally comprising: repeating the selecting at least one frame until a latency period is about to elapse.

16. A computer-implemented method for selecting frames from a video stream, the method comprising: receiving a video streaming data comprising a first plurality of video frames; performing at least one of: selecting at least one frame from the first plurality of video frames to form a reduced frame set; and removing at least one frame from the first plurality' of video frames to form a reduced frame set; and communicating the reduced frame set to a Large Language Model.

17. The method according to claim 16, additionally comprising at least one of: performing the at least one of selecting and removing in real-time; performing the at least one of selecting and removing within a latency period; and repeating the performing the at least one of selecting and removing until a latency period is about to elapse.

Citation Information

Patent Citations

  • Creating a reliable configuration specification using an llm

    EP4546230A1

  • Artificial intelligence for intent-based networking

    US11968088B1

  • Systems for controllable summarization of content

    US12008332B1

  • Optimization of a media processing system based on latency performance

    US20190026105A1

  • Systems and methods for dynamic large language model prompt generation

    US20240330579A1