Method to extract video context from external description
By using natural language input from industrial domain experts to calculate domain context and automatically label images in industrial videos, the method addresses the scalability issues of existing video context extraction processes, reducing deployment time and costs for AI solutions in new industrial environments.
Patent Information
- Application Number
- JP2024178680
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-08
- Filing Date
- 2024-10-11
- Publication Date
- 2025-06-19
- Estimated Expiration
- 2044-10-11
AI Technical Summary
The existing process of extracting video context for industrial applications is not scalable due to manual interaction between domain experts and data engineers, leading to inefficiencies and increased costs when deploying AI solutions in new industrial environments.
A method that utilizes natural language input from industrial domain experts to calculate domain context, and automatically labels images in videos observing industrial activities, eliminating the need for manual data engineering.
This approach reduces the time and cost associated with expanding AI solutions in new industrial environments by automating the image labeling process, thereby enhancing scalability and efficiency.
Smart Images

Figure 2025092415000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure is generally directed to video systems, and more particularly, to systems and methods for extracting video context from external descriptions.
Background Art
[0002] In many industrial applications, it is necessary to observe activities including the interaction between workers and industrial assets or tools and draw inferences. For example, it is necessary to observe a worker performing a repetitive task and calculate the time it takes for the worker to complete each cycle of the task and the variation in such times over multiple cycles of the same task. This is called cycle time calculation.
[0003] In another example, there is a need to observe whether a worker is maintaining or violating the standard operating procedure (SOP) of an industrial activity. For example, the SOP can instruct a precisely defined procedure for assembling parts, if there is a violation and it is necessary to report it to the line supervisor or appropriate authorities.
[0004] As a more general example, by observing the industrial environment and deriving certain conclusions about the occurrence of an event from the combination of the states of assets and workers. Conventionally, such a process of observation and inference has been carried out by the line supervisors themselves. However, since supervisors cannot be present at all workplaces simultaneously, it is inefficient. Also, human errors and biases are likely to occur. Against this backdrop, with the emergence of Industrial Internet of Things (IIoT), the number of cases where sensors (such as video cameras) for recording the industrial environment (such as factories) are installed and artificial intelligence (AI) is used to infer what is happening in the factory from the data for observation is increasing. The inference can include some conclusion about the events occurring in the factory based on the combination of the states of assets and workers. The inference can also infer the actions of workers, and the key performance indicators (KPIs) related to the tasks of workers, such as compliance or violation of SOPs and cycle times.
[0005] A common approach using AI for image data is to use image classification. Different classifications are, for example, "cycle start" and "cycle stop" for cycle time calculation, "normal operation", "prohibited operation", "unknown operation" for SOP compliance, or "event A" start for generalized event-based analysis. To train such a classification machine learning (ML) model, it is necessary to prepare training data including labeled images. To label an image, it is necessary to understand the context of that image. In many cases, this is a specialized task that can only be performed by experts in the industrial field.
[0006] FIG. 1 is a diagram showing an example of a related art system for image labeling for training a machine learning (ML) model. As shown in FIG. 1, data engineer 401 needs to conduct a joint session with industry domain expert 101. Expert 101 looks at the images from input video 201 and describes what industrial activities are taking place in each image. Labels are attached to the images of the input video. Based on the context information of the description by expert 101, data engineer 401 labels the images of the input video with context information and creates context-attached image label 301. SUMMARY OF THE INVENTION PROBLEMS TO BE SOLVED BY THE INVENTION
[0007] As expected, the above-described process is not scalable due to the manual interaction between the industry domain expert and the data engineer. For example, since industrial activities are highly specialized, every time a video analysis solution is introduced to a new industrial assembly line or a new factory, it is necessary to conduct manual interaction work between the industry domain expert and the data engineer. This is because even if there is a common problem of SOP compliance, what SOP means varies depending on the assembly line or factory, and thus it is necessary to train data accordingly. Although there is a publicly available database of images, it should be borne in mind that a large-scale labeled image database containing all data variations is not available for industrial scenarios. The main reason why a large-scale labeled image database is not available for industrial scenarios is the confidential information contained in the images. Therefore, it is necessary to conduct training for each new industrial scenario.
[0008] The implementation of the related art is troubled by the increased time for expanding AI solutions in new industrial environments such as new assembly lines and new factories, and the associated costs for hiring data engineers. MEANS FOR SOLVING THE PROBLEM
[0009] In an exemplary implementation, a domain context for industrial activities calculated based on natural language input from experts in the industrial domain is utilized, and further, a method is included for automatically labeling images in a video that observes subsequent instances of the industrial activities without requiring a data engineer.
[0010] One aspect of a method for extracting video context from an external description to solve the above problems is in a method for extracting video context from an external description, which receives an input video for processing, executes a zero-shot image labeler that generates a plurality of labels corresponding to a plurality of images of the input video, calculates an image embedding vector for the plurality of labels, and refers to a context database with the image embedding vector to determine a context label from an inspected event graph, and generates a context label for replacing each of the labels corresponding to each of the images based on the context labels determined temporally for the current image and previous images. The nodes of the event graph indicate events that can occur during the duration of the video. Replace each of the plurality of labels with the generated context label.
[0011] Aspects of the present disclosure include, in a method for extracting video context from external descriptions, means for receiving an input video for processing, means for executing a zero-shot image labeler that generates a plurality of labels corresponding to a plurality of images of the input video, means for calculating an image embedding vector for the plurality of labels, and means for generating a context label for replacing each of the labels corresponding to each of the images based on context labels determined temporally for the current and previous images by referring to a context database with the image embedding vectors in order to determine a context label from an inspected event graph, wherein the nodes of the event graph indicate events that can occur during the duration of the video, and means for replacing each of the plurality of labels with the generated context labels.
[0012] Aspects of the present disclosure can include a computer program that causes a processor to execute instructions including receiving a video for processing, executing an image labeler on the video to generate a plurality of labels corresponding to the images of the video, and calculating an image embedding vector for each of the labels of the images. By the program, the processor generates a context label for replacing each of the labels corresponding to each of the images based on context labels determined temporally for the current and previous images by referring to a context database with the image embedding vectors in order to determine a context label from an inspected event graph, wherein the nodes of the event graph indicate events that can occur during the duration of the video, and replaces each of the plurality of labels with the generated context label. The computer program and instructions are stored on a non-transitory computer-readable medium and can be executed by one or more processors.
[0013] Aspects of the present disclosure can include a processor configured to execute an image labeler on a video to generate a plurality of labels corresponding to the images of the video for receiving the video for processing, and calculate an image embedding vector for each of the labels of the images. The processor generates a context label for replacing each of the labels corresponding to each of the images based on the context labels determined for the current and previous images in time by referring to a context database with the image embedding vectors to determine the context labels from the inspected event graph, where the nodes of the event graph indicate events that can occur during the duration of the video, and replaces each of the plurality of labels with the generated context labels.
[0014] Aspects of the present disclosure can be a system having a context database and a processor, which can include a processor configured to execute an image labeler on a video to generate a plurality of labels corresponding to the images of the video for receiving the video for processing, and calculate an image embedding vector for each of the labels of the images. The system generates a context label for replacing each of the labels corresponding to each of the images based on the context labels determined for the current and previous images in time by referring to a context database with the image embedding vectors to determine the context labels from the inspected event graph, where the nodes of the event graph indicate events that can occur during the duration of the video, and replaces each of the plurality of labels with the generated context labels.
Advantages of the Invention
[0015] According to the present invention, there is provided a method for extracting video context from external descriptions, which can suppress an increase in time for expanding AI solutions in new industrial environments such as new assembly lines and new factories, and reduce the associated costs for hiring data engineers.
Brief Description of the Drawings
[0016]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
[0017] The following embodiments will be described in detail. The figures and examples are for explaining the embodiments in detail. The reference numerals between the figures and the descriptions of redundant elements are omitted for clarity. The terms used throughout this specification are provided as examples and are not intended to be limiting. For example, the use of the term "automatic" may include fully automatic or semi-automatic embodiments involving user or administrator control over specific aspects of the embodiment, depending on the desired embodiment of those skilled in the art practicing the embodiments of the present invention. The selection can be implemented by the user through a user interface or other input means, or can be implemented through a desired algorithm. The exemplary implementations described herein can be utilized either alone or in combination, and the functions of the exemplary implementations can be implemented in any manner according to the desired implementation.
[0018] FIG. 2 is a diagram showing an exemplary learning mode and operation mode executed by the processor 1210 (FIG. 12).
[0019] In the learning mode 100, the raw-form data context 101a in natural language, which is explained by experts (domain experts) in each industrial domain, is used. The forms that the raw-form data context 101a can take are various. For example, the raw-form data context 101a can include text input by experts, or can also include text records of what experts say. Instead of human experts, it may also be a manual in which important information is extracted via computer algorithms to construct the raw-form data (raw-form data context) 101a.
[0020] The input of data context 101a by an expert in the industrial domain is output to and supplied to the context description module 102. The context description module 102 outputs an event graph 103 that documents the graphical relationships between various events described by the input context 101a. This can only be done manually for one sample of the input context 101a. For multiple samples (e.g., inputs from multiple domain experts 101), a generalized approach can be followed as known in the related art. One way to generalize is to consider not only tasks but also events. In the case of tasks, the time taken to execute each of these steps can also be potentially extracted from the input of the industrial domain expert 101.
[0021] FIG. 3 is a diagram showing an example of an inspection event graph of the event graph 103 (see FIG. 2) executed by the processor 1210 (FIG. 12). This is an example where the event sequence describes a typical workflow related to the inspection of parts at inspection points. In the example of FIG. 3, the nodes of the event graph can include, but are not limited to, both the operator and the part not being present at inspection point 103 - a, the part arriving at the inspection point but the operator not being present 103 - b, the part arriving at the inspection point with the operator present but not performing an inspection 103 - c, the part arriving at the inspection point with the operator present and performing an inspection 103 - d, and the part leaving the inspection point with the operator present and not performing an inspection 103 - e.
[0022] The text embedding module 104 captures the text within the event graph 103, generates high-dimensional vectors, and embeds them into the event graph. This is performed by taking a sentence as input as a character string and outputting a real-valued N-dimensional vector such that sentences with similar meanings have embedding vectors that are close to each other in a Euclidean sense. The N-dimensional vector can be output by various embedding solutions available in the related art. For example, the sentences "I like Lionel Messi" and "My favorite sport is soccer" generate embedding vectors "u" and "v" such that (u - v) 2 results in a small value. "u" and "v" represent the embedding vectors of their respective sentences.
[0023] The event graph 105 with embeddings generated by the text embedding module 104 is stored in the context database 106. The context database 106 includes two tables. The context database 106 manages the association between the embedding vectors and the context information of each node of the inspected event graph, which will be described later, and can be implemented by any type of hardware database according to the desired implementation, such as a storage system, a cloud-based system, etc. (but not limited to these).
[0024] Figure 4 is a diagram showing an example of the context information 107 (node table) of the inspection event graph according to an embodiment. The first table in Figure 4 is the node table for the example of the inspection event graph in Figure 3. This table can include the following information.
[0025] An ID field that counts the number of entries.
[0026] A NodeID field that uniquely identifies the node of the graph containing the event information. For example, the nodes of the event graph represent events that can occur during the duration of a video.
[0027] The Description field contains text that describes the content of the event, generated from the input of domain experts. The content of a typical workflow related to the inspection of parts at an inspection point in the event sequence of Figure 3 is described.
[0028] The Embedding field is an N-dimensional real vector that indicates the embedding value corresponding to the text of the Description field. The value of N depends on the specific embedding method. For example, N = 512 is the default value of the related art embedding CLIP. This Embedding field corresponds to, for example, the embedding vector generated by the text embedding module 104.
[0029] The Estimated Duration field indicates the estimated duration of the event. This can be calculated from the input given by the domain expert if the estimated duration is specified. If not specified, the field is denoted as N / A or Not Available.
[0030] The Level field indicates the depth of the graph in which the node is located. That is, it indicates the hierarchy of each node.
[0031] Figure 5 is a diagram showing an example of the context information 107 (edge table) of the inspection event graph according to the embodiment. The second table is an edge table as shown in Figure 5 for the example of the inspection event graph of Figure 3. This table can include the following information.
[0032] An ID field that counts the number of entries.
[0033] An EdgeID field that uniquely names the edges of the graph connecting the nodes containing event information. The EdgeID field corresponds to, for example, the inspection points of the workflow shown in Figure 3.
[0034] Starting Node field that specifies the starting node of a given edge.
[0035] End Node field that indicates the end node of the specified edge.
[0036] If any, Conditions field that captures special conditions that are additionally required to capture dependencies between events. In the above example, there are no special conditions.
[0037] FIG. 6 is a diagram showing an example of an inspection event graph 103 according to an embodiment.
[0038] For more generalized events, consider a more detailed inspection event graph as shown in FIG. 6 to consider how the solution unfolds. Here, the events capture only the detailed tasks that an operator must perform. This is referred to herein as an inspection task graph, and the "task" is used interchangeably with "event" only for this type of situation. Thus, FIG. 6 can be referred to as an inspection task graph.
[0039] In the first step 103-f, the operator stands in front of the belt of the assembly line and waits for the part to arrive and stop under the magnifying glass for inspection. Next, the operator refers to the work instruction and determines the specifications to be checked for both the right side 103-h and the left side 103-g of the part. After checking the specifications on one side, the operator proceeds to inspect the sides of the part (103-i, 103-j). It doesn't matter which part the operator inspects first. Only after both sides have been inspected can the operator proceed to the next step. This step is to sign in the manufacturing execution system (MES) kiosk indicating that the work is completed (103-k). After the work is completed, the operator takes one of two actions depending on whether the entire part is judged to be a defective part: either discard the part in the defective part bin 103-l or press the button 103-m to advance the part.
[0040] Furthermore, the event list also describes tasks that an operator must not perform, such as picking up parts (step 103-x).
[0041] Note that steps that must not be taken, such as step 103-x, are depicted in a novel way as compared to the related art.
[0042] FIG. 7 is a diagram showing context information 107 (node table) corresponding to the inspection task graph shown in FIG. 6. Note that some events (tasks) are at the same level indicating tasks that can be executed in any order. Also, prohibited events (tasks, procedures) are assigned a special level number (for example, the node indicating the prohibited action of picking up a part has a level of "-1").
[0043] FIG. 8 is a diagram showing context information 107 (edge table) of the inspection task graph corresponding to the inspection task graph shown in FIG. 6. As described above, the Condition field currently captures special conditions that are additionally required to capture dependencies between tasks such as AND conditions and NOT conditions.
[0044] In the operation mode 200 (FIG. 2) of the computer device 1205 (FIG. 12), the input video 201, which is a video observing the same activities as those provided by the expert in the industrial domain in the learning mode 100, is automatically labeled.
[0045] When the zero-shot image labeler 202 receives the input video 201, a text-based image label 203 is generated. The zero-shot image labeler 202 executes the generation of a plurality of labels corresponding to a plurality of images of the input video 201. These labels are based on a model (hence the name zero-shot) pre-trained on a large corpus of publicly available datasets. Therefore, the image label 203 is somewhat correlated with the activities occurring in the input video 201, but cannot fully capture the domain context. This is because although the training data of the zero-shot image labeler 202 is publicly available, it does not contain domain-specific data.
[0046] The same text embedding module 204 as the text embedding module 104 receives a plurality of image labels 203, calculates an embedding vector for each image label, and outputs an embedded image label 205 with the embedding vector attached. Then, the embedded image label 205 is output to the context generation module 206.
[0047] Figures 9 and 10 illustrate the process of determining context labels from the inspected event graph. The determination of context labels generates context labels based on the image embedding vectors calculated by the text embedding module 204, with reference to the context database 106, for the current image and the context labels determined temporally for previous images. Replace each label corresponding to each of the images with the generated context label. The nodes of the event graph represent events that can occur during the duration of the video. Replace each of the plurality of labels with the generated context label.
[0048] Figure 9 is a diagram showing the processing of the context generation module 206 executed by the processor 1210 (Figure 12). The context generation module 206 includes the following steps.
[0049] In step 206-1, the context generation module 206 calculates a task embedding list R of events prohibited from the event graph 105. These include examples such as the events at level = -1 in FIG. 7.
[0050] In step 206-2, the context generation module 206 inputs an image embedding IK (embedded image label 205). Here, 0 ≤ K ≤ M - 1, and M is the total number of images in the input video.
[0051] In step 206-3, the context generation module 206 inputs an event embedding tn. Here, 0 ≤ n ≤ T - 1, and T is the total number of nodes in the task graph having valid actions such that tn does not belong to the list R from step 206-1.
[0052] In step 206-4, the context generation module 206 checks whether the image embedding IK is close to the entry where the list R is set with high reliability. To do this, the context generation module 206 calculates the following.
[0053]
Equation
[0054] Here, D(x, y) calculates the distance between two vectors x and y such as the Euclidean distance. If this condition is true, a warning is issued. To understand the meaning of this condition, consider the case where the overall application is for SOP compliance. This condition means that an SOP violation has occurred.
[0055] In step 206-5, the context generation module 206 determines the event embedding t K-1 closest to the image embedding IK considering both the distance in the embedding space and the calculated tasks for the previous image embedding I nCalculate this. To execute this, calculate the following.
[0056]
Number
[0057] Here, d(a, b) calculates the distance between nodes a and b of the event graph. Therefore, this term attempts to set n at time K K to the node of the task graph that is the next node determined at the previous time instant. That is, the node at the previous time is n K-1 and the next node of the es graph is n (K-1 +1. This is because the task graph is constructed considering the order of tasks. The parameter α K indicates how much weight should be given to this term. This is determined as follows as a function of the estimated duration value from the node table.
[0058] Find n K-1 ≠n K and set α K = α0. Here, α0 is a certain initial number. n K+i =n K sets α K+i =α0g(i). Here, g(i) is an increasing function of i, and the growth rate is directly proportional to the estimated duration value. This idea captures the intuition that when a new task is started and enough time has passed until its estimated duration value, the next task is expected to start.
[0059] In this way, in order to determine the context label from the inspected event graph, by referring to the context database with the image embedding vector, based on the context labels determined temporally for the current image and the previous image, generate context labels for replacing each of the labels corresponding to each of the images, and replace each of the plurality of labels with the generated context labels.
[0060] FIG. 10 shows an example of assigning a value to the weighting parameter α K executed by the processor 1210 (FIG. 12) according to the example described in
[0067] . FIG. 10 shows an example of event selection based on graphical considerations.
[0061] In step 207-1, event selection determines that a new task has started at time K.
[0062] In step 207-2, event selection initializes a K to a low value a0.
[0063] In step 207-3, event selection reads the estimated duration, which is the estimated time of the new task T.
[0064] In step 207-4, for each subsequent time instance K+i, event selection K increases a K as a = a0exp (iw / T) (where w is a constant). In this case, g(i)=exp (iw / T) In step 206-6, the content generation module 206 replaces the label of image k with the task description of task n calculated in step 206-5, and generates an image label with context 301.
[0065] As an example, consider the event graph of FIG. 3 to see how the solution functions for the situation shown in FIG. 11. FIG. 11 is an illustrative diagram of an available solution example according to the embodiment. As shown in FIG. 11, for the three image frames I1, I2, I3, the zero-shot image classifier generates a description that, although not inaccurate, cannot capture the domain context mentioned in the event graph of FIG. 3. This can be done as the output of the context generation module by appropriately selecting the following parameters. FIG. 11 also shows α K=0 is indicated by a dashed frame. In this case, the determination wrongly predicts that the operator is still performing the inspection without considering the graphical relationship between events. With an appropriate α K incorporated, the solution correctly predicts that the next task should have started.
[0066] Through the examples described in this specification, the time to deploy an AI solution that requires image data labeled in a domain context can be shortened. Additionally, the exemplary implementation reduces costs as it requires fewer resources for manual data engineering.
[0067] FIG. 12 shows an exemplary computing environment having an exemplary computer device suitable for use in some exemplary implementations. The computer device 1205 in the computing environment 1200 can include one or more processing units, cores, or processors 1210, memory 1215 (e.g., RAM, ROM, and / or the like), internal storage 1220 (e.g., magnetic, optical, solid-state storage, and / or organic), and / or an IO interface 1225, any of which can be coupled on a communication mechanism or bus 1230 for communicating information or can be embedded within the computer device 1205. The IO interface 1225 can also be configured to receive images from a camera or provide images to a projector or display, depending on the desired implementation.
[0068] The computer device 1205 can be communicatively coupled to an input / user interface 1235 and an output device / interface 1240. Either or both of the input / user interface 1235 and the output device / interface 1240 can be a wired or wireless interface and can be removable. The input / user interface 1235 can include any physical or virtual device, component, sensor, or interface (e.g., buttons, touch screen interface, keyboard, pointing / cursor control, microphone, camera, braille, motion sensor, accelerometer, optical reader, and / or the like) that can be used to provide input. The output device / interface 1240 can include a display, television, monitor, printer, speaker, braille, etc. In some exemplary implementations, the input / user interface 1235 and the output device / interface 1240 can be embedded in or physically coupled to the computer device 1205. In other exemplary implementations, other computer devices can function as or provide the functionality of the input / user interface 1235 and the output device / interface 1240 of the computer device 1205.
[0069] Examples of the computer device 1205 can include, but are not limited to, highly mobile devices (e.g., smartphones, devices mounted on vehicles and other machines, devices carried by humans and animals, etc.), mobile devices (e.g., tablets, notebooks, laptops, personal computers, portable televisions, radios, etc.), and devices not designed for mobility (e.g., desktop computers, other computers, information kiosks, televisions, radios, etc. having one or more processors embedded in and / or coupled thereto).
[0070] Computer device 1205 can be communicatively coupled to external storage 1245 and network 1250 (e.g., via IO interface 1225) to communicate with any number of network-connected components, devices, and systems, including one or more computer devices of the same or different configurations. Computer device 1205 or any connected computer device can function as, provide services as, or be referred to as a server, client, synserver, general machine, special-purpose machine, or other label.
[0071] IO interface 1225 can include wired and / or wireless interfaces that use any communication or IO protocol or standard (e.g., Ethernet, 802.11x, Universal System Bus, WiMAX, modem, cellular network protocol, etc.) for communicating information between at least all connected components, devices, and networks within computing environment 1200, but is not limited thereto. Network 1250 can be any network or combination of networks (e.g., the Internet, local area network, wide area network, telephone network, cellular network, satellite network, etc.).
[0072] Computer device 1205 can use and / or communicate with computer-usable media or computer-readable media, including transient media and non-transient media. Transient media includes transmission media (e.g., metal cables, optical fibers), signals, carrier waves, etc. Non-transient media includes magnetic media (disks, tapes, etc.), optical media (CD ROM, digital video disk, Blu-ray disk, etc.), solid media (RAM, ROM, flash memory, solid state storage, etc.), and other non-volatile storage or memory.
[0073] Computer device 1205 can be used to implement techniques, methods, applications, processes, or computer-executable instructions in some exemplary computing environments. The computer-executable instructions can be obtained from a transient medium, stored in a non-transient medium, and can be obtained from the non-transient medium. The executable instructions can be derived from one or more of any programming language, scripting language, and machine language (e.g., C, C++, C#, Java®, Visual Basic, Python, Perl, JavaScript®, etc.).
[0074] Processor(s) 1210 can execute under any operating system (OS) (not shown) in a native or virtual environment. One or more applications can be deployed that include logical unit 1260, application programming interface (API) unit 1265, input unit 1270, output unit 1275, and an inter-unit communication mechanism 1295 for the different units to communicate with each other, with the OS, and with other applications (not shown). The units and elements described can vary in design, function, configuration, or implementation and are not limited to the description provided. Processor(s) 1210 can be in the form of a hardware processor such as a central processing unit (CPU) or in the form of a combination of hardware units and software units.
[0075] In some exemplary implementations, when information or execution instructions are received by the API unit 1265, they can be transmitted to one or more other units (e.g., the logic unit 1260, the input unit 1270, the output unit 1275). In some embodiments, the logic unit 1260 can be configured to control the flow of information between units and direct the services provided by the API unit 1265, the input unit 1270, and the output unit 1275 in some of the exemplary implementations described above. For example, the flow of one or more processes or implementations can be controlled by the logic unit 1260 alone or in cooperation with the API unit 1265. The input unit 1270 can be configured to obtain inputs for the calculations described in the exemplary embodiments, and the output unit 1275 can be configured to provide outputs based on the calculations described in the exemplary embodiments.
[0076] The processor(s) 1210 is configured to execute a method or instructions including receiving the processing video 201, executing an image labeler 202 on the video to generate a plurality of image labels 203 corresponding to the images of the video, and calculating an image embedding vector for each of the image labels. Further, the processor 1210 generates a context label 206 to replace each of the labels corresponding to each of the images based on the context labels determined for the current and previous images in time by referring to a context database with the image embedding vectors to determine context labels from the inspected event graph, where the nodes of the event graph indicate events that can occur during the duration of the video, and as shown in FIG. 2, each of the plurality of labels is replaced with the generated context label 301. Also, as described herein with respect to FIG. 11, each of the plurality of labels is replaced with the generated context label 301.
[0077] Depending on the desired implementation, the context database manages the association between the event embedding vector and the context information of each node of the inspected event graph, as shown in FIG. 5, via the event graph having the embedding 105.
[0078] The processor(s) 1210 is configured to execute the methods or instructions described herein. Here, generating a context label that replaces each of the labels corresponding to each of the images further includes determining a node from the inspected event graph associated with the event embedding vector in the context database having the minimum distance from the image embedding vector, and generating a context label from the context information associated with the node, as described with respect to FIG. 9.
[0079] The processor(s) 1210 is configured to execute the methods or instructions described herein. Here, each node is associated with an estimated duration in the context database. Here, the calculated distance between the image embedding vector and the event embedding vector is weighted based on the duration between the images of the video within the time associated with the same context label compared to the estimated duration, as described with respect to FIG. 10, between the current and previous images.
[0080] Depending on the desired implementation, the context database can manage prohibited procedures within the inspected event graph. In such an exemplary implementation, the processor(s) 1210 can be configured to execute the methods or instructions as described above, and further include raising an alert regarding non - compliance with the standard operating procedure for a generated context label that includes one of the prohibited procedures, as described with respect to FIG. 9.
[0081] The processor(s) 1210 is configured to execute the methods or instructions described herein, and further includes identifying a task from the generated context label and determining a cycle time of the identified task from the length of the video over the generated context label associated with the identified task as described with respect to FIG. 10. In an exemplary implementation, the cycle time can be determined by aggregating the estimated durations between each task.
[0082] The processor(s) 1210 is configured to execute the methods or instructions as described above, and further includes providing an indication that an abnormal event has occurred for one or more of the generated context labels indicating an abnormal event. For example, if a warning needs to be issued considering an abnormal event associated with a node in the inspected event graph, the event can be detected through the exemplary implementation described herein and controlled to indicate that an abnormal event has occurred.
[0083] The processor(s) 1210 can be configured to execute the methods or instructions described herein, and further includes storing the generated context label with the video and indexing the video with a search engine. In an exemplary implementation, since each of the generated context labels is associated with a time within the video, the context label can be provided to the search engine to index the video based on time and / or label according to the desired implementation, along with the corresponding timestamp and duration.
[0084] Some portions of the detailed description are presented from the perspective of symbolic representations of algorithms and operations within a computer. These algorithmic descriptions and symbolic representations are means used by those skilled in the data processing arts to convey the essence of their technological innovation to others skilled in the art. An algorithm is a series of defined steps that lead to a desired final state or result. In an embodiment, the steps executed require physical manipulation of physical quantities to achieve a visible result.
[0085] Unless otherwise specified, as will be apparent from the discussion, throughout this specification, discussions using terms such as "processing," "computing," "calculating," "determining," "displaying," etc., may include the operations and processes of a computer system or other information processing device that manipulate and transform data represented as physical (electronic) quantities within the registers and memories of the computer system into other data similarly represented as physical quantities within the memories or registers of the computer system or other information storage, transmission, or display devices.
[0086] The exemplary embodiments also relate to an apparatus for performing the operations herein. This apparatus may be specially configured for the required purposes or may include one or more general-purpose computers selectively activated or reconfigured by one or more computer programs. Such computer programs can be stored on a computer-readable medium such as a computer-readable storage medium or a computer-readable signal medium. The computer-readable storage medium can include, but is not limited to, tangible media such as optical disks, magnetic disks, read-only memory, random access memory, solid-state devices and drives, and can include any other type of tangible or non-transitory media suitable for storing electronic information. The computer-readable signal medium can include media such as a carrier wave. The algorithms and displays presented herein are not inherently related to a particular computer or other device. The computer program can include a pure software implementation including instructions to perform the operations of the desired implementation.
[0087] Various general-purpose systems may be used with the programs and modules according to the examples herein, or it may prove convenient to construct more specialized apparatus for performing the desired method steps. Further, the examples are not described with reference to a particular programming language. It will be understood that various programming languages may be used to implement the teachings of the examples described herein. The instructions of the programming language may be executed by one or more processing devices, such as a central processing unit (CPU), a processor, or a controller.
[0088] As is known in the art, the above-described operations can be performed by hardware, software, or some combination of software and hardware. Various aspects of the exemplary implementations may be implemented using circuits and logic devices (hardware), while other aspects may be implemented using instructions stored on a machine-readable medium (software) that, when executed by a processor, cause the processor to execute a method for implementing the implementations of the present application. Further, some exemplary implementations of the present application may be executed by hardware only, while other exemplary implementations may be executed by software only. Further, the various functions described can be performed by a single unit or can span multiple components in any number of ways. When executed by software, the method can be executed by a processor, such as a general-purpose computer, based on instructions stored on a computer-readable medium. Optionally, the instructions can be stored on the medium in a compressed and / or encrypted format.
[0089] Furthermore, other embodiments of the present application will be apparent to those skilled in the art from a consideration of the specification and practice of the teachings of the present application. The various aspects and / or components of the described exemplary embodiments can be used alone or in any combination. The specification and exemplary embodiments are intended to be considered as exemplary only, and the true scope and spirit of the present application are indicated by the following claims.
Description of Reference Numerals
[0090] 101 Expert in the industrial domain 101a Raw form data context 102 Context description module 103 Event graph 104 Text embedding module 105 Event graph with embeddings 106 Context database 107 Context information 201 Input video 202 Zero-shot Image Labeler 203 Image Label 204 Text Embedding Module 205 Embedded Image Label 206 Context Generation Module 301 Image Label with Context
Claims
1. 1. A method for extracting video context from an external description, comprising: Receives an input video for processing; running a zero-shot image labeller on the input video to generate a plurality of labels corresponding to a plurality of images of the input video; Calculate an image embedding vector for the plurality of labels; generating context labels for replacing each of the labels corresponding to each of the plurality of images based on the temporally determined context labels for the current image and the previous image by referencing a context database with the image embedding vectors to determine context labels from the inspected event graph, the nodes of the event graph indicating possible events occurring during a duration of the video; replacing each of the plurality of labels with the generated context label; A method for extracting video context from external descriptions.
2. 2. The method of extracting video context from an external description of claim 1, further comprising: the context database manages associations between the image embedding vectors and context information for each node of the inspected event graph; A method for extracting video context from external descriptions.
3. 2. The method of extracting video context from an external description of claim 1, further comprising: Generating a context label to replace each of the labels corresponding to each of the plurality of images includes: determining a node from the examined event graph associated with an event embedding vector in the context database that has a minimum distance from the image embedding vector; Generate context labels from context information associated with the nodes; A method for extracting video context from external descriptions.
4. The method of extracting video context from an external description of claim 3, further comprising: Each of the nodes is associated with an estimated duration of the context database; The calculated distance between the image embedding vector and the event embedding vector is weighted based on the duration of time between the images of the video in time and the current and previous images associated with the same context label compared to the estimated duration. A method for extracting video context from external descriptions.
5. 2. The method of extracting video context from an external description according to claim 1, further comprising: the context database manages prohibited procedures within the inspected event graph; If the generated context label contains one of the prohibited procedures, issue a warning about non-compliance with standard operating procedures. A method for extracting video context from external descriptions.
6. 2. The method of extracting video context from an external description according to claim 1, further comprising: Further, a task is identified from the generated context label; Calculate the cycle time for the identified tasks based on the length of the video across the generated context labels associated with the identified tasks. A method for extracting video context from external descriptions.
7. 2. The method of extracting video context from an external description according to claim 1, further comprising: For one or more of the generated context labels that indicate an anomalous event, provide an indication that an anomalous event has occurred. A method for extracting video context from external descriptions.
8. 2. The method of extracting video context from an external description according to claim 1, further comprising: The generated contextual labels are stored together with the video and the video is indexed by search engines. A method for extracting video context from external descriptions.
9. In a system for extracting video context from external descriptions, a context database and a processor; The processor, Receives an input video for processing; running an image labeller on the input video to generate a plurality of labels corresponding to a plurality of images in the input video; Compute a zero-shot image embedding vector for each label of the plurality of images; generating a context label to replace each of the labels corresponding to each of the plurality of images based on the temporally determined context labels for the current image and the previous image by referencing a context database with the image embedding vector to determine a context label from the inspected event graph; the nodes of the event graph represent events that may occur during the duration of the video; replacing each of the plurality of labels with the generated context label; A system for extracting video context from external descriptions.
10. 10. A system for extracting video context from an external description according to claim 9, comprising: The context database manages associations between the event embedding vector and context information for each node in the examined event graph. A system for extracting video context from external descriptions.
11. 10. A system for extracting video context from an external description according to claim 9, comprising: the processor is configured to generate a context label to replace each of the labels corresponding to each of the images by determining a node from the examined event graph associated with an event embedding vector in the context database that has a minimum distance from the image embedding vector; Generate a context label from context information associated with a node of the event graph. A system for extracting video context from external descriptions.
12. 12. A system for extracting video context from an external description according to claim 11, comprising: a node of the event graph is associated with an estimated duration in a context database; The processor weights the calculated distance between the image embedding vector and the event embedding vector based on the duration of time between the images of the video in time associated with the same context label and the current and previous images compared to the estimated duration. A system for extracting video context from external descriptions.
13. 10. A system for extracting video context from an external description according to claim 9, comprising: the context database manages prohibited procedures within the inspected event graph; The processor further comprises: If the generated context label contains one of the prohibited procedures, a warning is issued regarding non-compliance with standard operating procedures. A system for extracting video context from external descriptions.
14. 10. A system for extracting video context from an external description according to claim 9, comprising: The processor further comprises: Identify the task from the generated context label, Calculate the cycle time for an identified task based on the overall video length of the generated context labels associated with the identified task. A system for extracting video context from external descriptions.
15. 10. A system for extracting video context from an external description according to claim 9, comprising: The processor further comprises: Controlling the display of one or more of the generated context labels that indicate an anomalous event to indicate that an anomalous event has occurred A system for extracting video context from external descriptions.
16. 10. A system for extracting video context from an external description according to claim 9, comprising: The processor further comprises: Save the generated contextual labels along with the video to index the video in search engines A system for extracting video context from external descriptions.
Citation Information
Patent Citations
Behavior recognition device, control method therefor, and program
JP2021082137A
Information processing apparatus, information processing method, and information processing program
JP2023135777A
Text-conditioned video representation
US20230351753A1