Interactive time series analysis system and method of the same
The integration of CLIP and LLMs in an interactive time series analysis system addresses the limitations of existing video analytics by enabling dynamic event interpretation and conversational querying, improving manufacturing efficiency and accuracy.
Patent Information
- Application Number
- JP2025036203
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-16
- Filing Date
- 2025-03-07
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-03-07
AI Technical Summary
Existing video analytics technologies in manufacturing lack the ability to perform detailed, context-aware analysis of sequences of events over time and provide conversational interfaces for abstract queries, leading to inefficiencies and human errors.
An interactive time series analysis system integrating CLIP for visual data interpretation and large-scale language models (LLMs) for context-rich natural language interaction, enabling dynamic event interpretation and conversational querying.
Minimizes human error and enhances manufacturing efficiency by providing accurate, context-aware analysis and interactive responses to complex queries.
Smart Images

Figure 2025162971000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure is directed generally to factory systems, and more specifically to video content interpretation and text-context-based augmented retrieval using large-scale language models (LLMs). In particular, the present disclosure is directed to an interactive time series analysis system and method for video content interpretation using zero-shot inference and text-context-based augmented retrieval. [Background technology]
[0002] In the manufacturing industry, frequent production stoppages due to human error are a major challenge. Historically, records of individual worker actions and patterns have been kept on paper and not digitized, leaving gaps in efficiently understanding and preventing such human errors.
[0003] Expectations for digitizing human behavior patterns on manufacturing floors are rising with the aim of mitigating factory downtime and improving operational efficiency. Recent advances in artificial intelligence (AI), particularly machine learning models for video analytics, are beginning to address this need. These advances go beyond the analysis of single images to enable contextual analysis of video frames, providing nuanced and accurate interpretations of visual data in real time. Summary of the Invention [Problem to be solved by the invention]
[0004] The application of supervised learning AI, which has long been studied, has shown some effectiveness in digitizing these patterns, leading to the analysis of production bottlenecks and the potential for maximizing productivity. However, this approach faces challenges, such as the significant effort required to optimize the AI model and the difficulty of horizontal deployment across different locations.
[0005] Furthermore, foundational models such as large-scale language models (LLMs) and contrastive language-image pre-training (CLIP) offer exciting opportunities for zero-shot learning, allowing them to be used on new data without specific training. These have demonstrated promising applications in classification, object recognition, and image captioning. This can significantly reduce the time and resources required to train and deploy models. However, because they depend on the quality of the input data, the output may contain irrelevant or inaccurate information, posing challenges to accuracy and reliability. Accuracy is particularly poor when analyzing video with complex backgrounds that may contain undetected objects.
[0006] Against this backdrop, the manufacturing industry is undergoing a transformation where digitizing human behavior patterns through optimized AI models can unlock new levels of productivity and operational insights. The balance between high accuracy of shop-floor-optimized AI and broad applicability but lower accuracy of basic models is a crucial area of development.
[0007] Existing technologies that use basic models primarily focus on object detection and image classification without deeply integrating contextual or temporal analysis. For example, traditional machine learning models can identify objects and anomalies within a single frame, but struggle to understand the significance of sequence and change over time. Related technology implementations include various approaches to video analytics and anomaly detection, but often lack the integration of natural language processing (NLP) to enhance contextual understanding and interactivity. While products and services on the market may offer basic video analytics, they do not fully leverage the synergy between visual data interpretation and natural language understanding.
[0008] In a related art implementation, there is a method for querying video data. The video data is divided into shots based on image frames, audio data, and caption data associated with the same caption, and feature quantities for each shot are extracted as vector information. The vector information for each shot is processed collectively using a multi-layer neural network to generate a feature vector for the entire video data. The optimal video data is selected from video storage based on the similarity with the comparison feature vector. In such a related art implementation, time series analysis is not performed on a frame-by-frame basis.
[0009] Another related-art implementation is a computer vision system that learns directly from text descriptions, bypassing the need for labeled data. By pre-training on 400 million image-text pairs collected from the web, these related-art models use natural language to identify and describe visual concepts, enabling zero-shot classification across diverse tasks without task-specific training. This related-art method rivals the performance of traditional fully supervised models like ResNet-50 on ImageNet, demonstrating significant adaptability and efficiency. However, this approach does not involve time-series information processing, instead focusing on leveraging natural language for visual recognition.
[0010] The implementations described herein address these challenges and provide a novel solution that leverages the advantages of both approaches to minimize human error and improve manufacturing efficiency. The embodiments described herein can be applied not only to the digitization of human activities, but also to the digitization of other devices, materials, autonomous vehicles (AGVs), etc. For ease of understanding, the embodiments described herein are described with respect to the digitization of human activities, but are not limited thereto.
[0011] A key challenge unresolved by related technologies is their limited ability to perform detailed, context-aware analysis of sequences of events in video data. While existing solutions can perform comprehensive semantic extraction and scene classification of videos, they lack the ability to dynamically interpret the meaning of events as they change over time or provide a conversational interface for abstract, unambiguous queries about time. The implementation described in this paper aims to fill this gap by providing time series analysis of video data by integrating CLIP for visual data interpretation and LLM for context-rich natural language interaction. Specifically, we aim to provide an interactive time series analysis system and method that utilizes a conversational interface for dynamically interpreting the meaning of events as they change over time and for abstract, unambiguous queries about time. [Effects of the Invention]
[0012] According to the present invention, it is possible to minimize human error and improve manufacturing efficiency. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 1 is a diagram illustrating an example of an outline of the operation of an interactive time series analysis system according to an embodiment. [Figure 2] FIG. 2 illustrates an interactive time series analysis system tuned to analyze the mean time to detection (MTTD) of a specified process, according to an embodiment. [Figure 3A] FIG. 3A illustrates a sequence diagram associated with the system described herein, according to an embodiment. [Figure 3B] FIG. 3B illustrates an example question and answer tree for obtaining related information, according to an embodiment. [Figure 4] FIG. 4 is a diagram illustrating an example of preprocessing according to the embodiment. [Figure 5] FIG. 5 shows the MTTD measurement being queried according to an embodiment. [Figure 6] FIG. 6 is a diagram illustrating an example of a user interface of the interactive time series analysis system according to the embodiment. [Figure 7A] FIG. 7A is a diagram illustrating an example of execution of the context information calculation unit according to the embodiment. [Figure 7B] FIG. 7B shows a probability column when this scene is converted into text by each event probability calculation unit according to the embodiment. [Figure 7C] FIG. 7C illustrates an example of a calculation by the context information calculation unit, according to an embodiment. [Figure 7D] FIG. 7D illustrates an example of a calculation by the context information calculation unit, according to an embodiment. [Figure 8] FIG. 8 shows another example of an interactive time series system designed for inventory control within the XYZ process. [Figure 9] FIG. 9 illustrates an exemplary computing environment with an exemplary computer device in an exemplary implementation of a system for interactive time series analysis. DETAILED DESCRIPTION OF THE INVENTION
[0014] The following detailed description provides details of the figures and examples of the present application. Reference numerals and redundant element descriptions between figures have been omitted for clarity. Terms used throughout this specification are provided by way of example and are not intended to be limiting. For example, the use of the term "automatic" can include fully automatic or semi-automatic implementations with user or administrator control over certain aspects of the implementation, depending on the desired implementation of one skilled in the art practicing embodiments of the present invention. Selection can be performed by a user through a user interface or other input means, or can be performed through a desired algorithm. The exemplary implementations described herein can be utilized either alone or in combination, and the functionality of the exemplary implementations can be implemented in any manner according to the desired implementation.
[0015] FIG. 1 illustrates an exemplary system for interactive time series analysis (interactive time series analysis system) according to an embodiment, and is a diagram illustrating an example of an outline of the operation of a processing unit of the interactive time series analysis system. The embodiments described herein include an innovative system that combines CLIP (Corresponding Event Probability Calculator 103) for advanced image analysis with large-scale language models (LLM-based Analysis 107) to provide a retriever-augmented generation (RAG)-based chat system for interactive time series analysis. For example, by identifying peaks and valleys in event probability across video frames in the contextual information calculator 105 and enriching these insights with manufacturing contextual information, the system allows users to interactively query the system using natural language. This dual approach not only improves the accuracy of event detection and classification in video data, but also revolutionizes the way users interact with and understand the analysis, enabling comprehensive, conversational answers to queries such as, "Show me frames where there may be equipment issues" or "When was the last time a part was not present?"
[0016] As shown in Figure 1, in this embodiment, an interactive time series analysis system includes a step of calculating probability information of objects in each frame of video data, a step of calculating the state of the object at a specified time based on the probability information from the past to the present, and a step of inputting the state at the specified time into a natural language model (LLM), enabling analysis and prediction based on natural language.
[0017] Depending on the desired implementation, the calculation of the state at the specified time can include a function that integrates past and current probability information. The processing unit of the interactive time series analysis system is configured to calculate the state of the at least one object at the specified time by integrating past and current probability information.
[0018] Depending on the desired implementation, the LLM can be configured to generate an interactive response based on input probability information and state information, i.e., the LLM generates an interactive response based on input probability information and the state of the at least one object at a specified time.
[0019] Depending on the desired implementation, the step of calculating the state at the specified time uses a probabilistic model that takes into account dynamic changes of the object. The processing unit of the interactive time series analysis system is configured to calculate the state at the specified time using a probabilistic model that incorporates dynamic changes of the at least one object.
[0020] Depending on the desired implementation, calculating the state at the specified time may include calculating / predicting future probability information, and based on this future probability information, an LLM is used to facilitate analysis and prediction of future events or states. The processing unit of the interactive time series analysis system is configured to calculate the state at the specified time by predicting the future probability information and use the future probability information as input to the LLM, thereby facilitating analysis and prediction of future events.
[0021] Depending on the desired implementation, the LLM can be configured to dynamically adjust its responses depending on the context of the generated dialogue response (hereinafter also referred to as context) and the user's request for additional information.
[0022] Depending on the desired implementation, the LLM is configured to present predictive information based on future probability information to the user as warnings, suggestions, or instructions for action.
[0023] Depending on the desired implementation, there may be a preprocessing module that optimizes the label information before calculating the object probability information, thereby improving the accuracy of subsequent analysis and prediction. The processing unit of the interactive time series analysis system is configured to optimize the label information through a preprocessing procedure before calculating the probability information.
[0024] Depending on the desired implementation, LLM can utilize a Retriever-Augmented Generation (RAG) approach to handle complex queries and can integrate contextual information from external knowledge bases to enrich dialogue responses.
[0025] Depending on the desired implementation, there may also be a feedback mechanism that allows the system to learn from user interactions and improve the predictive model over time, thereby increasing the relevance and accuracy of the output.
[0026] In the context of image processing and computer vision, an object in an image frame refers to a distinct item, shape, or region that is the subject of analysis and classification. These objects can be anything from people, vehicles, and animals to more abstract concepts like shapes and text. Labels, on the other hand, are tags or names assigned to these objects to identify them as belonging to a particular category or class. For example, in a street scene, objects such as cars, pedestrians, and traffic lights are labeled based on their appearance and characteristics in the image.
[0027] In classification problems, probability information refers to the likelihood or confidence that a given object or instance belongs to a particular class or category. This information is typically output by a classification model, such as a neural network, which processes input data (such as images or a set of features) and predicts the class membership of each object. Probability is often expressed as a value between 0 and 1, with higher values indicating greater confidence in the classification. For example, a model might predict that an image of a cat has a 95% probability of belonging to the "cat" category and a 5% probability of belonging to the "dog" category.
[0028] State information derived from time series data includes the state and attributes of a system or process at different points in time based on past and current data. In the context of video analytics and sequential data processing, this involves understanding how object attributes (such as position, movement, and appearance) change over time. Analyzing these dynamic changes can infer the current state of a system and predict its future state. For example, by tracking the movement of a vehicle across successive frames of video, its speed and direction can be calculated and its future position can be predicted. When using moving cameras, such as those on AGVs, synchronizing the extracted positions from the AGV with probabilistic information can also correct for camera-to-subject relationships between the camera and the vehicle.
[0029] Retriever-Augmented Generation (RAG) is a natural language processing (NLP) technique that combines the retrieval of relevant information from large text corpora (the retrieval part) with a generative model that can generate human-like text based on the retrieved information (the generation part). This approach allows the model to incorporate external knowledge relevant to the current context or query, thereby improving the quality and relevance of the generated response. In practical applications, RAG can be used to answer complex questions, generate detailed explanations, or create content by accessing and synthesizing information from diverse sources. For example, when asked a specific question, a RAG system can search a database of documents to find relevant information and use that information to construct a coherent and useful answer.
[0030] Figure 2 illustrates an interactive time series analysis system tailored to analyze the mean time to detection (MTTD) of a specified process, according to an embodiment. The example in Figure 2 has a specified process referred to as the "ABC" process, specifically for the month of May. The system consists of three main components: a time series analysis component 100, a data communications component 200, and mass data storage 300.
[0031] The entry of a user prompt (1) is the starting point. If the initial data entry is insufficient, the system can request additional information via the LLM-based user interface (UI) 101. This interactive Q&A (if necessary) ensures the system has all the information it needs to proceed with the analysis. The UI 101 queries the mass data storage 300 for relevant video data (3) related to the ABC process. The mass data storage 300 stores and manages multiple videos. The query (2) is entered into the video data storage 301 via the data communication component 200, which facilitates the transfer of all data from the mass data storage 300 to the analysis component 100.
[0032] The video frame extraction unit 102 divides the video data into individual image frames (4). These frames, along with MTTD labels (5), enter the event probability calculation unit 103, where the probability (6) of each event is determined, i.e., the probability of at least one object in each frame. In the case of MTTD analysis, the labels (5) for MTTD might be a red (or green) traffic light and a worker responding to a problem.
[0033] The system also incorporates a time series probability storage unit 104 that stores time series probability strings from the past to the present. The time series probability strings stored in this time series probability storage unit 104 are combined with context strings (7) and processed by a context information calculation unit 105 to create comprehensive information that encompasses both probability and contextual nuances. For example, it can identify important moments such as sharp peaks and valleys in event probability and identify worker response times, which are essential context data for MTTD evaluation. This information is stored in a context information calculation and storage unit 106.
[0034] The LLM-based analysis unit 107 then utilizes this rich context information (8) stored in the context information calculation and storage unit 106, along with the initial user prompt (9) containing relevant information, to perform a detailed time series analysis, which may generate analytical data such as MTTD-related insights (10).
[0035] Finally, the LLM-based UI 101 employs and visualizes the analytical results, generating dynamic interactive responses (11) for the user. This can include interactive feedback such as clarifying the significance of traffic light colors in the operational context or explaining the MTTD metric within the system. Furthermore, the system's user-friendly interface allows for easy input and interpretation of complex time-series data, thereby supporting optimization of the decision-making process associated with the ABC process. While not discussed here, in addition to probability information, external data, such as data from a programmable logic controller (PLC), can also be used as input to the system.
[0036] 3A shows a sequence diagram associated with the system described herein according to an exemplary implementation. The externally referenced sections ("ref") describe pre-processing to add information required later in the process to the ambiguous user prompt.
[0037] In the example flow of FIG. 3A, first, a user provides a user prompt (1) to the UI 101. The UI 101 can then conduct a Q&A (question and answer) session and generate a query to gather more information about the provided prompt. The query (2), which in this example is a video related to the ABC process, is sent to the video data storage 301. The related video (3) is retrieved from the video data storage 301 and processed by the video frame extraction unit 102 to extract frames (4). Each extracted frame (4) is processed by the event probability calculation unit 103 to calculate probability information for at least one object in each frame. The event probability calculation unit 103 also receives labels for the MTTD (5) generated by the UI 101. The frames and labels are processed by the event probability calculation unit 103 to determine the probability of each event. This process is repeated for each frame.
[0038] The probability of each event is provided to a context information calculation unit 105 configured to determine indexed probabilities of the time series events (7). The indexed time series event probabilities are processed by the context information calculation unit 105 to generate context information, which is stored in a context information calculation and storage unit 106 and input for processing by an LLM-based analysis unit 107 (8).
[0039] The LLM-based analysis unit 107 is configured to take in relevant information (9) from the UI 101 as well as user prompts, and context information (8) from the context information calculation and storage unit 106, and return analyzed data (10). In this example, the relevant information (9) included in the user prompt is "Green light indicates normal operation, red light indicates abnormal event. MTTD indicates the average time it took for the worker to discover the problem." The LLM-based analysis unit 107 returns the analyzed data (10) to the UI 101, which provides the data in a visualization (11) that can be viewed by the user.
[0040] FIG. 3B illustrates an example of a question and answer tree for obtaining relevant information according to an embodiment. FIG. 4 illustrates an example of preprocessing according to an embodiment. In the example of FIG. 4, multiple questions are asked to add information required for each subsequent processing unit to the information included in the user prompt of the LLM-based UI. In this example, questions #1 to #5 shown in FIG. 3B are implemented to enhance the RAG system to realize the relevant information shown in the fourth column of FIG. 3B. The UI may re-ask the user in a pre-fixed format if an unexpected prompt is encountered, but the present disclosure is not limited thereto, and other implementations may be utilized to facilitate a desired implementation.
[0041] As shown in FIG. 4, a user prompt (step 400, hereafter referred to as simply S) is provided, which in this example is "Please analyze the MTTD of the ABC process during May." In S401, the preprocessing of FIG. 3B is executed starting with question #1, "Does the user prompt contain 'analyze'?" If yes (YES), question #2 is skipped; if not (NO), the flow proceeds to S402 to ask a second question. In S402, question #2 is asked: "Does the user prompt contain 'acquire'?" If yes (YES), the flow proceeds to S403; if not (NO), the flow proceeds to S406.
[0042] In S403, question #3 is asked: "Does the user prompt include 'MTTD'?" If yes (YES), flow continues to S405; if not (NO), flow continues to S404. In S404, question #4 is asked: "Does the user prompt include 'SOP'?" If yes (YES), flow continues to S405; if not (NO), flow continues to S406.
[0043] At S405, question #5 is asked: "Does the user prompt include a specific process and a specific month?" If so (YES), the flow ends; if not (NO), the flow proceeds to S406. At S406, the flow generates a "Please re-ask the following" on the LLM-based UI: (1) Analyze SOP compliance of the AZ process (2) Get video related to the NM process.
[0044] Figure 5 illustrates how an MTTD measurement may be queried by a user in accordance with an embodiment. Specifically, Figure 5 illustrates an MTTD measurement queried by a user in the context information calculation and storage unit 106 by detecting a change in probability information from the past to the present that exceeds a predetermined threshold.
[0045] 6 is a diagram illustrating an example of a user interface (UI) of a system for interactive time series analysis, according to an embodiment. As shown in FIG. 6, the UI can also display associated video data (3), probabilities of each event (7), and interactive responses (11) through an LLM-based UI 101 that takes user prompts (1).
[0046] FIG. 7A is a diagram illustrating an example of execution of the context information calculation unit 105 according to the embodiment. The example in FIG. 7A is an example of execution in the case of MTTD, in which a red light turns on at the exact center time k between Frame-k-1, Frame-k, and Frame-k+1. The worker confirms that the red light is on at the center time m between Frame-m-1, Frame-m, and Frame-m+1. The red light turns off and changes to a green light at the center time n between Frame-n-1, Frame-n, and Frame-n+1. FIG. 7B shows a probability column when this scene is converted into text by the event probability calculation unit 103, and FIGS. 7C and 7D show examples of calculations by the context information calculation unit 105. FIG. 7C shows a case where only frames with large changes in probability are extracted, and FIG. 7D shows a case where the change in probability from the immediately preceding frame is calculated and displayed.
[0047] Figure 8 shows another example of an interactive time series analysis system designed for inventory management within the XYZ process. This use case demonstrates how the system can detect probability information for parts identified with the classification label "Parts" and provide guidance on material delivery timing, as well as warnings of future stockouts, and other suggestions or action directives according to the desired implementation.
[0048] The details of the role of each component in this use case are as follows:
[0049] User Prompt (1): The user asks the system, "By when should the part be delivered to station XYZ?" This input starts the analysis process.
[0050] LLM-Based UI 101: Driven by a large-scale language model (LLM), the system's user interface interprets the user's prompts and determines whether additional information is needed. If necessary, it can engage in Q&A to clarify or expand on the user's request.
[0051] Related video query (2): The UI sends a query to the video frame extraction unit 102 to retrieve video data related to the part from the mass data storage 300.
[0052] Mass Data Storage 300: Mass data storage 300 stores a wide range of image data, including time-lapse images of XYZ processes.
[0053] Video Frame Extractor 102: Once the relevant video data storage 301 has been identified, the Video Frame Extractor 102 extracts frames from the video for analysis.
[0054] Frame and event probability calculation and time series probability storage unit: The probability calculation unit for each event 103 determines the probability of each event based on the individual frame (4) and the part label (5) provided by the UI, and inputs the time series probability string (6) into the time series probability storage unit 104. The time series probability string (6) is processed by the time series probability storage unit 104 to calculate the probability of each event (7) and the probability of the time series event (time series probability).
[0055] Context information calculation unit 105 and context information storage unit 106: The calculated probability and label are combined with the context information string and output (8), providing a comprehensive understanding of the event within its operational context.
[0056] LLM-based analysis unit 107: The LLM processes all the above information and performs a detailed analysis (9). This analysis can include approximate expressions of the parts in terms of their probabilities over time, visualized in a graph.
[0057] Interactive Responses (10 and 11): Based on the analysis results, the LLM-based UI provides an interactive response to the user, e.g., "You should aim to deliver the part no later than frame 10."
[0058] The components work in concert to enable the system to not only detect current inventory status but also predict future needs, enabling effective inventory management and optimization in the XYZ process. The system's ability to process and analyze video data through the integration of frame extraction, event probability calculation, and contextual analysis, resulting in an LLM-based predictive response, exemplifies a cutting-edge approach to parts and materials management in industrial environments.
[0059] The implementation described herein builds on the seamless integration of CLIP and LLM technologies for interactive video analysis, as illustrated. The system's ability to analyze video data with CLIP by identifying relevant objects and events, and the detailed process of passing information to LLM to generate context-rich interactive responses, provides a solid foundation. The innovative integration of visual data analysis with natural language processing and search augmentation highlights the unique interactive and insightful analytical tools provided by this invention. The addition of preprocessing for data quality improvement, RAG for complex query processing, and feedback mechanisms for model refinement further expands the system's capabilities for real-time monitoring, predictive maintenance, SOP compliance, and defect detection.
[0060] One possible implementation example is a prediction use case based on multiple data sets, such as data from CLIP (keyword search or similar image search) and data from PLC, which is useful for events that cannot be resolved using image information alone.
[0061] Because CLIP allows for the use of long sentences rather than single words to select labels, the RAG system can be used to label the "worker" organizing luggage in the image as "worker organizing luggage." By modifying the flowchart in Figure 4, labeling can be optimized semi-automatically. It is possible to refer to the work procedures registered in PLM from the process name and perform labeling based on the work procedures written in the work procedures.
[0062] 9 illustrates an exemplary computing environment having an exemplary computer device suitable for use in some exemplary implementations of a system for interactive time series analysis, a database for managing multiple videos, etc. The computing device 905 of the computing environment 900 can include one or more processing units, cores, or processors (processing unit) 910, memory 915 (e.g., RAM, ROM, and / or the like), internal storage 920 (e.g., magnetic, optical, solid-state storage, and / or organic), and / or an I / O interface 925, any of which can be coupled over a communication mechanism or bus 930 for communicating information or embedded in the computing device 905. The I / O interface 925 can also be configured to receive images from a camera or provide images to a projector or display, depending on the desired implementation.
[0063] The computing device 905 can be communicatively coupled to an input / user interface 935 and an output device / interface 940. Either or both of the input / user interface 935 and the output device / interface 940 can be wired or wireless interfaces and can be detachable. The input / user interface 935 can include any device, component, sensor, or interface, physical or virtual, that can be used to provide input (e.g., buttons, touchscreen interfaces, keyboards, pointing / cursor control, microphones, cameras, Braille, motion sensors, optical readers, and / or the like). The output device / interface 940 can include displays, televisions, monitors, printers, speakers, Braille, etc. In some exemplary implementations, the input / user interface 935 and the output device / interface 940 can be embedded in or physically coupled to the computing device 905. In other exemplary implementations, other computing devices can function as or provide the functionality of the input / user interface 935 and the output device / interface 940 of the computing device 905.
[0064] Examples of computing devices 905 may include, but are not limited to, highly mobile devices (e.g., smartphones, devices mounted on vehicles and other machines, devices carried by humans and animals, etc.), mobile devices (e.g., tablets, notebooks, laptops, personal computers, portable televisions, radios, etc.), and devices not designed for mobility (e.g., desktop computers, other computers, information kiosks, televisions with one or more processors embedded therein and / or coupled thereto, radios, etc.).
[0065] Computing device 905 may be communicatively coupled (e.g., via I / O interface 925) to external storage 945 and a network 950 to communicate with any number of networked components, devices, and systems, including one or more computing devices of the same or different configurations. Computing device 905 or any connected computing device may function, provide a service, or be referred to as a server, a client, a thin server, a general machine, a special purpose machine, or another label.
[0066] I / O interface 925 can include, but is not limited to, wired and / or wireless interfaces using any communication or I / O protocol or standard (e.g., Ethernet, 802.11x, Universal System Bus, WiMAX, modem, cellular network protocol, etc.) for communicating information to and from at least all connected components, devices, and networks in computing environment 900. Network 950 can be any network or combination of networks (e.g., the Internet, a local area network, a wide area network, a telephone network, a cellular network, a satellite network, etc.).
[0067] The computing device 905 can use and / or communicate with computer-usable or computer-readable media, including transient and non-transitory media. Transitory media include transmission media (e.g., metallic cables, optical fibers), signals, carrier waves, etc. Non-transitory media include magnetic media (disks, tapes, etc.), optical media (CD ROM, digital video disks, Blu-ray disks, etc.), solid-state media (RAM, ROM, flash memory, solid-state storage, etc.), and other non-volatile storage or memory.
[0068] The computing device 905 can be used to implement techniques, methods, applications, processes, or computer-executable instructions in some exemplary computing environments. The computer-executable instructions can be retrieved from a transitory medium or stored on and retrieved from a non-transitory medium. The executable instructions can originate from one or more of a programming language, a scripting language, and a machine language (e.g., C, C++, C#, Java, Visual Basic, Python, Perl, JavaScript, etc. (Java is a registered trademark)).
[0069] The processor(s) 910 can run under any operating system (OS) (not shown) in a native or virtual environment. One or more applications can be deployed, including a logic unit 960, an application programming interface (API) unit 965, an input unit 970, an output unit 975, and an inter-unit communication mechanism 995 through which different units communicate with each other, with the OS, and with other applications (not shown). The described units and elements may vary in design, function, configuration, or implementation and are not limited to the provided description. The processor(s) 910 can be in the form of a hardware processor, such as a central processing unit (CPU), or a combination of hardware and software units.
[0070] In some example implementations, when information or instructions for execution are received by API unit 965, it may be communicated to one or more other units (e.g., logic unit 960, input unit 970, output unit 975). In some examples, logic unit 960 may be configured to control the flow of information between units and direct the services provided by API unit 965, input unit 970, and output unit 975 in some example implementations described above. For example, the flow of one or more processes or implementations may be controlled by logic unit 960 alone or in cooperation with API unit 965. Input unit 970 may be configured to obtain inputs for the calculations described in the example embodiments, and output unit 975 may be configured to provide outputs based on the calculations described in the example embodiments.
[0071] The processor(s) 910 are configured to receive a query, and when a query such as that shown in (2) is received, the processor(s) can be configured to calculate probability information of at least one object on each frame of a video from a plurality of videos related to the query, as shown in (3) to (6). Based on the probability information from the past up to a specified time, the processor(s) calculates a state of the at least one object at the specified time, as shown in (7), and inputs the state at the specified time into a large-scale language model (LLM) configured to output analysis and prediction in a natural language output responsive to the query, as shown in (8) to (10).
[0072] The processor(s) 910 may be configured to calculate the state of at least one object at a specified time by integrating past and present probability information, as described with respect to Figures 1 and 2.
[0073] Depending on the desired implementation, the LLM can be configured to generate an interactive response based on input of probability information and the state of at least one object at a specified time (11).
[0074] The processor(s) 910 may be configured to calculate the state at a specified time using a probabilistic model that incorporates dynamic changes of at least one object.
[0075] The processor(s) 910 can be configured to calculate states at specified times by predicting future probability information, as shown in FIG. 5, and use the future probability information as input to the LLM, thereby facilitating analysis and prediction of future events.
[0076] Depending on the desired implementation, the LLM can be configured to dynamically adjust its responses according to the context of the generated dialogue response and the user request for additional information, as shown in (9) through (11).
[0077] Depending on the desired implementation, the LLM can be configured to output predictions in natural language output based on future probability information as one or more of a warning, a suggestion, or an instruction to act, as shown in FIG. 8.
[0078] The processor(s) 910 may be configured to optimize the label information through a preprocessing procedure before calculating the probability information, as shown in (5), which may improve the accuracy of subsequent analysis and prediction.
[0079] Depending on the desired implementation, the LLM can be configured to perform a retriever-augmentation-generation (RAG) based approach in response to input to integrate contextual information from external knowledge bases, as described herein.
[0080] The processor(s) 910 may be configured to implement a feedback mechanism to refine the model used to calculate the probability information and state of the at least one object at a specified time from user interaction, as shown in (10) and (11).
[0081] Some portions of the detailed description are presented in terms of algorithms and symbolic representations of operations within a computer. These algorithmic descriptions and symbolic representations are the means used by those skilled in the data processing arts to convey the essence of their innovations to others skilled in the art. An algorithm is a sequence of defined steps leading to a desired end state or result. In the examples, the steps performed require measurable quantities of physical manipulations to achieve a measurable result.
[0082] Unless specifically stated otherwise, as will be apparent from the discussion, it will be understood that throughout this specification discussions using terms such as "processing," "computing," "calculating," "determining," "displaying," and the like may include operations and processes of a computer system or other information processing device that manipulate and transform data represented as physical (electronic) quantities in the registers and memory of the computer system into other data similarly represented as physical quantities in the memory or registers of the computer system or other information storage, transmission, or display device.
[0083] Exemplary embodiments also relate to apparatuses for performing the operations herein. This apparatus may be specially constructed for the required purposes, or it may include one or more general-purpose computers selectively activated or reconfigured by one or more computer programs. Such computer programs may be stored on computer-readable media, such as computer-readable storage media or computer-readable signal media. Computer-readable storage media may include tangible media, such as, but not limited to, optical disks, magnetic disks, read-only memory, random-access memory, solid-state devices and drives. Computer-readable signal media may include media such as carrier waves. The algorithms and displays presented herein are not inherently related to any particular computer or other apparatus. A computer program may include a pure software implementation containing instructions for performing the operations of a desired implementation.
[0084] Various general-purpose systems may be used with the programs and modules according to the embodiments herein, or it may prove convenient to construct more specialized apparatus to perform the desired method steps. Moreover, the embodiments are not described with reference to a particular programming language. It will be understood that a variety of programming languages may be used to implement the teachings of the embodiments described herein. Instructions in the programming language may be executed by one or more processing units, such as a central processing unit (CPU), processor, or controller.
[0085] As is known in the art, the operations described above can be performed by hardware, software, or some combination of software and hardware. Various aspects of the exemplary implementations may be implemented using circuits and logic devices (hardware), while other aspects may be implemented using instructions stored on a machine-readable medium that stores software that, when executed by a processor, causes the processor to perform methods for implementing the present application. Furthermore, some exemplary implementations of the present application may be performed solely in hardware, while other exemplary implementations may be performed solely in software. Furthermore, the various functions described may be performed in a single unit or may be spread across multiple components in any number of ways. When performed by software, the methods may be executed by a processor, such as a general-purpose computer, based on instructions stored on a computer-readable medium. If desired, the instructions may be stored on the medium in a compressed and / or encrypted format.
[0086] Additionally, other embodiments of the present application will be apparent to those skilled in the art from consideration of the specification and practice of the teachings herein. Various aspects and / or components of the described exemplary embodiments may be used alone or in any combination. It is intended that the specification and exemplary embodiments be considered exemplary only, with the true scope and spirit of the present application being indicated by the following claims. [Explanation of symbols]
[0087] 100 Time Series Analysis Components 101 User Interface (UI) 102 Video Frame Extraction Unit 103 CLIP (Each event probability calculation part) 104 Time series probability memory unit 105 Context information calculation section 106 Context information calculation / storage unit 107 LLM-Based Analysis Department 200 Data Communication Components 300 Mass Data Storage 301 Video Data Storage
Claims
1. a database for managing multiple videos; a processing unit configured to receive a query; The processing unit Calculating probability information of at least one object for each frame of a video related to the query from the plurality of videos; calculating a state of at least one object at a specified time based on probability information from the past up to the specified time; Inputting the state at a specified time into a large-scale language model (LLM) configured to output analysis and predictions in natural language output responsive to a query. An interactive time series analysis system.
2. 2. The interactive time series analysis system according to claim 1, the processing unit is configured to calculate a state of the at least one object at a specified time by integrating past and present probability information. An interactive time series analysis system.
3. 2. The interactive time series analysis system according to claim 1, The LLM generates an interactive response based on the input of the probability information and the state of the at least one object at the specified time. An interactive time series analysis system.
4. 2. The interactive time series analysis system according to claim 1, The processing unit and configured to calculate the state at the specified time using a probabilistic model that incorporates dynamic changes of the at least one object. An interactive time series analysis system.
5. 2. The interactive time series analysis system according to claim 1, The processing unit The system is configured to calculate the state at a specified time by predicting future probability information, and use the future probability information as input to the LLM, thereby facilitating the analysis and prediction of future events. An interactive time series analysis system.
6. 2. The interactive time series analysis system according to claim 1, The LLM: configured to dynamically adjust responses depending on the context of the generated dialogue response and requests for additional information from the user An interactive time series analysis system.
7. 2. The interactive time series analysis system according to claim 1, The LLM: and configured to output a prediction in the natural language output based on the future probability information as one or more of a warning, a suggestion, or an instruction to act. An interactive time series analysis system.
8. 2. The interactive time series analysis system according to claim 1, The processing unit A pre-processing procedure is configured to optimize label information before calculating the probability information. An interactive time series analysis system.
9. 2. The interactive time series analysis system according to claim 1, The LLM: It is configured to execute a Retriever-Augmented Generation (RAG)-based approach in response to input to integrate contextual information from external knowledge bases. An interactive time series analysis system.
10. 2. The interactive time series analysis system according to claim 1, The processing unit Implementing a feedback mechanism to refine the model used to calculate the probability information and state of the at least one object at the specified time from user interaction. An interactive time series analysis system.
11. 1. An interactive time series analysis method for analysis of an interactive time series analysis system having a database for managing a plurality of videos and a processing unit configured to receive a query, the method comprising: Executed by the processing unit: Calculating probability information of at least one object on each frame of a plurality of videos related to a query; Calculating a state of at least one object at a specified time based on probability information from the past up to the specified time; inputting the states at specified times into a large scale language model (LLM) configured to output an analysis and a prediction in natural language output responsive to the query. Interactive time series analysis methods.
12. 12. The interactive time series analysis method according to claim 11, The processing unit Computing the state of at least one object at a specified time integrates past and present probability information. Interactive time series analysis methods.
13. 12. The interactive time series analysis method according to claim 11, The LLM: generating an interactive response based on the input of the probability information and the state of the at least one object at the specified time; Interactive time series analysis methods.
14. 12. The interactive time series analysis method according to claim 11, Calculating the state at the specified time uses a probabilistic model that incorporates dynamic changes of the at least one object. Interactive time series analysis methods.
15. 12. The interactive time series analysis method according to claim 11, The calculation of the state at the specified time by the processing unit is based on prediction of future probability information, and the use of future probability information as input to the LLM facilitates analysis and prediction of future events. Interactive time series analysis methods.
16. 12. The interactive time series analysis method according to claim 11, The LLM: configured to dynamically adjust responses according to the context of the generated interactive response and the user request for additional information. Interactive time series analysis methods.
17. 12. The interactive time series analysis method according to claim 11, The LLM: and configured to output a prediction in the natural language output as one or more of a warning, a suggestion, or an instruction to act based on future probability information. Interactive time series analysis methods.
18. 12. The interactive time series analysis method according to claim 11, The processing unit Before calculating the probability information, a preprocessing step is performed to optimize the label information. Interactive time series analysis methods.
19. 12. The interactive time series analysis method according to claim 11, The LLM: and configured to implement a Retriever-Augmented Generation (RAG) based approach to integrate contextual information from an external knowledge base in response to the input. Interactive time series analysis methods.
20. 12. The interactive time series analysis method according to claim 11, The processing unit Implementing a feedback mechanism that refines the model used to calculate the probability information and state of the at least one object at a specified time from user interaction. Interactive time series analysis methods.
Citation Information
Patent Citations
Text creation system, text creation program, and report database production method
JP2025127401A
Event determination method, determination device, and determination program using text video classification model
JP7409731B1
Information analysis system, information analysis method, and program
WO2024231989A1
Information processing device, information processing method, and recording medium
WO2024261895A1