Interactive time series analysis system and method

The integration of CLIP and LLM in an interactive time-series analysis system addresses the challenge of dynamic event interpretation in video data, enhancing manufacturing efficiency through accurate and conversational analysis.

JP7844695B2Active Publication Date: 2026-04-13HITACHI LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
HITACHI LTD
Filing Date
2025-03-07
Publication Date
2026-04-13

AI Technical Summary

Technical Problem

Existing technologies lack the ability to perform detailed, context-aware analysis of event sequences within video data, failing to dynamically interpret the meaning of events as they change over time and provide a conversational interface for abstract, unambiguous queries about time, leading to inefficiencies in manufacturing processes due to human error.

Method used

An interactive time-series analysis system integrating CLIP for visual data interpretation and LLM for context-rich natural language dialogue, enabling dynamic event interpretation and conversational querying.

Benefits of technology

Minimizes human error and enhances manufacturing efficiency by providing accurate, context-aware analysis and interactive responses to complex queries, improving productivity and operational insights.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007844695000001
    Figure 0007844695000001
  • Figure 0007844695000002
    Figure 0007844695000002
  • Figure 0007844695000003
    Figure 0007844695000003
Patent Text Reader

Abstract

To provide a system and a method for an interactive time series analysis utilizing a conversational interface for abstract and unambiguous queries that dynamically interpret a meaning of events that change from moment to moment over time.SOLUTION: An interactive time series analysis system includes: a mass data storage 300 that manages a plurality of videos; and a time series analysis component 100 that, when a query is received, calculates probability information of at least one object on each frame of a video from a plurality of videos associated with the query, calculates a state of the at least one object at a specified time based on probability information from a past to the specified time, and inputs a state of the specific time into a large-scale language model (LLM) that outputs an analysis and a prediction by a natural language output that responds the query.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure is generally directed to factory systems, and more specifically, to the interpretation of video content by the use of large language models (LLMs) and extended search based on text context. In particular, this disclosure is directed to an interactive time-series analysis system and method for video content interpretation by zero-shot inference and extended search based on text context.

Background Art

[0002] In the manufacturing industry, production stoppages due to human error occur frequently and pose a major problem. Historically, records of individual workers' actions and patterns were stored on paper and not digitized, leaving a gap in the efficient understanding and prevention of such human errors.

[0003] The expectation for digitizing human behavior patterns in manufacturing sites is increasing for the purpose of alleviating factory stoppages and improving work efficiency. Recent advances in artificial intelligence (AI), particularly machine learning models for video analysis, are beginning to address this need. These advances enable not only the analysis of a single image but also the contextual analysis of video frames, providing accurate interpretations rich in the nuances of visual data in real time.

Summary of the Invention

Problems to be Solved by the Invention

[0004] The application of supervised learning AI, which has been studied for a long time, has shown a certain effect in digitizing such patterns, leading to the analysis of production bottlenecks and the possibility of maximizing productivity. However, this approach has problems such as the need for a great deal of effort in optimizing the AI model and the difficulty of horizontal deployment between different sites.

[0005] ]> Furthermore, foundational models such as Large-Scale Language Models (LLMs) and Contrastive Language Image Pre-training (CLIP) offer exciting opportunities for zero-shot learning, allowing them to be used on new data without specific training. These have demonstrated promising applications in classification, object recognition, and image captioning. This can significantly reduce the time and resources required to train and deploy models. However, because they depend on the quality of the input data, the output may contain irrelevant or inaccurate information, posing challenges to accuracy and reliability. Accuracy is particularly poor when analyzing videos with complex backgrounds that may contain objects that should not be detected.

[0006] Against this backdrop, the manufacturing industry is undergoing a period of transformation, and by digitizing human behavior patterns through optimized AI models, it is possible to unlock new levels of productivity and operational insights. Balancing the high accuracy of AI optimized for the workplace with the broad applicability but lower accuracy of basic models is a critical area of ​​development.

[0007] Existing technologies using foundational models primarily focus on object detection and image classification, lacking deep integration of context and temporal analysis. For example, traditional machine learning models can identify objects or anomalies within a single frame, but struggle to understand the significance of sequences and changes over time. Implementations of related technologies include various approaches to video analysis and anomaly detection, but often lack the integration of natural language processing (NLP) to enhance contextual understanding and interactivity. While products and services on the market may offer basic video analysis, they do not fully leverage the synergy between visual data interpretation and natural language understanding.

[0008] One embodiment of the related technology involves querying video data. The video data is divided into shots based on image frames, audio data, and caption data associated with the same caption, and the features of each shot are extracted as vector information. The vector information of each shot is then processed together using a multilayer neural network to generate a feature vector for the entire video data. Based on the similarity to the comparison feature vector, the optimal video data is selected from the video storage. In such implementations of the related technology, frame-by-frame time-series analysis is not performed.

[0009] Another implementation of related arts involves computer vision systems that learn directly from text descriptions, avoiding the need for labeled data. By pre-training on 400 million image-text pairs collected from the web, such related arts models can use natural language to identify and describe visual concepts, enabling zero-shot classification across diverse tasks without task-specific training. This related arts approach has shown significant adaptability and efficiency, comparable in performance to traditional fully supervised models such as ResNet-50 on ImageNet. However, this approach does not involve time-series information processing, instead focusing on the use of natural language for visual recognition.

[0010] The implementation examples described herein address these challenges and provide novel solutions that leverage the strengths of both approaches to minimize human error and improve manufacturing efficiency. The embodiments described herein can be applied not only to the digitalization of human behavior but also to the digitalization of other devices, materials, and autonomous vehicles (AGVs). For ease of understanding, the embodiments described herein will focus on the digitalization of human behavior, but are not limited to this.

[0011] A major challenge that remains unresolved by related technologies is the limited ability to perform detailed, context-aware analysis of event sequences within video data. While existing solutions can perform comprehensive semantic extraction and scene classification of video, they cannot dynamically interpret the meaning of events as they change over time or provide a conversational interface for abstract, unambiguous queries about time. The implementation examples described in this book aim to bridge this gap by providing time-series analysis of video data by integrating CLIP for visual data interpretation and LLM for context-rich natural language dialogue. In other words, the goal is to provide an interactive time-series analysis system and method that can dynamically interpret the meaning of events as they change over time and utilize a conversational interface for abstract, unambiguous queries about time. [Effects of the Invention]

[0012] According to the present invention, it is possible to minimize human error and improve manufacturing efficiency. [Brief explanation of the drawing]

[0013] [Figure 1] Figure 1 shows an example of the schematic operation of an interactive time series analysis system according to an embodiment. [Figure 2] Figure 2 shows an interactive time series analysis system adapted to analyze the mean detection time (MTTD) of a specified process, according to an example. [Figure 3A] Figure 3A shows a sequence diagram related to the system described herein, according to an example. [Figure 3B] Figure 3B shows an example of a question and answer tree for obtaining relevant information according to the embodiment. [Figure 4] Figure 4 shows an example of pretreatment according to the embodiment. [Figure 5] Figure 5 shows the queried MTTD measurement according to the example. [Figure 6] Figure 6 shows an example of a user interface for an interactive time series analysis system according to an embodiment. [Figure 7A] Figure 7A shows an example of the execution of the context information calculation unit according to the embodiment. [Figure 7B] Figure 7B shows the probability column when the event probability calculation unit converts this scene into text, according to the embodiment. [Figure 7C] Figure 7C shows an example of a calculation performed by the context information calculation unit according to an embodiment. [Figure 7D] Figure 7D shows an example of a calculation performed by the context information calculation unit according to an embodiment. [Figure 8] Figure 8 shows another example of an interactive time-series system designed for inventory management within an XYZ process. [Figure 9] Figure 9 shows an exemplary computing environment with exemplary computer devices in an exemplary implementation of a system for interactive time series analysis. [Modes for carrying out the invention]

[0014] The following detailed description provides details of the figures and embodiments of this application. Reference figures between figures and descriptions of redundant elements have been omitted for clarity. Terms used throughout this specification are provided as examples and are not intended to limit them. For example, the use of the term “automatic” may include fully automatic or semi-automatic embodiments with user or administrator control over certain aspects of the embodiment, depending on a desired embodiment for those skilled in the art practicing embodiments of the present invention. Selection may be performed by a user through a user interface or other input means, or through a desired algorithm. The exemplary implementations described herein may be used individually or in combination, and the functions of the exemplary implementations may be implemented in any way according to a desired implementation.

[0015] FIG. 1 shows an exemplary system (interactive time-series analysis system) for interactive time-series analysis according to an embodiment, and is a diagram showing an example of an outline of operations by a processing unit of the interactive time-series analysis system. In the embodiments described herein, an innovative system is included that combines CLIP (each event probability calculation unit 103) for advanced image analysis and a large language model (LLM: LLM-based Analysis LLM-based analysis unit 107) to provide a RAG (Retriever-Augmented Generation)-based chat system for interactive time-series analysis. For example, the context information calculation unit 105 identifies peaks and valleys of event probabilities between video frames and enriches these insights with context information in the manufacturing industry, enabling the system to allow users to interactively query the system using natural language. This dual approach not only improves the accuracy of event detection and classification of video data, but also innovates the way users analyze and interact with and understand, enabling comprehensive and conversational responses to queries such as "Please display frames that may have problems with the device" and "When was the part not present".

[0016] As shown in FIG. 1, in the embodiment, an interactive time-series analysis system includes steps of calculating probability information of an object on each frame of video data, calculating a state of the object at a specified time based on probability information from the past to the present, and inputting the state at the specified time into a natural language model (LLM), enabling analysis and prediction based on natural language.

[0017] Depending on the desired implementation, a function for integrating past and current probability information can be included in the calculation of the state at the specified time. The processing unit of the interactive time-series analysis system is configured to calculate the state of the at least one object at the specified time by integrating past and current probability information.

[0018] <于 Depending on the desired implementation, the LLM can be configured to generate a dialogue response based on the input probability information and state information. That is, the LLM generates a dialogue response based on the input of probability information and the state of the at least one object at a specified time.

[0019] Depending on the desired implementation, in the step of calculating the state at a specified time, a probability model that takes into account the dynamic changes of the object is used. The processing unit of the interactive time series analysis system is configured to calculate the state at a specified time using a probability model incorporating the dynamic changes of at least one object.

[0020] Depending on the desired implementation, the calculation of the state at a specified time can include the calculation / prediction of future probability information, and based on this future probability information, by using the LLM, it facilitates the analysis and prediction of future events or states. The processing unit of the interactive time series analysis system calculates the state at a specified time by predicting future probability information and uses the future probability information as an input to the LLM, so as to be configured to facilitate the analysis and prediction of future events.

[0021] Depending on the desired implementation, the LLM can be configured to dynamically adjust the response according to the context of the generated dialogue response (hereinafter also referred to as context) and the user's request for additional information.

[0022] Depending on the desired implementation, the LLM is configured to present prediction information based on future probability information to the user as a warning, a suggestion, or an action instruction.

[0023] Depending on the desired implementation, there is a preprocessing module that optimizes the label information before calculating the object probability information, thereby improving the accuracy of subsequent analysis and prediction. The processing unit of the interactive time series analysis system is configured to optimize the label information by a preprocessing procedure before calculating the probability information.

[0024] Depending on the desired implementation, LLM can utilize the RAG (Retriever-Augmented Generation) approach to handle complex queries, integrating contextual information from external knowledge bases to enrich interactive responses.

[0025] Depending on the desired implementation, a feedback mechanism can be included that allows the system to learn from user interactions, improve its predictive model over time, and thereby increase the relevance and accuracy of its outputs.

[0026] In the context of image processing and computer vision, objects within an image frame refer to distinct items, shapes, or areas that are the subject of analysis or classification. These objects can be anything from people, vehicles, and animals to more abstract concepts like shapes or text. Labels, on the other hand, are tags or names assigned to identify that these objects belong to a particular category or class. For example, in a street scene, objects such as cars, pedestrians, and traffic lights are labeled based on their appearance and characteristics within the image.

[0027] In classification problems, probabilistic information refers to the likelihood or confidence that a given object or instance belongs to a particular class or category. This information is typically output by classification models such as neural networks, which process input data (such as images or sets of features) and predict the class membership of each object. Probabilities are often expressed as values ​​between 0 and 1, with higher values ​​indicating greater confidence in the classification. For example, a model might predict that an image of a cat has a 95% probability of belonging to the "cat" category and a 5% probability of belonging to the "dog" category.

[0028] State information obtained from time-series data includes the state and attributes of a system or process at different points in time, based on past and present data. In the context of video analysis and sequential data processing, this involves understanding how an object's attributes (position, motion, appearance, etc.) change over time. By analyzing these dynamic changes, it is possible to infer the current state of a system and predict its future state. For example, by tracking the movement of a vehicle across consecutive frames of video, its speed and direction can be calculated, and its future position can be predicted. When using moving cameras, such as those mounted on AGVs, the camera-subject relationship between the camera and the vehicle can also be corrected by synchronizing the position extracted from the AGV with probabilistic information.

[0029] RAG (Retriever-Augmented Generation) is a natural language processing (NLP) technique that combines a retrieval component (search component) with a generative component (generation component) that can generate human-like text based on the retrieved information. This approach allows the model to incorporate external knowledge related to the current context and query, thereby improving the quality and relevance of the generated response. In practical applications, RAG can be used to answer complex questions, generate detailed explanations, or create content by accessing and synthesizing information from diverse sources. For example, given a specific question, a RAG system can search a document database to find relevant information and use that information to construct a coherent and helpful answer.

[0030] Figure 2 shows an interactive time series analysis system tuned to analyze the mean time to detection (MTTD) of a specified process, according to an embodiment. The example in Figure 2 has a specified process referred to as the "ABC" process, specifically during May. The system consists of three main components: a time series analysis component 100, a data communication component 200, and a large-capacity data storage 300.

[0031] The user prompt (1) is the starting point. If the initial data input is insufficient, the system can request additional information via the LLM-based user interface (UI) 101. This interactive Q&A (if necessary) ensures that the system obtains all the information necessary to proceed with the analysis. The UI 101 queries the high-capacity data storage 300 for relevant video data (3) related to the ABC process. The high-capacity data storage 300 stores and manages multiple videos. The query (2) is input to the video data storage 301 via the data communication component 200, facilitating the transfer of any data from the high-capacity data storage 300 to the analysis component 100.

[0032] The video frame extraction unit 102 divides the video data into individual image frames (4). These frames, along with the MTTD labels (5), enter the probability calculation unit 103 for each event. Here, the probability (6) of each event (occurrence) is determined. In other words, the probability of at least one object in each frame is determined. In the case of MTTD analysis, the labels (5) for MTTD could be a red (or green) signal and a worker responding to the problem.

[0033] The system further incorporates a time-series probability storage unit 104 that stores time-series probability strings from the past to the present. These time-series probability strings stored in the time-series probability storage unit 104 are combined with context strings (7) and processed by the context information calculation unit 105 to create comprehensive information encompassing both probability and contextual nuances. For example, it can identify critical moments such as sharp peaks and troughs in event probability, or identify worker response times, making it essential contextual data for MTTD (Mean Time To Date) evaluation. This information is stored in the context information calculation and storage unit 106.

[0034] Next, the LLM-based analysis unit 107 uses the rich contextual information (8) stored in the contextual information calculation unit / storage unit 106, along with the initial user prompt (9) containing related information, to perform a detailed time-series analysis. This analysis may generate analytical data such as MTTD-related insights (10).

[0035] Ultimately, the LLM-based UI101 adopts and visualizes the analysis results and generates dynamic interactive responses (11) to the user. This may include interactive feedback such as clarifying the significance of signal colors in the operational context or explaining MTTD metrics within the system. Furthermore, the system's user-friendly interface allows for easy input and interpretation of complex time-series data, thereby assisting in the optimization of decision-making processes related to the ABC process. Although not described here, in addition to probabilistic information, external data such as programmable logic controllers (PLCs) can also be used as input to the system.

[0036] Figure 3A shows a sequence diagram relating to the system described herein, following an exemplary implementation. Externally referenced sections ("ref") describe preprocessing to add information required later in the process to ambiguous user prompts.

[0037] In the example flow shown in Figure 3A, the user first provides a user prompt (1) to the UI 101. The UI 101 can then run a Q&A (question and answer) session and generate queries to gather further information about the provided prompt. The query (2), in this example, is a video related to the ABC process and is sent to the video data storage 301. The related video (3) is retrieved from the video data storage 301 and processed by the video frame extraction unit 102 to extract frames (4). Each extracted frame (4) is processed by the event probability calculation unit 103 to calculate the probability information of at least one object in each frame. The event probability calculation unit 103 is also input with a label for the MTTD (5) generated by the UI 101. The frames and labels are processed by the event probability calculation unit 103 to determine the probability of each event. This process is repeated for each frame.

[0038] The probability of each event is provided to a context information calculation unit 105 configured to determine the indexed probability of the time series event (7). The indexed probability of the time series event is processed by the context information calculation unit 105 to generate context information, which is stored in the context information calculation and storage unit 106 and input for processing by the LLM-based analysis unit 107 (8).

[0039] The LLM-based analysis unit 107 is configured to take relevant information (9) from the UI 101 and context information (8) from the context information calculation and storage unit 106, similar to the user prompt, and return the analyzed data (10). In this example, the relevant information (9) included in the user prompt is: "Green light indicates normal operation, and red light indicates an abnormal event. MTTD indicates the average time it took the worker to discover the problem." The LLM-based analysis unit 107 returns the analyzed data (10) to the UI 101, and this data is visualized (11) and provided to the user from the UI 101.

[0040] Figure 3B shows an example of a question and answer tree for obtaining relevant information according to an embodiment. Figure 4 shows an example of preprocessing according to an embodiment. In the example in Figure 4, multiple questions are asked to add information necessary for each subsequent processing unit to the information contained in the user prompt of the LLM-based UI. In this example, questions #1 to #5 as shown in Figure 3B are asked to augment the RAG system to realize the relevant information shown in the fourth column of Figure 3B. The UI may ask the user again in a predetermined fixed format if there is an unexpected prompt, but this disclosure is not limited thereto, and other implementations may be used to facilitate the desired implementation.

[0041] As shown in Figure 4, a user prompt (step 400, hereafter steps will simply be referred to as S) is provided, which in this example is "Analyze the MTTD of the ABC process in May". In S401, the preprocessing shown in Figure 3B is executed, starting with question #1, "Does the user prompt contain 'analyze'?". If yes, question #2 is skipped; otherwise, the flow proceeds to S402 to ask the second question. In S402, question #2 is asked: "Does the user prompt contain 'get'?". If yes, the flow proceeds to S403; otherwise, the flow proceeds to S406.

[0042] In S403, question #3 is asked: "Does the user prompt include 'MTTD'?" If yes, the flow proceeds to S405; otherwise, the flow proceeds to S404. In S404, question #4 is asked: "Does the user prompt include 'SOP'?" If yes, the flow proceeds to S405; otherwise, the flow proceeds to S406.

[0043] In S405, question #5 is asked: "Does the user prompt include a specific process and a specific month?" If yes, the flow ends; otherwise, the flow proceeds to S406. In S406, the flow generates the following on the LLM-based UI: "Please ask again: (1) Analyze the SOP compliance of the AZ process (2) Get the video related to the NM process."

[0044] Figure 5 shows how the MTTD measurement is queried according to the embodiment. Specifically, Figure 5 shows the MTTD measurement queried by the user by detecting changes in probability information from the past to the present that exceed a predetermined threshold in the context information calculation and storage unit 106.

[0045] Figure 6 shows an example of a user interface (UI) for a system for interactive time series analysis according to an embodiment. As shown in Figure 6, the UI can also take user prompts (1) and display relevant video data (3), the probability of each event (7), and interactive responses (11) using an LLM-based UI 101.

[0046] Figure 7A shows an example of the execution of the context information calculation unit 105 according to the embodiment. The example in Figure 7A is an execution example in the case of MTTD, where a red light turns on at the exact center time k of Frame-k-1, Frame-k, and Frame-k+1, the operator confirms that the red light turns on at the center time m of Frame-m-1, Frame-m, and Frame-m+1, and the red light turns off and changes to a green light at the center time n of Frame-n-1, Frame-n, and Frame-n+1. Figure 7B shows the probability column when this scene is converted into text by each event probability calculation unit 103, and Figures 7C and 7D show examples of calculations by the context information calculation unit 105. Figure 7C shows the case where only frames with large probability changes are extracted, and Figure 7D shows the case where the change in probability from the previous frame is calculated and displayed.

[0047] Figure 8 shows another embodiment of an interactive time-series analysis system designed for inventory management within the XYZ process. This use case demonstrates how the system can detect probabilistic information for parts identified by the classification label "parts" and provide guidance on material delivery timing, as well as warnings of future stockouts, and other suggestions or action directives according to the desired implementation.

[0048] The details of the role of each component in this use case are as follows:

[0049] User prompt (1): The user asks the system, "By when should I provide the parts to the XYZ station?" This input initiates the analysis process.

[0050] LLM-based UI101: A user interface for a system driven by a Large-Scale Language Model (LLM) interprets user prompts and determines whether additional information is needed. It can then conduct Q&A as needed to clarify or expand on user requirements.

[0051] Related video query (2): The UI sends a query to the video frame extraction unit 102 to retrieve video data related to the relevant part from the large-capacity data storage 300.

[0052] High-capacity data storage 300: High-capacity data storage 300 stores a wide range of video data, including video over time from the XYZ process.

[0053] Video frame extraction unit 102: Once the associated video data storage 301 is identified, the video frame extraction unit 102 extracts frames from the video for analysis.

[0054] Frame and event probability calculation and time series probability storage unit: The probability calculation unit 103 for each event determines the probability of each event based on the individual frame (4) and the component labels (5) provided by the UI, and inputs a time series probability string (6) to the time series probability storage unit 104. The time series probability string (6) is processed by the time series probability storage unit 104 to calculate the probability (7) of each event and the probability of the time series event (time series probability).

[0055] Context information calculation unit 105 and context information storage unit 106: The calculated probabilities and labels are output in combination with the context information string (8), allowing for a comprehensive understanding of the event within its operational context.

[0056] LLM-based analysis unit 107: The LLM processes all of the above information and performs a detailed analysis (9). This analysis may include approximation formulas for the parts in terms of probability over time, visualized in a graph.

[0057] Interactive responses (10 and 11): Based on the analysis results, the LLM-based UI provides interactive responses to the user. For example, "You should aim to deliver the parts later than frame 10."

[0058] Each component works in coordination to enable the system to not only detect current inventory levels but also predict future needs, allowing for effective inventory management and optimization in the XYZ process. The system's ability to process and analyze video data through integrated frame extraction, event probability calculation, and contextual analysis, culminating in LLM-based predictive responses, exemplifies a cutting-edge approach to parts and materials management in industrial environments.

[0059] The implementation examples described herein are based on the seamless integration of CLIP and LLM technologies for interactive video analysis, as described. The system's ability to analyze video data with CLIP by identifying relevant objects and events, and the detailed process of passing information to LLM to generate contextually rich conversational responses, provides a robust foundation. The unique interactive and insightful analytical tools offered by the present invention are highlighted by the emphasis on the innovative integration of visual data analysis with natural language processing and search augmentation. The addition of preprocessing for data quality improvement, RAG for complex query processing, and feedback mechanisms for model refinement further enhances the system's capabilities for real-time monitoring, predictive maintenance, SOP compliance, and defect detection.

[0060] Possible implementation examples include predictive use cases based on multiple datasets, such as data from CLIP (keyword search or similar image search) and data from PLC. This is effective for issues that cannot be resolved with image information alone.

[0061] CLIP allows the use of long sentences rather than single words for label selection, so by using the RAG system, a "worker" organizing packages in an image can be labeled as "worker organizing packages." By modifying the flowchart in Figure 4, labeling can be optimized semi-automatically. It is possible to refer to the work procedures registered in the PLM from the process name and label based on the work procedures written in the work procedures.

[0062] Figure 9 shows an exemplary computing environment having exemplary computer devices suitable for use in several exemplary implementations, such as a system for interactive time-series analysis and a database for managing multiple videos. The computer device 905 of the computing environment 900 may include one or more processing units, cores, or processors (processing units) 910, memory 915 (e.g., RAM, ROM, and / or similar), internal storage 920 (e.g., magnetic, optical, solid-state storage, and / or organic), and / or I / O interfaces 925, any of which may be coupled on a communication mechanism or bus 930 for communicating information or embedded in the computer device 905. The I / O interface 925 may also be configured, depending on the desired implementation, to receive images from a camera or provide images to a projector or display.

[0063] Computer device 905 can be communicatively coupled to an input / user interface 935 and an output device / interface 940. Either or both of the input / user interface 935 and the output device / interface 940 can be wired or wireless interfaces and can be detachable. The input / user interface 935 may include any physical or virtual device, component, sensor, or interface that can be used to provide input (e.g., buttons, touchscreen interfaces, keyboards, pointing / cursor controls, microphones, cameras, Braille, motion sensors, optical readers, and / or similar). The output device / interface 940 may include displays, televisions, monitors, printers, speakers, Braille, etc. In some exemplary implementations, the input / user interface 935 and the output device / interface 940 may be embedded in or physically coupled to the computer device 905. In other exemplary implementations, other computer devices may function as or provide the functions of the input / user interface 935 and the output device / interface 940 of the computer device 905.

[0064] Examples of computer devices 905 include, but are not limited to, highly mobile devices (e.g., smartphones, devices mounted on vehicles and other machines, devices carried by humans and animals), mobile devices (e.g., tablets, notebooks, laptops, personal computers, portable televisions, radios, etc.), and devices not designed for mobility (e.g., desktop computers, other computers, information kiosks, televisions, radios, etc., having one or more processors embedded therein and / or coupled thereto).

[0065] Computer device 905 may be communicatively coupled (for example, via I / O interface 925) to external storage 945 and network 950 for communication with any number of network-connected components, devices, and systems, including one or more computer devices of the same or different configurations. Computer device 905 or any connected computer device may function, provide, or be referred to as a server, client, thin server, general machine, special-purpose machine, or another label.

[0066] The I / O interface 925 may include, but is not limited to, wired and / or wireless interfaces using any communication or I / O protocol or standard (e.g., Ethernet, 802.11x, Universal System Bus, WiMAX, modem, cellular network protocol, etc.) for communicating information to and from at least all connected components, devices, and networks within the computing environment 900. The network 950 may be any network or combination of networks (e.g., the Internet, local area network, wide area network, telephone network, cellular network, satellite network, etc.).

[0067] Computer device 905 may use and / or communicate using computer-usable media or computer-readable media, including transient media and non-transient media. Transient media include transmission media (e.g., metal cables, optical fibers), signals, carrier waves, etc. Non-transient media include magnetic media (disks, tapes, etc.), optical media (CD-ROMs, digital video discs, Blu-ray discs, etc.), solid-state media (RAM, ROMs, flash memory, solid-state storage, etc.), and other non-volatile storage or memory.

[0068] Computer device 905 can be used to implement technologies, methods, applications, processes, or computer executable instructions in several exemplary computing environments. Computer executable instructions can be retrieved from transient media and stored in and retrieved from non-transient media. Executable instructions can originate from one or more programming languages, scripting languages, and machine languages ​​(e.g., C, C++, C#, Java, Visual Basic, Python, Perl, JavaScript, etc. (Java is a registered trademark)).

[0069] The processor(s) 910 can run under any operating system (OS) (not shown) in a native or virtual environment. One or more applications can be deployed, including logical units 960, application programming interface (API) units 965, input units 970, output units 975, and inter-unit communication mechanisms 995 for different units to communicate with each other, with the OS, and with other applications (not shown). The described units and elements may vary in design, function, configuration, or implementation, and are not limited to the description provided. The processor(s) 910 may take the form of a hardware processor such as a central processing unit (CPU), or a combination of hardware and software units.

[0070] In some exemplary implementations, once information or execution instructions are received by the API unit 965, they may be transmitted to one or more other units (e.g., a logic unit 960, an input unit 970, and an output unit 975). In some embodiments, the logic unit 960 may be configured to control the flow of information between units and to direct the services provided by the API unit 965, the input unit 970, and the output unit 975, in some exemplary implementations described above. For example, one or more processes or implementation flows may be controlled by the logic unit 960 alone or in conjunction with the API unit 965. The input unit 970 may be configured to receive input for the computations described in the exemplary embodiments, and the output unit 975 may be configured to provide outputs based on the computations described in the exemplary embodiments.

[0071] The processor(s) 910 are configured to receive queries, and when a query such as the one shown in (2) is received, they can be configured to compute probability information for at least one object on each frame of a video from multiple videos related to the query, as shown in (3) to (6). Based on the probability information from the past up to a specified time, as shown in (7), the state of at least one object at a specified time is computed, and the state at the specified time is input into a large language model (LLM) configured to output analysis and predictions in natural language output in response to the query, as shown in (8) to (10).

[0072] The processor(s) 910 may be configured to calculate the state of at least one object at a given time by integrating past and present probability information, as described with respect to Figures 1 and 2.

[0073] Depending on the desired implementation, the LLM can be configured to generate an interactive response based on the input of probabilistic information and the state of at least one object at a given time (11).

[0074] The processor(s) 910 can be configured to compute the state at a specified time using a probabilistic model that incorporates the dynamic changes of at least one object.

[0075] As shown in Figure 5, the processor(s) 910 can be configured to facilitate the analysis and prediction of future events by calculating the state at a specified time based on predictions of future probability information and using the future probability information as input to the LLM.

[0076] Depending on the desired implementation, the LLM can be configured to dynamically adjust the response according to the context of the generated dialogue response and the user request seeking additional information, as shown in (9) to (11).

[0077] Depending on the desired implementation, the LLM can be configured to output predictions in natural language output based on future probability information as one or more warnings, suggestions, or action instructions, as shown in Figure 8.

[0078] The processor(s) 910(s) can be configured to optimize the label information through a preprocessing step before calculating the probability information, as shown in (5), thereby improving the accuracy of subsequent analysis and prediction.

[0079] Depending on the desired implementation, the LLM can be configured to perform a retriever-extension generation (RAG) based approach in response to input to integrate contextual information from an external knowledge base, as described herein.

[0080] The processor(s) 910 can be configured to perform a feedback mechanism to refine the model used to calculate probabilistic information for a specified time and the state of at least one object from the interaction with the user, as shown in (10) and (11).

[0081] Some parts of the detailed explanation are presented in terms of symbolic representations of algorithms and computer operations. These algorithmic descriptions and symbolic representations are means used by those skilled in the field of data processing technology to convey the essence of the innovation. An algorithm is a set of defined steps that lead to a desired final state or result. In the examples, the steps performed require a visible amount of physical operation to achieve the visible result.

[0082] Unless otherwise stated, as will be evident from the discussions, discussions throughout this specification using terms such as “processing,” “calculation,” “computation,” “determination,” and “display” may include the operations and processes of a computer system or other information processing device that manipulate and convert data represented as physical (electronic) quantities in the registers and memory of a computer system into other data similarly represented as physical quantities in the memory or registers of a computer system or other information storage, transmission, or display devices.

[0083] Exemplary embodiments also relate to apparatus for performing the operations described herein. This apparatus may be specifically configured for a particular purpose and may include one or more general-purpose computers that are selectively started or reconfigured by one or more computer programs. Such computer programs may be stored on computer-readable media such as computer-readable storage media or computer-readable signal media. Computer-readable storage media may include, but are not limited to, tangible media such as optical disks, magnetic disks, read-only memory, random-access memory, solid-state devices, and drives. Computer-readable signal media may include media such as carrier waves. The algorithms and representations presented herein are not inherently related to any particular computer or other apparatus. Computer programs may include purely software implementations containing instructions that perform the operation of a desired implementation.

[0084] Various general-purpose systems may be used with the programs and modules according to the embodiments herein, or it may be convenient to construct more specialized devices for performing the desired method steps. Furthermore, the embodiments are not described with reference to any particular programming language. It will be understood that various programming languages ​​may be used to carry out the teachings of the embodiments described herein. Instructions in a programming language may be executed by one or more processing units, such as a central processing unit (CPU), a processor, or a controller.

[0085] As is known in the art, the operations described above can be performed by hardware, software, or any combination of software and hardware. Various embodiments of the exemplary implementations may be implemented using circuit and logic devices (hardware), while other embodiments, when performed by a processor, may be implemented using instructions stored on a machine-readable medium containing software for causing the processor to perform the method of performing the implementation of this application. Furthermore, some exemplary implementations of this application may be performed by hardware alone, while other exemplary implementations may be performed by software alone. Moreover, the various functions described may be performed by a single unit or may span a number of components in any number of ways. When performed by software, the method may be performed by a processor such as a general-purpose computer based on instructions stored on a computer-readable medium. If desired, the instructions may be stored on the medium in a compressed and / or encrypted format.

[0086] Furthermore, other embodiments of the present application will be apparent to those skilled in the art from the considerations of the present specification and the practice of the teachings of the present application. Various aspects and / or components of the exemplary embodiments described herein can be used individually or in any combination. This specification and the exemplary embodiments are intended to be considered illustrative only, and the true scope and spirit of the present application are shown by the following claims. [Explanation of symbols]

[0087] 100 Time Series Analysis Components 101 User Interface (UI) 102 Video frame extraction unit 103 CLIP (Event Probability Calculation Unit) 104 Time-series probability memory unit 105 Context information calculation unit 106 Context information calculation / storage unit 107 LLM-Based Analysis Department 200 Data Communication Components 300 Large-capacity data storage 301 Video Data Storage

Claims

1. A database that manages multiple videos, It has a processing unit configured to receive queries, The aforementioned processing unit, Calculate probability information for at least one object in each frame of the video relevant to the query from multiple videos. Based on probability information from the past up to a specified time, calculate the state of at least one object at the specified time. A large-scale language model (LLM) is configured to output analysis and predictions in natural language output in response to queries, and to perform a Retriever-Augmented Generation (RAG) based approach in response to inputs to integrate contextual information from an external knowledge base, with the state at a specified time point being input. An interactive time series analysis system.

2. In the interactive time series analysis system described in claim 1, The processing unit is configured to calculate the state of at least one object at a specified time by integrating past and present probability information. An interactive time series analysis system.

3. In the interactive time series analysis system described in claim 1, The LLM generates an interactive response based on the input of the probability information and the state of the at least one object at the specified time. An interactive time series analysis system.

4. In the interactive time series analysis system described in claim 1, The aforementioned processing unit, The system is configured to calculate the state of the specified object at a given time using a probabilistic model that incorporates the dynamic changes of at least one of the objects. An interactive time series analysis system.

5. In the interactive time series analysis system described in claim 1, The aforementioned processing unit, By predicting future probability information, it calculates the state at a specified time and uses that future probability information as input to the LLM, thus facilitating the analysis and prediction of future events. An interactive time series analysis system.

6. In the interactive time series analysis system described in claim 1, The aforementioned LLM is, It is configured to dynamically adjust the response based on the context of the generated dialogue response and any additional information requests from the user. An interactive time series analysis system.

7. In the interactive time series analysis system described in claim 1, The aforementioned LLM is, It is configured to output predictions in natural language output based on future probability information as one or more of the following: warnings, suggestions, or action instructions. An interactive time series analysis system.

8. In the interactive time series analysis system described in claim 1, The aforementioned processing unit, The system is configured to optimize the label information by a preprocessing step before calculating the aforementioned probability information. An interactive time series analysis system.

9. In the interactive time series analysis system described in claim 1, The aforementioned processing unit, A feedback mechanism is implemented to improve the model used to calculate the probability information and the state of at least one object at a specified time based on user interaction. An interactive time series analysis system.

10. An interactive time series analysis method for analyzing an interactive time series analysis system having a database for managing multiple videos and a processing unit configured to receive queries, The processing unit performs the following: A step of calculating probability information for at least one object on each frame of a video from multiple videos related to a query, A step of calculating the state of at least one object at a specified time based on probability information from the past up to a specified time, The step of inputting a state at a specified time into a Large Language Model (LLM) configured to output analysis and predictions in natural language output in response to the aforementioned query, and configured to perform a Retriever-Augmented Generation (RAG) based approach in response to the input in order to integrate contextual information from an external knowledge base. Interactive time series analysis method.

11. In the interactive time series analysis method described in claim 10, The aforementioned processing unit, Calculating the state of at least one object at a given time integrates past and present probabilistic information. Interactive time series analysis method.

12. In the interactive time series analysis method described in claim 10, The aforementioned LLM is, Based on the input of the aforementioned probability information and the state of the at least one object at the specified time, an interactive response is generated. Interactive time series analysis method.

13. In the interactive time series analysis method described in claim 10, Calculating the state at the specified time uses a probabilistic model that incorporates the dynamic changes of at least one object. Interactive time series analysis method.

14. In the interactive time series analysis method described in claim 10, The calculation of the state at the specified time by the processing unit is performed based on predictions of future probability information, and since future probability information is used as input to the LLM, it facilitates the analysis and prediction of future events. Interactive time series analysis method.

15. In the interactive time series analysis method described in claim 10, The aforementioned LLM is, It is configured to dynamically adjust the response according to the context of the generated dialogue response and the user's request for additional information. Interactive time series analysis method.

16. In the interactive time series analysis method described in claim 10, The aforementioned LLM is, Based on future probability information, the system is configured to output predictions in the natural language output as one or more of the following: warnings, suggestions, or action instructions. Interactive time series analysis method.

17. In the interactive time series analysis method described in claim 10, The aforementioned processing unit, Before calculating probability information, the label information is optimized through a preprocessing procedure. Interactive time series analysis method.

18. In the interactive time series analysis method described in claim 10, The aforementioned processing unit, A feedback mechanism is implemented to improve the model used to calculate the probability information and the state of at least one object over a specified time period, based on user interaction. Interactive time series analysis method.

Citation Information

Patent Citations

  • Text creation system, text creation program, and report database production method

    JP2025127401A

  • JPP7409731B

  • Information analysis system, information analysis method, and program

    WO2024231989A1

  • Information processing device, information processing method, and recording medium

    WO2024261895A1