Information processing program, information processing method, and information processing device
The information processing program addresses AI chatbot hallucination by generating conditional documents and utilizing scene graphs to select appropriate documents, enhancing answer accuracy and reliability in AI chatbot responses.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-20
- Publication Date
- 2026-03-26
AI Technical Summary
Existing AI chatbot systems using large language models suffer from hallucination, where inaccurate or irrelevant outputs are generated, and the selection of optimal prompts for improving answer accuracy is not effectively managed, leading to suboptimal document generation.
An information processing program that utilizes a computer to acquire video footage, generate a conditional document, and select appropriate documents based on detected events and ground truth data, employing knowledge and action scene graphs to enhance answer accuracy.
Enables the generation of accurate answers to user questions by selecting optimal documents and patterns, thereby reducing hallucination and improving the reliability of AI chatbot responses.
Smart Images

Figure JP2024033767_26032026_PF_FP_ABST
Abstract
Description
Information Processing Program, Information Processing Method, and Information Processing Apparatus
[0001] The present invention relates to an information processing program and the like.
[0002] In recent years, AI chatbot services that answer user questions using AI (Artificial Intelligence) have been increasing. For example, in the prior art, there is a dialogue system that answers questions from users using large language models such as LLMs (Large Language Models).
[0003] Here, when generating an answer to a question using a large language model, a phenomenon (hallucination) may occur where outputs of content different from facts or content unrelated to the context are plausibly generated.
[0004] For example, in order to suppress hallucination, a technique called RAG (Retrieval Augmented Generation) is used. RAG improves the answer accuracy by combining external information search when performing text generation using a large language model.
[0005] Patent No. 7509972
[0006] However, the above-mentioned prior art has a problem that it cannot generate an accurate answer to a question.
[0007] For example, even when using RAG in the prior art, external information that can improve the answer accuracy cannot be set. Also, in a large language model, the answer obtained depends on how the instruction (prompt) is given, so it is required to input a more optimal one. However, if the prompt cannot be used properly, a document as intended cannot be obtained.
[0008] In one aspect, an object of the present invention is to provide an information processing program, an information processing method, and an information processing apparatus that can assist in selecting an appropriate document.
[0009] In the first proposal, the computer is instructed to perform the following processes: The computer acquires video footage, a first conditional document describing an event in the video, and ground truth data indicating the time or location where the event occurred. By inputting the first conditional document and a prompt indicating the document transformation into a large-scale language model, the computer generates a second conditional document in which the first conditional document has been transformed. Based on the generated second conditional document, the computer detects the location of the event in the video and selects a second conditional document based on the detected location of the event and the ground truth data.
[0010] It can assist in selecting the appropriate documents.
[0011] Figure 1 is a diagram illustrating the overall processing of the information processing device according to this embodiment 1. Figure 2 is a diagram illustrating an example of the data structure of the detection pattern and the matching pattern. Figure 3 is a diagram illustrating an example of a knowledge graph. Figure 4 is a diagram illustrating an example of an action scene graph. Figure 5 is a functional block diagram showing the configuration of the information processing device according to this embodiment 1. Figure 6 is a flowchart illustrating the processing procedure of the information processing device according to this embodiment 1. Figure 7 is a diagram illustrating the processing of the information processing device according to this embodiment 2. Figure 8 is a diagram illustrating the processing of the extraction unit. Figure 9A is a diagram (1) illustrating the processing of the candidate list generation unit. Figure 9B is a diagram (2) illustrating the processing of the candidate list generation unit. Figure 9C is a diagram (3) illustrating the processing of the candidate list generation unit. Figure 9D is a diagram (4) illustrating the processing of the candidate list generation unit. Figure 9E is a diagram (5) illustrating the processing of the candidate list generation unit. Figure 10 is a diagram supplementing the processing of the candidate list generation unit. Figure 11 is a diagram illustrating the processing of the video processing unit. Figure 12 is a diagram illustrating the processing of the evaluation unit. Figure 13A is a diagram illustrating the processing of the selection unit. Figure 13B is a diagram illustrating the processing of the selection unit. Figure 14 is a functional block diagram showing the configuration of the information processing device according to this embodiment 2. Figure 15 is a flowchart showing the processing procedure by the information processing device according to this embodiment 2 to generate a knowledge graph. Figure 16 is a flowchart showing the processing procedure by the information processing device to generate an answer when it receives a question. Figure 17 is a diagram (1) illustrating the processing of the information processing device according to this embodiment 3. Figure 18 is a diagram illustrating the processing of the template generation unit. Figure 19 is a diagram (2) illustrating the processing of the information processing device according to this embodiment 3. Figure 20A is a diagram (1) illustrating the processing during operation of the candidate list generation unit. Figure 20B is a diagram (2) illustrating the processing during operation of the candidate list generation unit. Figure 21 is a functional block diagram showing the configuration of the information processing device according to this embodiment 3. Figure 22 is a flowchart showing the pre-processing performed by the information processing device according to this embodiment 3. Figure 23 is a flowchart showing the processing during operation performed by the information processing device according to this embodiment 3.Figure 24 shows an example of a computer hardware configuration that achieves similar functions to the information processing device in the embodiment.
[0012] The following describes in detail, with reference to the drawings, embodiments of the information processing program, information processing method, and information processing apparatus disclosed in this application. However, this invention is not limited to these embodiments.
[0013] Figure 1 is a diagram illustrating the overall processing of the information processing device according to this embodiment 1. For example, the information processing device 100 of this embodiment 1 is a device that outputs an answer to a question sentence 11 related to a video 10 when it receives such a question sentence 11 from a user U1. The video 10 is a time-series frame (still image).
[0014] The information processing device 100 performs KG generation processing, ASG generation processing, and graph analysis processing. For example, the KG generation processing and ASG generation processing are performed in advance. The graph analysis processing is performed to generate an answer when a question 11 is received from the user. In the following description, the KG generation processing, ASG generation processing, and graph analysis processing will be described in order.
[0015] The KG generation process performed by the information processing device 100 will now be described. The KG generation process is the process of generating a Knowledge Graph 50 that shows the conditions for detecting a certain event in the video 10. For example, the Knowledge Graph 50 is a graph corresponding to detection patterns and matching patterns.
[0016] For example, the information processing device 100 acquires text 12 related to the domain of the detected object included in the video 10. The text 12 is such as "dangerous behavior with accident risk". The information processing device 100 generates a list of detected objects from the text 12 using LLM (Large Language Models) or the like. The list of detected objects is such as "approaching a moving forklift without wearing a vest", "carrying a load for a long time", and "entering the road without checking left and right".
[0017] The information processing device 100 generates multiple candidate detection patterns and matching patterns by setting a list of detection targets as a prompt for generating detection patterns and matching patterns and inputting it into the LLM.
[0018] Figure 2 shows an example of the data structure for detection patterns and matching patterns. The example shown in Figure 2 includes detection patterns 5-1, 5-2, 5-3 and matching pattern 5-4. Detection patterns 5-1 to 5-3 each define the conditions for the object to be detected. Detection pattern 5-1 defines "Subject," "Object," and "Relationship." For example, detection pattern 5-1 shows a relationship where a person corresponding to "Subject" approaches a forklift corresponding to "Object." "Relationship" is an example of interaction information.
[0019] Detection patterns 5-2 and 5-3 define "Subject" and "Attribute." For example, detection pattern 5-2 indicates that the person corresponding to "Subject" is wearing a vest. "Attribute" is an example of attribute information.
[0020] Matching pattern 5-4 further defines the conditions for the matching target for each detection target that matches the conditions of detection patterns 5-1, 5-2, and 5-3. For example, matching pattern 5-4 defines "Detection target" and "Pattern". "Pattern" defines a pattern in which a person is approaching a forklift and the forklift is moving. In such a "Pattern", whether or not the person is approaching the forklift is determined based on detection pattern 5-1. Whether or not the forklift is moving is determined based on detection pattern 5-3. In addition, as defined in detection pattern 5-2, information that the target person is wearing a vest may be further set in "Pattern".
[0021] If video 10 matches the "Pattern" of matching pattern 5-4, it is determined that the matching conditions shown in "Detection target" are met.
[0022] The information processing device 100 evaluates multiple candidate detection patterns and matching patterns, and selects the optimal detection pattern and matching pattern based on the evaluation results. The information processing device 100 generates a knowledge graph 50 based on the selected detection pattern and matching pattern.
[0023] Figure 3 shows an example of a knowledge graph. For example, the knowledge graph 50 shown in Figure 3 is generated based on detection patterns 5-1 to 5-3 and matching pattern 5-4. The knowledge graph 50 includes nodes n1-1, n1-2, n1-3, n1-4, and n1-5. Node n1-1 is the node corresponding to "Subject is wearing a vest". Node n1-2 is the node corresponding to Person. An arrow is set from node n1-1 to node n1-2, indicating that the Subject of node n1-1 is defined in node n1-2.
[0024] Nodes n1-3 correspond to the "Subject is moving" node. Node n1-2 corresponds to the forklift node. An arrow is set from node n1-3 to node n1-4, indicating that the Subject of node n1-3 is defined in node n1-4.
[0025] Nodes n1-5 are nodes corresponding to "Subject is approaching Object". An arrow is set from node n1-5 to node n1-2, indicating that the Subject of node n1-5 is defined in node n1-2. An arrow is set from node n1-5 to node n1-4, indicating that the Object of node n1-5 is defined in node n1-4. Note that the knowledge graph 50 may be generated from detection patterns only. In that case, the knowledge graph 50 may be represented using the data structures 5-1 to 5-3. Furthermore, when the knowledge graph 50 is generated from both detection patterns and matching patterns, it may be represented using the data structures 5-1 to 5-4.
[0026] The KG generation process performed by the information processing device 100 has been described above.
[0027] Returning to the explanation of Figure 1, the ASG generation process performed by the information processing device 100 will be described. The ASG generation process is a process that generates an Action Scene Graph 60 from the video 10 using the detection patterns of the knowledge graph 50. The ASG is also called a Video Scene Graph or Spatio-temporal scene graph.
[0028] For example, the information processing device 100 performs object detection using a detection pattern on time-series frames of the video 10 and tracks the detected objects. The information processing device 100 generates video clips by summarizing the detection results and tracking results for a predetermined number of frames. The information processing device 100 inputs the video clips and prompts for relationship and attribute detection generated from the detection patterns into a visual detection model such as a Vision Language Model (VLM), thereby identifying the attribute information of the detected objects contained in the video clips, interaction information between detected objects, and the time when the attribute information and interaction information occurred.
[0029] The information processing device 100 generates an action scene graph 60 based on a video clip, attribute information of the detection target identified from the video clip, interaction information between the detection targets, and time. The action scene graph 60 maintains the relationship between Subject, object, and relation, or the relationship between Subject, object, and attribute, on an event basis (attribute information, relation information).
[0030] Figure 4 shows an example of an action scene graph. As shown in Figure 4, the action scene graph 60 has time nodes n2-1, n2-2, n2-3, n2-4, n2-5, and n2-6. The action scene graph 60 has event nodes n3-1, n3-2, n3-3, n3-4, n3-5, and n3-6. The action scene graph 60 has concrete object nodes n4-1, n4-2, n4-3, n4-4, and n4-5.
[0031] Time nodes n2-1 to n2-6 are nodes that indicate time, and correspond to times T1, T2, T3, T4, T5, and T6, respectively. For example, times T1, T2, T3, T4, T5, and T6 are associated with the time (frame number) of each frame contained in the video clip.
[0032] Event nodes n3-1 to n3-6 are nodes corresponding to attribute information and interaction information. For example, event nodes n3-1 to n3-3 correspond to "wearing a vest". Event nodes n3-4 and n3-6 correspond to "moving". Event node n3-5 corresponds to "approaching".
[0033] Specific object nodes n4-1 to n4-5 are nodes corresponding to the detection target. For example, specific object nodes n4-1 to n4-4 correspond to people P1, P2, P3, and P4, respectively. Specific object node 4-5 corresponds to a forklift.
[0034] By using the action scene graph 60, it becomes possible to grasp various information about the video 10. For example, event node n3-1, which is connected to time nodes n2-1 and n2-6, is connected to concrete object node n4-2. This indicates that person P2, who is wearing a vest, is present in the video 10 during times T1 to T6.
[0035] The event node n3-2, connected to time nodes n2-1 and n2-6, is connected to the concrete object node n4-3. This indicates that person P3, wearing a vest, is present in video 10 during times T1 to T6.
[0036] The event node n3-3, connected to time nodes n2-1 and n2-6, is connected to the concrete object node n4-4. This indicates that person P4, wearing a vest, is present in video 10 during times T1 to T6.
[0037] Event nodes 3-4, connected to time nodes n2-1 and n2-3, are connected to object node n4-5. This indicates that the moving forklift is present in video 10 during times T1 to T3.
[0038] The event node n3-5, connected to time nodes n2-2 and n2-3, is connected to concrete object nodes n4-1 and n4-5. This indicates that the event of person P1 approaching a moving forklift occurred in time T2-T3 of video 10.
[0039] Event node n3-6, connected to time nodes n2-5 and n2-6, is connected to concrete object n4-5. This indicates that the moving forklift is present in video 10 at times T5-T6.
[0040] The above describes the ASG generation process performed by the information processing device 100.
[0041] Returning to the explanation of Figure 1, the graph analysis process performed by the information processing device 100 will be described. The graph analysis process is a process that, when a question sentence 11 related to the video 10 is received from user U1, uses LLM to analyze the behavior scene graph 60 and generates an answer.
[0042] For example, the information processing apparatus 100 generates a search query based on the question sentence 11 and the knowledge graph 50, and performs a data search on the action scene graph 60 using such a search query. The information processing apparatus 100 generates an answer using the result of the data search.
[0043] Next, a configuration example of the information processing apparatus 100 according to the first embodiment will be described. FIG. 5 is a functional block diagram showing the configuration of the information processing apparatus according to the first embodiment. As shown in FIG. 5, the information processing apparatus 100 includes a communication unit 110, an input unit 120, a display unit 130, a storage unit 140, and a control unit 150.
[0044] The communication unit 110 executes data communication with a camera via a network. For example, the communication unit 110 receives video data from the camera. The video data is the data of the video 10 described in FIG. 1.
[0045] Further, the communication unit 110 executes data communication with an external server via a network. For example, the communication unit 110 receives text data from the external server. The text data is the data of the text 12 described in FIG. 1.
[0046] The input unit 120 is an input device that inputs various types of information to the control unit 150 of the information processing apparatus 100. The user U1 may operate the input unit 120 to input the question sentence 11.
[0047] The display unit 130 is a display device that displays information output from the control unit 150.
[0048] The storage unit 140 includes a knowledge graph 50, an action scene graph 60, a video buffer 70, and a text table 80. The storage unit 140 is a memory or the like.
[0049] The knowledge graph 50 is graph data generated based on detection patterns and matching patterns. For example, the description of the knowledge graph 50 is the same as the description of the knowledge graph 50 in FIG. 3.
[0050] The behavioral scene graph 60 is graph data generated based on the knowledge graph 50 and video data. The explanation of the behavioral scene graph 60 is the same as the explanation of the behavioral scene graph 60 in Figure 4.
[0051] The video buffer 70 is a buffer for storing video data.
[0052] Text table 80 stores text data.
[0053] We will now move on to the explanation of the control unit 150. The control unit 150 includes an acquisition unit 151, a KG generation unit 152, an ASG generation unit 153, and a graph analysis unit 154. The control unit 150 is a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), etc.
[0054] The acquisition unit 151 acquires video data captured by the camera. The acquisition unit 151 stores the video data in the video buffer 70. The acquisition unit 151 also acquires text data from an external server. The acquisition unit 151 stores the text data in the text table 80.
[0055] The KG generation unit 152 generates a knowledge graph 50 by executing the above-described KG generation process. The KG generation unit 152 stores the knowledge graph 50 in the storage unit 140.
[0056] The ASG generation unit 153 generates an action scene graph 60 by executing the above ASG generation process. The ASG generation unit 153 stores the action scene graph 60 in the storage unit 140.
[0057] When the Graph analysis unit 154 receives the input of the question text 11, it generates an answer by executing the above-described Graph analysis process. The Graph analysis unit 154 outputs the answer to the display unit 130 for display.
[0058] Next, an example of the processing procedure of the information processing device 100 according to this embodiment 1 will be described. Figure 6 is a flowchart of the processing procedure of the information processing device according to this embodiment 1. As shown in Figure 6, the acquisition unit 151 of the information processing device 100 acquires text data and stores it in the text table 80 (step S10). The KG generation unit 152 of the information processing device 100 generates a knowledge graph 50 based on the text data stored in the text table 80 (step S11). The acquisition unit 151 acquires video data and stores it in the video buffer 70 (step S12).
[0059] The ASG generation unit 153 of the information processing device 100 generates an action scene graph 60 based on the video data stored in the video buffer 70 and the knowledge graph 50 (step S13).
[0060] The graph analysis unit 154 of the information processing device 100 receives the question 11 (step S14). The graph analysis unit 154 performs graph analysis and generates an answer (step S15). The graph analysis unit 154 outputs the answer (step S16).
[0061] Next, the effects of the information processing device 100 according to this embodiment 1 will be described. The information processing device 100 generates a knowledge graph 50 through a KG generation process and generates an action scene graph 60 through an ASG generation process. When the information processing device 100 receives a question 11 from user U1, it generates an answer based on the knowledge graph 50 and the action scene graph 60. This makes it possible to generate an accurate answer to the question.
[0062] Before describing this second embodiment, let's briefly explain the problems that this second embodiment of information processing device aims to solve. In the information processing device 100 described in Figure 1, the ASG generation process is executed using the knowledge graph 50 (detection pattern) to generate the behavior scene graph 60. Furthermore, when the information processing device 100 receives a question 11 from the user, it executes a graph analysis process to generate an answer. For example, in the graph analysis process, the structure of the behavior scene graph 60 is grasped and analyzed based on the detection pattern. Therefore, if an appropriate detection pattern is not set, it becomes impossible to generate an accurate answer to the question.
[0063] The information processing device of this embodiment 2 enables the generation of accurate answers to questions by selecting an appropriate detection pattern from a plurality of candidate detection patterns.
[0064] Next, an information processing device according to this second embodiment will be described. Figure 7 is a diagram illustrating the processing of the information processing device according to this second embodiment. As shown in Figure 7, this information processing device 200 has a KG generation unit 252. The KG generation unit 252 generates a knowledge graph 50.
[0065] For example, the KG generation unit 252 includes an extraction unit 261, a candidate list generation unit 262, an image recognition unit 263, an evaluation unit 264, and a selection unit 265. Each processing unit will be described in order below.
[0066] The extraction unit 261 generates a list of detection targets 261a based on the text 12 relating to the domain of the detection target, which is included in the video 31a.
[0067] Figure 8 is a diagram illustrating the processing of the extraction unit 261. More specifically, the extraction unit 261 creates a prompt 13 for generating a list of targets 261a based on the text 12, as well as text 12a indicating the domain to be detected, meeting minutes 12b of the results of interviews with experts on the target of detection, etc.
[0068] For example, area 13a of prompt 13 contains a summary of the task to be given to LLM 25. Area 13b of prompt 13 contains text 12a representing the domain to be detected. Area 13c of prompt 13 contains text 12 related to the domain to be detected, or meeting minutes 12b.
[0069] The extraction unit 261 generates a detection target list 261a by inputting prompt 13 to the LLM 25. The extraction unit 261 outputs the detection target list 261a to the candidate list generation unit 262. For example, the detection target list 261a includes multiple condition documents. These condition documents include "approaching a forklift in operation," "carrying a load for an extended period," and "entering the road without checking left and right." The condition documents included in the detection target list 261a are examples of "first condition documents." In the following explanation, the condition documents included in the detection target list 261a may be referred to as "first condition documents" as appropriate.
[0070] Next, the processing of the candidate list generation unit 262 in Figure 7 will be explained. The candidate list generation unit 262 generates a candidate list 262a based on the detection target list 261a. The processing of the candidate list generation unit will be explained using Figures 9A, 9B, 9C, 9D, and 9E.
[0071] First, let's explain Figure 9A. Figure 9A is a diagram (1) illustrating the processing of the candidate list generation unit. The candidate list generation unit 262 creates a prompt 26 by setting the detection target list 261a in a pre-prepared template 14. The prompt 26 is a prompt that instructs the LLM 25 to generate a detection pattern and a matching pattern from the condition documents of the detection target list 261a. An example of the prompt 26 is shown in Figure 9B.
[0072] For example, template 14 corresponds to a "prompt indicating document conversion." As described above, the detection target list 261a contains multiple conditional documents, and one prompt is generated from each conditional document. For convenience, multiple prompts generated from multiple conditional documents are collectively referred to as prompt 26.
[0073] The candidate list generation unit 262 generates a candidate list 26a by inputting a prompt 26 to the LLM 25. An example of the candidate list 26a is shown in Figure 9C. For example, the candidate list 26a includes pairs of detection patterns 26a-1, 26a-2, 26a-3 and matching pattern 26a-4.
[0074] For example, prompt 26 contains multiple prompts, and from a single prompt, multiple pairs of detection patterns 26a-1 to 26a-3 and matching pattern 26a-4 are generated. For example, a detection pattern is a pattern that individually defines the pattern to be detected, and a matching pattern is a pattern that combines each of the detection patterns. The explanation of the relationship between detection patterns and matching patterns is the same as that explained in Figure 2.
[0075] For convenience, the pairs (multiple pairs) of detection patterns 26a-1 to 26a-3 generated from multiple prompts and matching pattern 26a-4 are collectively referred to as the candidate list 26a.
[0076] We will now move on to the explanation of Figure 9D. Figure 9D is a diagram (4) illustrating the processing of the candidate list generation unit. The candidate list generation unit 262 improves accuracy by performing the processing shown in Figure 9B to transform the representation of the candidate list 26a. For example, the candidate list generation unit 262 transforms "Wearing vest," which indicates wearing a vest, into "Wearing yellow safety vest." The candidate list generation unit 262 transforms "Forklift," which indicates a forklift, into "Forklift which is one of trucks and has an arm to lift something."
[0077] The candidate list generation unit 262 creates a prompt 40a by setting the candidate list 26a, as described in Figure 9A, into a pre-prepared template 40. The prompt 40a is a prompt that instructs the LLM 25 to perform a conversion of the representation.
[0078] The candidate list generation unit 262 outputs the relationship between the expression before conversion and the expression after conversion by inputting prompt 40a to the LLM 25. The candidate list generation unit 262 generates candidate list 262a by merging the output result of LLM 25 and candidate list 26a. For example, the candidate list generation unit 262 may generate candidate list 262a by converting the expressions included in candidate list 26a according to the output result of LLM 25. An example of candidate list 262a is shown in Figure 9E.
[0079] For example, prompt 40a contains multiple prompts, and from a single prompt, multiple pairs of detection patterns 262a-1 to 262a-3 and matching pattern 262a-4 are generated. For example, a detection pattern is a pattern that individually defines the pattern to be detected, and a matching pattern is a pattern that combines each of the detection patterns. The explanation of the relationship between detection patterns and matching patterns is the same as that explained in Figure 2.
[0080] For example, the processing of the candidate list generation unit 262 is conceptually as shown in Figure 10. Figure 10 is a diagram to supplement the processing of the candidate list generation unit. The candidate list generation unit 262 generates detection patterns and matching patterns 27a, 27b, 27c, and 27d for a single condition document 27 by executing the processing described in Figures 9A and 9B. A single condition document 27 is "approaching a forklift in operation," etc.
[0081] The detection pattern and matching pattern 27a is <"person" (=Subject) is "approaching a forklift in operation" (=Attribute)>.
[0082] The detection pattern and matching pattern 27b is <"person" (=Subject) is "approaching" (=Relationship) a "forklift in operation" (=Object)>.
[0083] The detection pattern and matching pattern 27c is defined as <"person" (=Subject) is "approaching" (=Relationship) the "forklift" (=Object) and the "forklift" (=Subject) is "moving" (=Attribute)>.
[0084] The detection pattern and matching pattern 27c is defined as <"person" (=Subject) is "approaching" (=Relationship) a "forklift" (=Object) and the "forklift" (=Subject) is "carrying a person" (=Attribute)>.
[0085] For example, in Figure 10, Subject is the object to be detected. Relationship is interaction information. Attribute is attribute information.
[0086] The candidate list generation unit 262 generates various candidate detection patterns and matching patterns for detecting a target defined in a single condition document from video footage. For example, the candidate list generation unit 262 uses the LLM 25 to convert the condition document 27 into detection patterns and matching patterns 27a, 27b, 27c, and 27d. This set of detection patterns and matching patterns corresponds to a "second condition document." In the following description, the set of detection patterns and matching patterns generated by the candidate list generation unit 262 may be referred to as a second condition document as appropriate.
[0087] The candidate list generation unit 262 generates various detection pattern and matching pattern candidates by performing the above processing for each condition document in the detection target list 261a. The candidate list generation unit 262 sets a plurality of second condition documents for the first condition document in the candidate list 262a and outputs the candidate list 262a to the video recognition unit 263 and the selection unit 265.
[0088] Next, the processing of the image recognition unit 263, evaluation unit 264, and selection unit 265 shown in Figure 7 will be explained. The KG generation unit 252 uses the image recognition unit 263, evaluation unit 264, and selection unit 265 to select a second condition document capable of generating an appropriate answer from a plurality of second condition documents generated from a single first condition document. The KG generation unit 252 performs this processing for each of the plurality of second condition documents generated from each first condition document. For the convenience of explanation, in the following explanation, the first condition document will be the condition document 27 shown in Figure 10, and the second condition document will be explained using the detection pattern and matching pattern 27a, 27b, 27c, and 27d shown in Figure 10.
[0089] First, the video recognition unit 263 and the evaluation unit 264 perform the following processing using the detection pattern and the matching pattern 27a.
[0090] Figure 11 is a diagram illustrating the processing of the video processing unit. The video recognition unit 263 receives video 31a and candidate list 262a (for example, detection pattern and matching pattern 27a) as input and executes ASG generation processing and graph analysis processing.
[0091] First, an example of the ASG generation process performed by the video recognition unit 263 will be described. The ASG generation process is the process of generating the action scene graph 61.
[0092] For example, the video recognition unit 263 performs object detection for each frame of the video 31a. Object detection uses detection algorithms such as YOLO (You Only Look Once). The video recognition unit 263 also tracks the detected objects.
[0093] The video recognition unit 263 analyzes the object tracking results to identify the interactions between objects where the events of the detection pattern and matching pattern 27a occur (e.g., interaction information), the object attribute information, and the time when such events occurred. The video recognition unit 263 generates an action scene graph 61 by connecting the node of the object to be detected, the node of interaction information and attribute information, and the node of time with edges.
[0094] For example, the ASG generation process performed by the image recognition unit 263 corresponds to the ASG generation process described in Figure 1.
[0095] Next, an example of the graph analysis process performed by the video recognition unit 263 will be described. The graph analysis process generates the detection result 263a based on the question text. The question text is generated based on the candidate list 262a. For example, if the candidate list 262a contains the matching pattern "approaching a forklift in operation", the question text may be "output the time when the event of approaching a forklift in operation occurred" or "output the time when the event of approaching a forklift in operation occurred and the location of the person who caused the event".
[0096] The video recognition unit 263 generates a search query based on the question text and the candidate list 262a, and uses this search query to perform a data search on the action scene graph 61. The video recognition unit 263 generates a detection result 263a using the results of the data search. For example, the detection result 263a includes the position of the detected person or forklift for each frame. The video recognition unit 263 may also set the coordinates of the bounding box of the detected object, interaction information of the detected event, and attribute information for each frame. The video recognition unit 263 outputs the detection result 263a to the evaluation unit 264.
[0097] Next, the processing of the evaluation unit 264 will be explained. Figure 12 is a diagram illustrating the processing of the evaluation unit. The evaluation unit 264 acquires the detection result 263a and the correct answer data 35. The correct answer data 35 is correct answer data generated in advance by a user or the like who is referring to the video 31a, and sets the frame-by-frame position of the person and forklift, etc. that are to be detected. The correct answer data 35 contains at least information on which frame number the event occurred in.
[0098] The evaluation unit 264 calculates an evaluation score based on the degree of positional match between the detection result 263a and the correct data 35. For example, the evaluation unit 264 increases the evaluation score the smaller the difference in position between the detection result 263a and the correct data 35. The evaluation unit 264 generates an evaluation score 36a by associating the detection pattern and matching pattern 27a with the evaluation score.
[0099] Alternatively, the coordinates of the bounding box of the subject may be set in the correct answer data 35. In this case, the evaluation unit 264 may quantify the degree of agreement between the coordinates of the bounding box of the subject in the detection result and the coordinates of the bounding box of the subject in the correct answer data 35 using IoU (Intersection over Union) and generate an evaluation score.
[0100] Furthermore, the evaluation unit 264 may interpret that detection patterns and matching patterns that can be broken down into many elements, rather than just the degree of simple matching, are easier for the user to understand in terms of the mechanism of image recognition, and may use the number of attribute information and relationship information to determine the evaluation score. For example, the evaluation unit 264 may calculate the evaluation score based on equation (1). The weights in equation (1) are predetermined values, such as "0.1".
[0101] Evaluation score = Evaluation score based on degree of match + Weight × (Number of attribute information and relationship information) ... (1)
[0102] The process by which the video recognition unit 263 and the evaluation unit 264 generate an evaluation score 36a using the detection pattern and the matching pattern 27a has been described above.
[0103] Here, the video recognition unit 263 and the evaluation unit 264 also perform the above processing for the remaining detection patterns and matching patterns 27b to 27d for the condition document 27. As a result, evaluation score information for the detection patterns and matching patterns 27b, 27c, and 27d is generated. The evaluation unit 264 sets the evaluation score information for each of the detection patterns and matching patterns 27a to 27d into the evaluation list 264a and outputs it to the selection unit 265.
[0104] Next, the processing of the selection unit 265 in Figure 7 will be explained. Based on the evaluation score information of each detection pattern and matching pattern 27a to 27d set in the evaluation list 264a, the selection unit 265 selects the detection pattern and matching pattern that have the highest evaluation score.
[0105] Figures 13A and 13B are diagrams illustrating the processing of the selection unit. As shown in Figure 13A, for example, the selection unit 265 obtains a candidate list 262a and an evaluation list 264a. Based on the evaluation list 264a, the selection unit 265 selects the detection pattern and matching pattern that have the highest evaluation score and outputs the selected detection pattern and matching pattern 265a. Figure 13B shows an example of a matching pattern 265a.
[0106] Furthermore, a knowledge graph 50 is generated based on the detection pattern and matching pattern 265a selected by the selection unit 265. The selection unit 265 may also generate the knowledge graph 50. The explanation of the knowledge graph 50 is the same as the explanation of the knowledge graph 50 described in Figure 3. The selected detection pattern and matching pattern 265a may also be configured to allow the user to manually modify them. In addition, the candidate list 262a may be configured to include candidates manually created by the user or candidates created by modifying candidates included in the candidate list 262a, so that the selection includes candidates modified by the user.
[0107] Next, an example of the configuration of the information processing device 200 that performs the above-described process will be explained. Figure 14 is a functional block diagram showing the configuration of the information processing device according to this embodiment 2. As shown in Figure 14, this information processing device 200 has a communication unit 210, an input unit 220, a display unit 230, a storage unit 240, and a control unit 250.
[0108] The communication unit 210 performs data communication with the camera via the network. For example, the communication unit 210 receives video data from the camera. The video data includes the video data 10 described in Figure 1, the video data 31a described in Figure 7, and so on.
[0109] The input unit 220 is an input device that inputs various types of information to the control unit 250 of the information processing device 200.
[0110] The display unit 230 is a display device that displays information output from the control unit 250.
[0111] The memory unit 240 includes a knowledge graph 50, an action scene graph 60, and a video buffer 70. The memory unit 240 is a memory, etc.
[0112] The knowledge graph 50 is generated based on the detection pattern and matching pattern 265a selected by the KG generation unit 252. For example, the knowledge graph 50 corresponds to the knowledge graph 50 described in Figure 3.
[0113] The behavioral scene graph 60 is graph data generated based on the detection patterns and video data included in the knowledge graph 50. The explanation of the behavioral scene graph 60 is the same as the explanation of the behavioral scene graph 60 in Figure 4.
[0114] The video buffer 70 is a buffer for storing video data.
[0115] We will now move on to the explanation of the control unit 250. The control unit 250 includes an acquisition unit 251, a KG generation unit 252, an ASG generation unit 253, and a graph analysis unit 254. The control unit 250 is a CPU, GPU, etc.
[0116] The acquisition unit 251 acquires video data captured by the camera. The acquisition unit 251 stores the video data in the video buffer 70.
[0117] The KG generation unit 252 generates the knowledge graph 50 by executing the processes described in Figures 7 to 13.
[0118] For example, the KG generation unit 252 acquires video data, a first condition document indicating an event in the video, and correct data for the time or predetermined location where the event occurred. The KG generation unit 252 inputs the first condition document and a prompt indicating document conversion to the LLM 25, thereby generating a second condition document in which the first condition document has been converted. Based on the second condition document, the KG generation unit 252 detects the location of the event in the video data, selects a second condition document based on the detected event location and the correct data, and generates a knowledge graph 50 using the selection result.
[0119] The ASG generation unit 253 executes the ASG generation process. The ASG generation process generates an action scene graph 60 from video data using the detection pattern of the knowledge graph 50. The explanation of the ASG generation process is the same as the ASG generation process described in Example 1. In addition, the ASG generation unit 253 generates the action scene graph 60 in advance after the knowledge graph 50 has been generated by the KG generation unit 252 and registers it in the storage unit 240.
[0120] The graph analysis unit 254 executes graph analysis processing. When a question related to video data is received from user U1, the graph analysis processing uses LLM to analyze the behavior scene graph 60 and generate an answer. The explanation of the graph analysis processing is the same as the graph analysis processing described in Example 1.
[0121] Next, an example of the processing procedure of the information processing device 200 according to this second embodiment will be described. The process by which the KG generation unit 252 of the information processing device 200 generates the knowledge graph 50, and the process by which the information processing device 200 generates an answer when it receives a question, will be described in order.
[0122] Figure 15 is a flowchart showing the processing procedure for generating a knowledge graph by the information processing device according to this embodiment 2. As shown in Figure 15, the KG generation unit 252 of the information processing device 200 generates a list of detection targets 261a based on the texts 12, 12a and the minutes 12b (step S100).
[0123] The KG generation unit 252 generates a candidate list 262a based on the detection target list 261a (step S101). The KG generation unit 252 generates a detection result 263a by performing image recognition (step S102). The processing related to image recognition corresponds to the processing performed by the image recognition unit 263.
[0124] The KG generation unit 252 generates an evaluation list 264a based on the correct answer data 35 and the detection result 263a (step S103). The KG generation unit 252 selects a detection pattern and a matching pattern based on the evaluation list 264a (step S104). The KG generation unit 252 generates a knowledge graph 50 based on the selection result (step S105).
[0125] Next, we will explain the process by which the information processing device 200 generates an answer when it receives a question. Figure 16 is a flowchart showing the procedure for generating an answer when the information processing device receives a question.
[0126] The acquisition unit 251 of the information processing device 200 acquires video data and stores it in the video buffer 70 (step S200). The ASG generation unit 253 of the information processing device 200 generates an action scene graph 60 based on the video data stored in the video buffer 70 and the knowledge graph 50 (step S201).
[0127] The Graph analysis unit 254 of the information processing device 200 receives the question (step S202). The Graph analysis unit 254 performs graph analysis and generates an answer (step S203). The Graph analysis unit 254 outputs the answer (step S204).
[0128] Next, the effects of the information processing device 200 according to this second embodiment will be described. The information processing device 200 acquires video data, a first condition document indicating an event in the video, and correct data for the time or predetermined location where the event occurred. The information processing device 200 inputs the first condition document and a prompt indicating document conversion to the LLM 25, thereby generating a second condition document in which the first condition document has been converted. Based on the second condition document, the information processing device 200 detects the location of the event in the video data and selects a second condition document based on the detected location of the event and the correct data. As a result, by selecting an appropriate detection pattern and a matching pattern from a plurality of detection patterns and matching patterns, an appropriate knowledge graph 50 can be generated, thereby enabling the generation of an accurate answer to a question.
[0129] Next, an information processing device according to this embodiment 3 will be described. Figure 17 is a diagram (1) illustrating the processing of the information processing device according to this embodiment 3. As shown in Figure 17, the information processing device 300 has a KG generation unit 352. For example, the KG generation unit 352 performs pre-processing and operation-based processing, respectively. In the following, the pre-processing and operation-based processing performed by the KG generation unit 352 will be described in order.
[0130] First, the preprocessing performed by the KG generation unit 352 will be explained using Figure 17. As shown in Figure 17, the KG generation unit 352 includes an extraction unit 261, a candidate list generation unit 262, an image recognition unit 263, an evaluation unit 264, a selection unit 365, and a template generation unit 366.
[0131] The descriptions of the extraction unit 261, candidate list generation unit 262, image recognition unit 263, and evaluation unit 264 shown in Figure 17 are the same as those described in Figure 7. Therefore, the same reference numerals are used and their descriptions are omitted.
[0132] The selection unit 365 obtains the candidate list 262a and the evaluation list 264a. Based on the evaluation list 264a, the selection unit 365 selects the second condition document (detection pattern and matching pattern) that has the highest evaluation score. The selection unit 365 sets information relating the first condition document from the candidate list 262a to the selected second condition document in the selection list 365a.
[0133] The selection unit 365 repeatedly performs the above process each time it obtains an evaluation list 264a for the second condition document of another first condition document included in the candidate list 262a. As a result, pairs of the first condition document and the second condition document with the highest evaluation score are set in the selection list 365a. The selection unit 365 outputs the selection list 365a to the template generation unit 366.
[0134] The template generation unit 366 generates a generation prompt 366a based on the selection list 365a. For example, the template generation unit 366 sets the relationship between the first condition document and the second condition document set in the selection list 365a in the generation prompt 366a. The generation prompt 366a is a prompt that instructs the LLM to generate the second condition document from the newly specified first condition document according to the relationship between the first and second condition documents.
[0135] Figure 18 is a diagram illustrating the processing of the template generation unit. For example, the template generation unit 366 generates a generation prompt 366a by setting the relationship between a first condition document and a second condition document included in the selection list 365a at a predetermined position in the pre-prepared template 365b. The relationship between the first condition document and the second condition document is training data for in-context learning.
[0136] The generation prompt 366a includes areas 366a-1, 366a-2, 366a-3, 366a-4, 366a-5, and 366a-6. For example, areas 366a-1 and 366a-5 contain a summary of the task to be given to the LLM. Area 366a-2 contains constraints on the <subject, relationship, object> and <subject, attribute> to be generated.
[0137] Area 366a-3 is provided with an area for setting the target to be detected.
[0138] In area 366a-4, training data for in-context learning is set. In area 366a-6, an area is provided for setting the target for which the expression should be modified.
[0139] The generation prompt 366a generated by the template generation unit 366 is used in the operation process described later.
[0140] The above describes the pre-processing performed by the KG generation unit 352.
[0141] Next, the operational processing performed by the KG generation unit 352 will be explained using Figure 19. Figure 19 is a diagram (2) for explaining the processing of the information processing device according to this embodiment 3. In Figure 19, only the extraction unit 261 and the candidate list generation unit 262, which are closely related to the operational processing, are shown among the processing units of the KG generation unit 352, and the other processing units are omitted from the illustration.
[0142] The extraction unit 261 acquires the text 12 and generates a list of items to be detected 261a. The extraction unit 261 outputs the list of items to be detected 261a to the candidate list generation unit 262. The description of the extraction unit 261 is the same as the processing content described in Figure 8.
[0143] The candidate list generation unit 262 generates a detection pattern and a matching pattern 366b based on the detection target list 261a and the generation prompt 366a. The generation prompt 366a is the same generation prompt 366a described in Figure 17.
[0144] Figure 20A is a diagram (1) illustrating the operation of the candidate list generation unit. The candidate list generation unit 262 generates a prompt 44 by setting the detection target list 261a to the generation prompt 366a. The candidate list generation unit 262 generates a detection pattern and a matching pattern 45 by inputting the prompt 44 to the LLM 25.
[0145] Figure 20B is a diagram (2) illustrating the processing performed by the candidate list generation unit during operation. The candidate list generation unit 262 improves accuracy by performing the processing shown in Figure 20B to convert the representation of the detection pattern and matching pattern 45. For example, the candidate list generation unit 262 converts "Wearing vest," which indicates wearing a vest, to "Wearing yellow safety vest." The candidate list generation unit 262 converts "Forklift," which indicates a forklift, to "Forklift which is one of trucks and has an arm to lift something."
[0146] The candidate list generation unit 262 creates a prompt 46a by setting the detection pattern and matching pattern 45 described in Figure 20A into a pre-prepared template 46. The prompt 46a is a prompt that instructs the LLM 25 to perform a representation conversion.
[0147] The candidate list generation unit 262 outputs the relationship between the expression before conversion and the expression after conversion by inputting prompt 46a to the LLM 25. The candidate list generation unit 262 generates a detection pattern and a matching pattern 366b by merging the output result of the LLM 25 with the detection pattern and matching pattern 45. The detection pattern and matching pattern 366b includes detection patterns 366b-1, 366b-2, 366b-3 and matching pattern 366b-4.
[0148] For example, the candidate list generation unit 262 may generate the detection pattern and the matching pattern 366b by converting the expression containing the detection pattern and the matching pattern 45 according to the output result of the LLM 25.
[0149] Returning to the explanation of Figure 19, the ASG generation unit 353 executes the ASG generation process. The ASG generation process is a process that generates an action scene graph 60 from the video 31a using the detection pattern and the matching pattern 366b. The explanation of the ASG generation process is the same as the ASG generation process explained in Example 1.
[0150] The graph analysis unit 354 executes graph analysis processing. When a question sentence 31b related to the video 31a is received from user U1, the graph analysis processing uses LLM to analyze the behavior scene graph 60 and generates an answer 31c. The explanation of the graph analysis processing is the same as the graph analysis processing described in Example 1.
[0151] Next, an example of the configuration of the information processing device 300 that performs the above-described process will be explained. Figure 21 is a functional block diagram showing the configuration of the information processing device according to this embodiment 3. As shown in Figure 21, this information processing device 300 has a communication unit 310, an input unit 320, a display unit 330, a storage unit 340, and a control unit 350.
[0152] The communication unit 310 performs data communication with the camera via the network. For example, the communication unit 310 receives video data from the camera. The video data includes the video data 10 described in Figure 1, the video data 31a described in Figure 19, and so on.
[0153] The input unit 320 is an input device that inputs various types of information to the control unit 250 of the information processing device 200.
[0154] The display unit 330 is a display device that displays information output from the control unit 250.
[0155] The memory unit 340 has an action scene graph 60 and a video buffer 70. The memory unit 340 is a memory, etc.
[0156] The behavior scene graph 60 is graph data generated based on the detection pattern and matching pattern 366b shown in Figure 19 and the video data. The explanation of the behavior scene graph 60 is the same as the explanation of the behavior scene graph 60 in Figure 4.
[0157] The video buffer 70 is a buffer for storing video data.
[0158] We will now move on to the explanation of the control unit 350. The control unit 350 includes an acquisition unit 351, a KG generation unit 352, an ASG generation unit 353, and a graph analysis unit 354. The control unit 350 is a CPU, GPU, etc.
[0159] The KG generation unit 352 generates detection patterns and matching patterns 366b by executing the processes described in Figures 17 to 19, and outputs the detection patterns and matching patterns 366b to the ASG generation unit 353. The KG generation unit 352 may also generate a knowledge graph 50 based on the detection patterns and matching patterns 366b.
[0160] For example, in preprocessing, the KG generation unit 352 acquires video data, a first condition document indicating an event in the video, and correct data for the time or predetermined location where the event occurred. The KG generation unit 352 inputs the first condition document and a prompt indicating document conversion to the LLM 25, thereby generating a second condition document in which the first condition document has been converted. Based on the second condition document, the KG generation unit 352 detects the location of the event in the video data and generates a generation prompt 366a indicating document conversion based on the detected event location and the correct data.
[0161] More specifically, the KG generation unit 352 generates an evaluation list 264a based on the location of the detected event and the correct answer data, and selects the second condition document with the highest evaluation score from among the multiple second condition documents included in the evaluation list 264a. The KG generation unit 352 sets the pair of the first condition document and the second condition document with the highest evaluation score as the generation prompt 366a.
[0162] During operation, the KG generation unit 352 generates a detection pattern and a matching pattern 366b based on the detection target list 261a and the generation prompt 366a. The KG generation unit 352 outputs the detection pattern and the matching pattern 366b to the ASG generation unit 535.
[0163] The ASG generation unit 353 executes the ASG generation process. The ASG generation process generates an action scene graph 60 from video data using the detection pattern and the matching pattern 366b. The description of the ASG generation process is the same as the ASG generation process described in Example 1.
[0164] The graph analysis unit 354 executes graph analysis processing. When a question sentence 31b related to the video 31a is received from user U1, the graph analysis processing uses LLM to analyze the behavior scene graph 60 and generates an answer 31c. The explanation of the graph analysis processing is the same as the graph analysis processing described in Example 1.
[0165] Next, an example of the processing procedure of the information processing device 300 according to this embodiment 3 will be described. In the following, the pre-processing by the information processing device 300 and the processing during operation by the information processing device 200 will be described in order.
[0166] Figure 22 is a flowchart showing the preprocessing performed by the information processing device according to this embodiment 3. As shown in Figure 22, the KG generation unit 352 of the information processing device 300 generates a detection target list 261a based on the texts 12, 12a and the minutes 12b (step S300).
[0167] The KG generation unit 352 generates a candidate list 262a based on the detection target list 261a (step S301). The KG generation unit 352 generates a detection result 263a by performing image recognition (step S302). The processing related to image recognition corresponds to the processing performed by the image recognition unit 263.
[0168] The KG generation unit 352 generates an evaluation list 264a based on the correct answer data 35 and the detection result 263a (step S303). The KG generation unit 352 selects a detection pattern and a matching pattern based on the evaluation list 264a (step S304). The KG generation unit 352 generates a selection list 365a based on the selection result (step S305).
[0169] The KG generation unit 352 generates a generation prompt 366a based on the selection list 365a (step S306).
[0170] Figure 23 is a flowchart showing the operational processing performed by the information processing device according to this embodiment 3. As shown in Figure 23, the KG generation unit 352 of the information processing device 300 generates a detection target list 261a based on the texts 12, 12a and the minutes 12b (step S400).
[0171] The KG generation unit 352 generates a detection pattern and a matching pattern 366b based on the detection target list 261a and the generation prompt 366a (step S401).
[0172] The acquisition unit 351 of the information processing device 300 acquires video data and stores it in the video buffer 70 (step S402). The ASG generation unit 353 of the information processing device 300 generates an action scene graph 60 based on the video data stored in the video buffer 70, the detection pattern and the matching pattern 366b (step S403).
[0173] The Graph analysis unit 354 of the information processing device 300 receives the question (step S404). The Graph analysis unit 354 performs graph analysis and generates an answer (step S405). The Graph analysis unit 354 outputs the answer (step S406).
[0174] Next, the effects of the information processing device 300 according to this embodiment 3 will be described. In preprocessing, the information processing device 300 acquires video data, a first condition document indicating an event in the video, and correct data for the time or predetermined location where the event occurred. The information processing device 300 inputs the first condition document and a prompt indicating document conversion to the LLM 25, thereby generating a second condition document in which the first condition document has been converted. Based on the second condition document, the information processing device 300 detects the location of the event in the video data and generates a generation prompt 366a indicating document conversion based on the detected location of the event and the correct data.
[0175] Here, the generation prompt 366a is a prompt that instructs the LLM to generate a second condition document from a newly specified first condition document, according to the relationship between the first condition document and the second condition document.
[0176] The information processing device 300 can generate detection patterns and matching patterns 366b based on the detection target list 261a and the generation prompt 366a during operation, and by using these detection patterns and matching patterns 366b, it becomes possible to generate accurate answers to questions.
[0177] The information processing device 300 generates an evaluation list 264a based on the second condition document, the location of the event detected in the video, and the correct answer data. From the multiple second condition documents included in the evaluation list 264a, it selects the second condition document with the highest evaluation score. This allows the selection of an appropriate second condition document from the multiple second condition documents generated from the first condition document to be used as a detection pattern and a matching pattern.
[0178] Next, an example of a computer hardware configuration that realizes the same functions as the information processing device 100 (200, 300) shown in the above embodiment will be described in order.
[0179] Figure 24 shows an example of a computer hardware configuration that realizes similar functions to the information processing apparatus of the embodiment. As shown in Figure 24, the computer 400 has a CPU 401 that performs various calculations, an input device 402 that receives data input from the user, and a display 403. The computer 400 also has a communication device 404 that exchanges data with cameras, external devices, etc. via a wired or wireless network, and an interface device 405. The interface device 405 may have a microphone, speaker, etc. connected to it. The computer 400 also has a RAM 406 that temporarily stores various information and a hard disk drive 407. Each of the devices 401 to 407 is connected to a bus 408.
[0180] The hard disk drive 407 includes an acquisition program 407a, a KG generation program 407b, an ASG generation program 407c, and a graph analysis program 407d. The CPU 401 reads each of the programs 407a to 407d and loads them into the RAM 406.
[0181] The acquisition program 407a functions as the acquisition process 406a. The KG generation program 407b functions as the KG generation process 406b. The ASG generation program 407c functions as the ASG generation process 406c. The graph analysis program 407d functions as the graph analysis process 406d.
[0182] The processing in acquisition process 406a corresponds to the processing in acquisition units 151, 251, and 351. The processing in KG generation process 406b corresponds to the processing in KG generation units 152, 252, and 352. The processing in ASG generation process 406c corresponds to the processing in ASG generation units 153, 253, and 353. The processing in graph analysis process 406d corresponds to the processing in graph analysis units 154, 254, and 354.
[0183] Furthermore, programs 407a to 407d do not necessarily have to be stored in the hard disk drive 407 from the beginning. For example, each program can be stored in a "portable physical medium" such as a flexible disk (FD), CD-ROM, DVD, magneto-optical disk, or IC card inserted into the computer 400. Then, the computer 400 can read and execute each program 407a to 407d.
[0184] 50 Knowledge Graph 60 Behavioral Scene Graph 70 Video Buffer 80 Text Table 100, 200, 300 Information Processing Unit 110, 210, 310 Communication Unit 120, 220, 320 Input Unit 130, 230, 330 Display Unit 140, 240, 340 Storage Unit 150, 250, 350 Control Unit 151, 251, 351 Acquisition Unit 152, 252, 352 KG Generation Unit 153, 253, 353 ASG Generation Unit 154, 254, 354 Graph Analysis Unit 261 Extraction Unit 262 Candidate List Generation Unit 263 Video Recognition Unit 264 Evaluation Unit 265, 365 Selection Unit 366 Template Generation Unit
Claims
1. An information processing program characterized by: acquiring video footage, a first condition document indicating an event in the video footage, and ground truth data for the time or location where the event occurred; inputting the first condition document and a prompt indicating document conversion into a large-scale language model to generate a second condition document obtained by converting the first condition document; detecting the location of the event in the video footage based on the generated second condition document; and causing a computer to perform a process of selecting a second condition document based on the detected location of the event and the ground truth data.
2. The information processing program according to claim 1, characterized in that it causes the computer to further execute a process to generate a generation prompt that sets the relationship between the first condition document and the second condition document selected by the selection process, and an instruction to generate a new second condition document from a new first condition document based on the relationship.
3. The information processing program according to claim 2, characterized in that it causes the computer to further perform a process of obtaining the new first condition document and generating the new second condition document by inputting the obtained new first condition document and the generation prompt into the large-scale language model.
4. An information processing method characterized by the following: acquiring video footage, a first condition document indicating an event in the video footage, and ground truth data for the time or location where the event occurred; inputting the first condition document and a prompt indicating the conversion of the document into a large-scale language model to generate a second condition document obtained by converting the first condition document; detecting the location of the event in the video footage based on the generated second condition document; and selecting a second condition document based on the detected location of the event and the ground truth data.
5. The information processing method according to claim 4, characterized in that the computer further executes a process to generate a generation prompt that sets the relationship between the first condition document and the second condition document selected by the selection process, and an instruction to generate a new second condition document from a new first condition document based on the relationship.
6. The information processing method according to claim 5, characterized in that the computer further performs a process of obtaining the new first condition document and generating the new second condition document by inputting the obtained new first condition document and the generation prompt into the large-scale language model.
7. An information processing device having a control unit that acquires video footage, a first condition document indicating an event in the video footage, and ground truth data for the time or location where the event occurred; inputs the first condition document and a prompt indicating document conversion into a large-scale language model to generate a second condition document obtained by converting the first condition document; detects the location of the event in the video footage based on the generated second condition document; and selects a second condition document based on the detected location of the event and the ground truth data.
8. The information processing apparatus according to claim 7, wherein the control unit further executes a process to generate a generation prompt that sets a relationship between the first condition document and the second condition document selected by the selection process, and an instruction to generate a new second condition document from a new first condition document based on the relationship.
9. The information processing apparatus according to claim 8, characterized in that the control unit further executes a process of acquiring the new first condition document and generating the new second condition document by inputting the acquired new first condition document and the generation prompt to the large-scale language model.
Citation Information
Patent Citations
Program, method, information processing device, and system
JP7488617B1
Systems and methods for visual question answering using image relevant textual prompts
US20240119257A1