Generation program, generation method, and information processing device
The system accurately generates facility-specific measures by analyzing images and integrating graph data with large language models to address AI hallucinations and improve answer precision.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- FUJITSU LTD
- Filing Date
- 2024-10-30
- Publication Date
- 2026-05-07
AI Technical Summary
Existing AI chatbot systems using large language models struggle with accurately generating answers related to facility-specific measures due to hallucinations and lack of context-based analysis.
A system that uses a computer to analyze facility images, identify behavior types, generate graph data associating on-site characteristics with behavior, and input prompts to a language model to generate accurate facility measures.
Enables precise generation of facility-specific measures by integrating image analysis with large language models, addressing hallucinations and improving answer accuracy.
Smart Images

Figure JP2024038769_07052026_PF_FP_ABST
Abstract
Description
Generation Program, Generation Method, and Information Processing Apparatus
[0001] The present invention relates to a generation program, a generation method, and an information processing apparatus.
[0002] In recent years, AI (Artificial Intelligence) chatbot services that answer users' questions have been increasing. For example, there is a dialogue system that uses large language models such as LLMs (Large Language Models) to answer questions from users. Here, when generating an answer to a question using a large language model, a phenomenon (hallucination) may occur where outputs of content different from facts or content unrelated to the context are plausibly generated.
[0003] Japanese Patent No. 7509972
[0004] However, with the above technology, it is not possible to accurately generate an answer to a question, and it is difficult to accurately generate information regarding measures to be applied to a facility.
[0005] In one aspect, an object is to provide a generation program, a generation method, and an information processing apparatus that can accurately generate information regarding measures to be applied to a facility.
[0006] In the first aspect, the generation program causes a computer to acquire an image of the inside of a facility, analyze the acquired image to identify the type of behavior of a person staying in the facility, generate graph data in which information regarding the on-site characteristics of the facility is associated with the identified type of behavior of the person, and input a prompt including the search result of the generated graph data and a requirement regarding a measure to be applied in the facility to a language model, thereby causing the language model to execute a process of generating information regarding the measure as an answer to the requirement.
[0007] According to one embodiment, information regarding measures to be applied to a facility can be accurately generated.
[0008] Figure 1 is a diagram illustrating the overall configuration of the system according to Example 1. Figure 2 is a diagram illustrating the information processing device according to Example 1. Figure 3 is a diagram illustrating the overall processing of the information processing device according to Example 1. Figure 4 is a diagram illustrating an example of the data structure of detection patterns and matching patterns. Figure 5 is a diagram illustrating an example of a knowledge graph. Figure 6 is a diagram illustrating an example of an action scene graph. Figure 7 is a functional block diagram showing the functional configuration of the information processing device according to Example 1. Figure 8 is a flowchart illustrating the processing procedure of the information processing device according to Example 1. Figure 9 is a diagram illustrating the graph analysis unit according to Example 2 in detail. Figure 10A is a diagram illustrating the processing of the first generation unit. Figure 10B is a diagram illustrating the processing of the third generation unit. Figure 10C is a diagram illustrating the processing of the third generation unit. Figure 11 is a diagram illustrating an example of a scene graph. Figure 12 is a diagram illustrating an example of generating a scene graph showing the relationship between people and objects. Figure 13 is a diagram illustrating the identification of relationships using a scene graph. Figure 14 is a diagram illustrating an example of relationship identification using a scene graph. Figure 15 is a diagram illustrating an example of a hardware configuration.
[0009] The following describes in detail, with reference to the drawings, embodiments of the generation program, generation method, and information processing apparatus disclosed herein. However, the present invention is not limited by these embodiments. Each embodiment can be combined as appropriate within a non-consistent range.
[0010] <System Configuration> Figure 1 is a diagram illustrating the overall configuration of the system according to Embodiment 1. As shown in Figure 1, this system comprises a store 1 and an information processing device 100, and each device within the store 1 and the information processing device 100 are connected to each other via a network N so that they can communicate with one another. The network N can be any communication network, such as the internet or a dedicated line, regardless of whether it is wired or wireless.
[0011] Store 1 is an example of a facility including warehouses and factories, and is connected to multiple surveillance cameras (hereinafter sometimes simply referred to as cameras). For example, each surveillance camera is installed in areas such as the aisles of a user (an example of a consumer), product shelves where goods are displayed, and self-checkout counters where products are purchased, and outputs the captured video data (hereinafter sometimes simply referred to as video) to the information processing device 100. Self-checkout counters are also called, for example, Self checkout, automated checkout, self-checkout machine, or self-checkout register.
[0012] The information processing device 100 is an example of a computer that acquires images captured by each camera in the store 1, analyzes the images, and notifies the store manager or others of measures to take against shoplifting and other fraudulent or suspicious behavior.
[0013] In recent years, shoplifting has become a major factor in store sales losses. There is a need for methods to prevent shoplifting by detecting shoplifting behavior and suspicious activity, understanding shoplifting methods and trends, and taking measures tailored to the specific characteristics of each store. However, most shoplifting detection technologies only aggregate and visualize detection results, and are unable to analyze shoplifting cases based on detection results and on-site characteristics, or to propose crime prevention measures.
[0014] Therefore, the information processing device 100 presents shoplifting prevention measures based on the on-site characteristics of the store. Figure 2 is a diagram illustrating the information processing device 100 according to Embodiment 1. The information processing device 100 shown in Figure 2 acquires video footage of the inside of store 1 and analyzes the acquired video footage to identify the types of behavior of people staying inside store 1. The information processing device 100 generates graph data in which information about the on-site characteristics of store 1 is associated with the identified types of behavior of people. The information processing device 100 inputs a prompt to a language model that includes the search results of the generated graph data and a request regarding measures to be applied inside store 1, and generates information about measures as a response to the request. More specifically, for example, the information processing device 100 inputs a prompt to a language model that includes the search results of the generated graph data and a request regarding measures to be applied inside store 1, and generates information about measures as a response to the request.
[0015] For example, as shown in Figure 2, the information processing device 100 uses Large Language Models to perform the following actions based on data such as the results of shoplifting recognition by the video recognition module, store design, and product information. Specifically, the information processing device 100 stores the relationships between shoplifting methods, products, time of day, etc., in a graph database based on the recognition results and facts, and constructs on-site-based graph data. Then, the information processing device 100 searches for the on-site-based graph data related to shoplifting based on the individual whose shoplifting behavior was recognized by analyzing the on-site video, and the domain information of the site, such as product and inventory information and store design information. After that, the information processing device 100 inputs the search results and requests regarding measures as prompts to the LLM, and proposes shift designs and personnel allocation.
[0016] For example, in response to a request regarding a policy, "Tell us about measures to prevent shoplifting of product A," the information processing device 100 can autonomously perform hypothesis testing, such as "There is a high risk of shoplifting targeting these products and areas" or "There is a high risk of shoplifting if customers behave in this way," and then propose a policy such as, "If a customer is holding a container of personal belongings with the hand opposite (different) the hand holding the product, there is a high risk of shoplifting product A, so it would be better to station a monitor to observe their behavior."
[0017] <Processing of the Information Processing Device 100> Figure 3 is a diagram illustrating the overall processing of the information processing device according to Embodiment 1. For example, the information processing device 100 of Embodiment 1 is a device that outputs an answer to a question 11 when it receives a question 11 from a user U1 that includes a request for measures such as "Please tell me about measures to prevent shoplifting of product A." The video 10 is a time-series frame (still image).
[0018] The information processing device 100 performs KG generation processing, ASG generation processing, and graph analysis processing. For example, the KG generation processing and ASG generation processing are performed in advance. The graph analysis processing is performed to generate an answer when a question 11 is received from the user. In the following description, the KG generation processing, ASG generation processing, and graph analysis processing will be described in order.
[0019] (KG generation process) The KG generation process performed by the information processing device 100 will now be described. The KG generation process is the process of generating a Knowledge Graph 50 that shows the conditions for detecting a certain event in the video 10. For example, the Knowledge Graph 50 is a graph corresponding to detection patterns and matching patterns.
[0020] For example, the information processing device 100 acquires text 12 related to the domain of the detected object included in the video 10. The text 12 is such as "dangerous behavior with accident risk". The information processing device 100 generates a list of detected objects from the text 12 using LLM or the like. For example, taking store 1 as an example, the list of detected objects would include abnormal behaviors and fraudulent actions such as "putting goods in personal belongings", "making movements that prevent the self-checkout machine from scanning barcodes", "staying in the camera's blind spot for a long time", and "hiding small items with large items".
[0021] The information processing device 100 generates multiple candidate detection patterns and matching patterns by setting a list of detection targets as a prompt for generating detection patterns and matching patterns and inputting it into the LLM.
[0022] Figure 4 shows an example of the data structure for detection patterns and matching patterns. The example shown in Figure 4 includes detection patterns 5-1, 5-2, 5-3 and matching pattern 5-4. Detection patterns 5-1 to 5-3 each define the conditions for the object to be detected. Detection pattern 5-1 defines "Subject," "Object," and "Relationship." For example, detection pattern 5-1 shows a relationship where a person corresponding to "Subject" approaches a forklift corresponding to "Object." "Relationship" is an example of interaction information.
[0023] For example, using the store in Example 1 as an example, in detection pattern 5-1, the person corresponding to "Subject" will indicate a relationship that shows actions (such as holding, putting in, or hiding) towards "Object" such as goods.
[0024] Detection patterns 5-2 and 5-3 define "Subject" and "Attribute." For example, detection pattern 5-2 indicates that the person corresponding to "Subject" is wearing a vest. "Attribute" is an example of attribute information.
[0025] For example, using the store in Example 1 as an example, detection pattern 5-2 would indicate the attribute that the person corresponding to "Subject" is holding a product. In other words, "the person is wearing a vest" can be reinterpreted as "the person is holding a product," "the person is holding a shopping basket (cart)," or "the person is holding a container of personal belongings."
[0026] Matching pattern 5-4 further defines the conditions for the matching target for each detection target that matches the conditions of detection patterns 5-1, 5-2, and 5-3. For example, matching pattern 5-4 defines "Detection target" and "Pattern". "Pattern" defines a pattern in which a person is approaching a forklift and the forklift is moving. In such a "Pattern", whether or not the person is approaching the forklift is determined based on detection pattern 5-1. Whether or not the forklift is moving is determined based on detection pattern 5-3. In addition, as defined in detection pattern 5-2, information that the target person is wearing a vest may be further set in "Pattern".
[0027] For example, taking the store in Example 1 as an example, "Pattern" defines a pattern in which a person approaches a product and is holding a personal container in the hand opposite to the hand holding the product. In such "Pattern," whether or not the person is holding a product is determined based on detection pattern 5-1. Whether or not the person is holding a personal container in the hand opposite to the hand holding the product is determined based on detection pattern 5-3. In other words, "a person is approaching a forklift" can be reinterpreted as "a person is approaching a product," and "the forklift is moving" can be reinterpreted as "the person is holding a personal container in the hand opposite to the hand holding the product."
[0028] If video 10 matches the "Pattern" of matching pattern 5-4, it is determined that the matching conditions shown in "Detection target" are met.
[0029] The information processing device 100 evaluates multiple candidate detection patterns and matching patterns, and selects the optimal detection pattern and matching pattern based on the evaluation results. The information processing device 100 generates a knowledge graph 50 based on the selected detection pattern and matching pattern.
[0030] Figure 5 shows an example of a knowledge graph. For example, the knowledge graph 50 shown in Figure 5 is generated based on detection patterns 5-1 to 5-3 and matching pattern 5-4. The knowledge graph 50 includes nodes n1-1, n1-2, n1-3, n1-4, and n1-5. Node n1-1 is the node corresponding to "Subject is wearing a vest". Node n1-2 is the node corresponding to Person. An arrow is set from node n1-1 to node n1-2, indicating that the Subject of node n1-1 is defined in node n1-2.
[0031] Nodes n1-3 correspond to the "Subject is moving" node. Node n1-2 corresponds to the forklift node. An arrow is set from node n1-3 to node n1-4, indicating that the Subject of node n1-3 is defined in node n1-4.
[0032] For example, taking the store in Example 1 as an example, nodes n1-3 are nodes corresponding to "Subject is holding a personal container (with the hand opposite to the one holding the product)." Node n1-2 is the node corresponding to the product.
[0033] Nodes n1-5 are nodes corresponding to "Subject is approaching Object". An arrow is set from node n1-5 to node n1-2, indicating that the Subject of node n1-5 is defined in node n1-2. An arrow is set from node n1-5 to node n1-4, indicating that the Object of node n1-5 is defined in node n1-4. Note that the knowledge graph 50 may be generated from detection patterns only. In that case, the knowledge graph 50 may be represented using the data structures 5-1 to 5-3. Furthermore, when the knowledge graph 50 is generated from both detection patterns and matching patterns, it may be represented using the data structures 5-1 to 5-4.
[0034] (ASG Generation Process) Next, the ASG generation process performed by the information processing device 100 will be described. The ASG generation process uses the detection patterns of the knowledge graph 50 to generate an Action Scene Graph 60 from the video 10 in which "information on the on-site characteristics of the store is associated with the type of person's action." The ASG is also called a Video Scene Graph or Spatio-temporal scene graph.
[0035] For example, the information processing device 100 performs object detection using a detection pattern on time-series frames of the video 10 and tracks the detected objects. The information processing device 100 generates video clips by summarizing the detection results and tracking results for a predetermined number of frames. The information processing device 100 inputs the video clips and prompts for relationship and attribute detection generated from the detection patterns into a visual detection model such as a Vision Language Model (VLM), thereby identifying the attribute information of the detected objects contained in the video clips, interaction information between detected objects, and the time when the attribute information and interaction information occurred.
[0036] The information processing device 100 generates an action scene graph 60 based on a video clip, attribute information of the detection target identified from the video clip, interaction information between the detection targets, and time. The action scene graph 60 maintains the relationship between Subject, object, and relation, or the relationship between Subject, object, and attribute, on an event basis (attribute information, relation information).
[0037] Figure 6 shows an example of an action scene graph. As shown in Figure 6, the action scene graph 60 has time nodes n2-1, n2-2, n2-3, n2-4, n2-5, and n2-6. The action scene graph 60 has event nodes n3-1, n3-2, n3-3, n3-4, n3-5, and n3-6. The action scene graph 60 has concrete object nodes n4-1, n4-2, n4-3, n4-4, and n4-5.
[0038] Time nodes n2-1 to n2-6 are nodes that indicate time, and correspond to times T1, T2, T3, T4, T5, and T6, respectively. For example, times T1, T2, T3, T4, T5, and T6 are associated with the time (frame number) of each frame contained in the video clip.
[0039] Event nodes n3-1 to n3-6 are nodes corresponding to attribute information and interaction information. For example, event nodes n3-1 to n3-3 correspond to "wearing a vest". Event nodes n3-4 and n3-6 correspond to "moving". Event node n3-5 corresponds to "approaching".
[0040] For example, using the store in Example 1, event nodes n3-1 to n3-3 correspond to "holding the product". Event nodes n3-4 and n3-6 correspond to "holding a personal item container with the hand opposite to the hand holding the product". Event node n3-5 corresponds to "person approaching the product".
[0041] The specific object nodes n4-1 to n4-5 are nodes corresponding to the detection target. For example, the specific object nodes n4-1 to n4-4 respectively correspond to the persons P1, P2, P3, and P4. The specific object node 4-5 corresponds to a forklift. For example, taking the store in the first embodiment as an example, the specific object node 4-5 corresponds to a commodity.
[0042] By using the action scene graph 60, it becomes possible to grasp various information regarding the video 10. For example, the event node n3-1 connected to the time nodes n2-1 and n2-6 is connected to the specific object node n4-2. This indicates that the person P2 wearing the best is present in the video 10 at times T1 to T6. For example, taking the store in the first embodiment as an example, it indicates that the person P2 having a commodity is present in the video 10 at times T1 to T6.
[0043] The event node n3-2 connected to the time nodes n2-1 and n2-6 is connected to the specific object node n4-3. This indicates that the person P3 wearing the best is present in the video 10 at times T1 to T6. For example, taking the store in the first embodiment as an example, it indicates that the person P3 having a commodity is present in the video 10 at times T1 to T6.
[0044] The event node n3-3 connected to the time nodes n2-1 and n2-6 is connected to the specific object node n4-4. This indicates that the person P4 wearing the best is present in the video 10 at times T1 to T6. For example, taking the store in the first embodiment as an example, it indicates that the person P4 having a commodity is present in the video 10 at times T1 to T6.
[0045] The event node 3-4 connected to the time nodes n2-1 and n2-3 is connected to the specific object node n4-5. This indicates that the moving forklift is present in the video 10 at times T1 to T3. For example, taking the store in the first embodiment as an example, it indicates that a person holding a personal belongings (with the hand opposite to the hand holding the commodity) is present in the video 10 at times T1 to T3.
[0046] The event node n3-5 connected to the time nodes n2-2 and n2-3 is connected to the concrete object nodes n4-1 and n4-5. This indicates that the event that person P1 approaches the moving forklift exists at times T2 to T3 in video 10. For example, taking the store in Example 1 as an example, it indicates that the event that person P1 holds a personal item (with the hand opposite to the hand holding the merchandise) exists at times T2 to T3 in video 10.
[0047] The event node n3-6 connected to the time nodes n2-5 and n2-6 is connected to the concrete object node n4-5. This indicates that the moving forklift exists at times T5 to T6 in video 10. For example, taking the store in Example 1 as an example, it indicates that the person holding a personal item (with the hand opposite to the hand holding the merchandise) exists at times T5 to T6 in video 10.
[0048] (Graph analysis process) Next, the Graph analysis process executed by the information processing device 100 will be described. The Graph analysis process is a process of analyzing the action scene graph 60 using the LLM and generating an answer when the information processing device 100 receives a question sentence 11 including a request regarding measures to be applied in store 1 from user U1. For example, when the information processing device 100 receives a question sentence 11 related to video 10 from user U1, the generation AI (for example, LLM) generates an answer to the question sentence 11 based on the generated action scene graph 60. More specifically, when the information processing device 100 receives a question sentence regarding the first object in the video from the user, based on the generated graph data, it identifies a result indicating interaction information associated with the first object, and based on the result indicating the identified interaction information, the generation AI generates an answer to the question sentence. For example, when the information processing device 100 receives a question sentence regarding the first object in the video, it searches the action scene graph 60 to identify a result indicating interaction information associated with the first object. Then, the information processing device 100 generates an answer to the question sentence by inputting a prompt composed of the question sentence and the interaction information into the LLM.
[0049] Furthermore, for example, the information processing device 100 generates a search query based on the question 11 and the knowledge graph 50, and uses this search query to perform a data search on the behavior scene graph 60. The information processing device 100 then generates an answer using the results of the data search.
[0050] <Functional Configuration> Figure 7 is a functional block diagram showing the functional configuration of the information processing device according to Embodiment 1. As shown in Figure 7, the information processing device 100 includes a communication unit 110, an input unit 120, a display unit 130, a storage unit 140, and a control unit 150.
[0051] The communication unit 110 performs data communication with the camera via the network. For example, the communication unit 110 receives video data from the camera. The video data is the video data 10 described above.
[0052] Furthermore, the communication unit 110 performs data communication with an external server via the network. For example, the communication unit 110 receives text data from an external server. This text data is the data of the text 12 described above.
[0053] The input unit 120 is an input device that inputs various types of information to the control unit 150 of the information processing device 100. User U1 may operate the input unit 120 to input the question text 11.
[0054] The display unit 130 is a display device that displays information output from the control unit 150.
[0055] The memory unit 140 includes a knowledge graph 50, an action scene graph 60, a video buffer 70, and a text table 80. The memory unit 140 is a memory, etc.
[0056] The knowledge graph 50 is graph data generated based on detection patterns and matching patterns. For example, the explanation of the knowledge graph 50 is the same as the explanation of the knowledge graph 50 in Figure 5.
[0057] The behavioral scene graph 60 is graph data generated based on the knowledge graph 50 and video data. The explanation of the behavioral scene graph 60 is the same as the explanation of the behavioral scene graph 60 in Figure 6.
[0058] The video buffer 70 is a buffer for storing video data.
[0059] We will now move on to the explanation of the control unit 150. The control unit 150 includes an acquisition unit 151, a KG generation unit 152, an ASG generation unit 153, and a graph analysis unit 154. The control unit 150 is a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), etc.
[0060] The acquisition unit 151 acquires video data captured by the camera. The acquisition unit 151 stores the video data in the video buffer 70. The acquisition unit 151 also acquires text data from an external server. The acquisition unit 151 stores the text data in the text table 80.
[0061] The KG generation unit 152 generates a knowledge graph 50 by executing the above-described KG generation process. The KG generation unit 152 stores the knowledge graph 50 in the storage unit 140.
[0062] The ASG generation unit 153 generates an action scene graph 60 by executing the above ASG generation process. The ASG generation unit 153 stores the action scene graph 60 in the storage unit 140.
[0063] When the Graph analysis unit 154 receives the input of the question text 11, it generates an answer by executing the above-described Graph analysis process. The Graph analysis unit 154 outputs the answer to the display unit 130 for display.
[0064] <Processing Flow> Next, an example of the processing procedure of the information processing device 100 according to Embodiment 1 will be described. Figure 8 is a flowchart of the processing procedure of the information processing device according to Embodiment 1. As shown in Figure 8, the acquisition unit 151 of the information processing device 100 acquires text data and stores it in the text table 80 (step S10). The KG generation unit 152 of the information processing device 100 generates a knowledge graph 50 based on the text data stored in the text table 80 (step S11). The acquisition unit 151 acquires video data and stores it in the video buffer 70 (step S12).
[0065] The ASG generation unit 153 of the information processing device 100 generates an action scene graph 60 based on the video data stored in the video buffer 70 and the knowledge graph 50 (step S13).
[0066] The graph analysis unit 154 of the information processing device 100 receives the question 11 (step S14). The graph analysis unit 154 performs graph analysis and generates an answer (step S15). The graph analysis unit 154 outputs the answer (step S16).
[0067] Next, the effects of the information processing device 100 according to Embodiment 1 will be described. The information processing device 100 generates a knowledge graph 50 through a KG generation process and generates an action scene graph 60 through an ASG generation process. When the information processing device 100 receives a question 11 from user U1, it generates an answer based on the knowledge graph 50 and the action scene graph 60. This makes it possible to generate an accurate answer to the question.
[0068] In Example 2, the generation of the graph analysis described in Example 1 will be explained in detail. Figure 9 is a diagram illustrating the graph analysis unit according to Example 2 in detail. As shown in Figure 9, when the graph analysis unit 154 of the information processing device 100 receives a question 11a from user U1, "Tell me about measures to prevent shoplifting of product A," which includes a request regarding measures to be applied in store 1, it uses LLM to analyze the behavior scene graph 60 and generates an answer 11b.
[0069] For example, the graph analysis unit 154 includes a first generation unit 261, a second generation unit 262, a search unit 263, a third generation unit 264, and a response generation unit 265. Each processing unit will be described in order below.
[0070] First, the processing of the first generation unit 261 will be explained. The first generation unit 261 obtains the question sentence 11a and generates a first prompt 261a based on the question sentence 11a and the detection pattern 16. The first prompt 261a is used to cause the LLM to execute the generation of a search query for searching the behavior scene graph 60. The detection pattern 16 is a natural language sentence that represents the target of detection.
[0071] Figure 10A is a diagram illustrating the processing of the first generation unit. As shown in Figure 10A, the first generation unit 261 generates a first prompt 261a by embedding the result of format conversion of the detection pattern 16 and the question text 11a into a pre-prepared template 17. For example, the first prompt 261a is set to information obtained from the detection pattern 16, which is information related to the structure of the behavior scene graph 60.
[0072] The first generation unit 261 outputs the first prompt 261a to the second generation unit 262.
[0073] Next, the processing of the second generation unit 262 in Figure 9 will be explained. The second generation unit 262 generates a search query 262a by inputting the first prompt 261a to the LLM. For example, the search query 262a includes the node to be searched, interaction information between nodes, and attribute information of the node, with respect to the behavior scene graph 60. The second generation unit 262 outputs the search query 262a to the search unit 263.
[0074] Next, the processing of the search unit 263 in Figure 9 will be explained. Based on the search query 262a, the search unit 263 performs a search on the behavior scene graph 60 and obtains the search result 263a. The search unit 263 outputs the search result 263a to the third generation unit 264.
[0075] Here, we will use Figure 6 to supplement the processing of the search unit 263. For example, suppose the search query 262a specifies "Person" and "Forklift" (or "Product" in the store example) as the nodes to be searched, and "Approaching" is specified as the interaction information. In this case, the search unit 263 identifies the specific object nodes n4-1, n4-5 and the event node n3-5 corresponding to the search query 262a, and searches for the start time and end time of the event node n3-5. In this case, the start time is T2 and the end time is T3. As a result, the search result shows that the time the person is approaching the forklift is T2 to T3.
[0076] For example, suppose the search query 262a specifies "Person" as the node to be searched and "wearing a vest" (or "person holding merchandise" in the store example) as attribute information. In this case, the search unit 263 identifies the specific object nodes n4-1 to n4-3 and the event nodes n3-1 to n3-3 that correspond to the search query 262a, and searches for the start time and end time of these event nodes n3-1 to n3-3, respectively.
[0077] For example, the start time of event node n3-1 is T1 and the end time is T6. The start time of event node n3-2 is T1 and the end time is T6. The start time of event node n3-3 is T1 and the end time is T6. As a result, the search results show that the time each person (P2, P3, P4) is wearing the vest (in the store example, "the time the person is holding the product") is between T1 and T6.
[0078] In addition, a search query 262a may contain multiple search items, and the search unit 263 performs the above search for each search item and generates the information obtained from each search as search result 263a.
[0079] Next, the processing of the third generation unit 264 in Figure 9 will be explained. The third generation unit 264 generates a second prompt 264a based on the search result 263a. The second prompt 264a is used when the LLM is to generate the answer.
[0080] Figures 10B and 10C are diagrams illustrating the processing of the third generation unit. As shown in Figure 10B, the third generation unit 264 generates the second prompt 264a by embedding the question text 11a and the search results 263a into a pre-prepared template 18. Figure 10C shows an example of the second prompt 264a.
[0081] The third generation unit 264 outputs the second prompt 264a to the answer generation unit 265.
[0082] Next, the processing of the answer generation unit 265 will be explained. The answer generation unit 265 generates the answer 11b by inputting the second prompt 264a to the LLM. The answer generation unit 265 outputs the generated answer 11b.
[0083] As described above, the information processing device 100 generates a search query 262a based on the question 11a and the detection pattern 16 related to the behavior scene graph 60, and searches the behavior scene graph 60 based on the search query 262a. The graph analysis unit 252 generates and outputs an answer 11b based on the search results of the behavior scene graph 60. This makes it possible to generate an accurate answer to the question.
[0084] The information processing device 100 generates a first prompt 251a based on the question text 11a and the detection pattern 16 related to the behavior scene graph 60, and inputs it to the LLM to generate a search query 262a. This allows for the efficient generation of the search query 262a.
[0085] The information processing device 100 identifies nodes corresponding to the search query from the behavior scene graph 60 and obtains information of nodes associated with the identified nodes as search results. For example, the information processing device 100 searches for the time at which an event related to attribute information occurred in the video, and the time at which an event related to the interaction information occurred in the video, based on the time node among the nodes corresponding to the search query. This makes it possible to find the time at which an event related to attribute information related to the search query occurred, and the time at which an event related to interaction information occurred.
[0086] By the way, in Example 1, an example of using an action scene graph as an example of graph data was explained, but it is not limited to this, and for example, a scene graph, which is an example of graph data that shows the relationships between each object included in video data, can also be used.Therefore, in Example 2, a description of the scene graph and the scene graph generation process executed by the control unit 150 will be explained in detail.Note that the search query and the like use the same method as in Example 1, only the graph data to be searched is different, so a detailed explanation will be omitted.
[0087] (Explanation of Scene Graph) Figure 11 shows an example of a scene graph. As shown in Figure 11, a scene graph is a directed graph in which objects in the image data are represented as nodes, each node has attributes (e.g., object type), and relationships between nodes are represented as directed edges. In the example in Figure 11, the relationship from the node "Person" with the attribute "Shopkeeper" to the node "Person" with the attribute "Customer" is shown to be "Talk". That is, it is defined that the relationship is "Shopkeeper talks to customer". Also, the relationship from the node "Person" with the attribute "Customer" to the node "Product" with the attribute "Large" is shown to be "Stand". That is, it is defined that the relationship is "Customer stands in front of the shelf of large products".
[0088] The relationships shown here are merely examples. For example, they include not only simple relationships such as "holding," but also complex relationships such as "holding product A in the right hand." It is also possible to store separate scene graphs corresponding to relationships between people and scenes corresponding to relationships between people and objects, or to store a single scene graph that includes all of these relationships. Furthermore, the scene graphs may be generated by the control unit 150 as described later, or they may be pre-generated.
[0089] (Scene Graph Generation) Next, we will explain how to generate a scene graph. Figure 12 is a diagram illustrating an example of scene graph generation showing the relationship between people and objects. As shown in Figure 12, the control unit 150 inputs image data into a trained recognition model and obtains the labels "person (male)", "drink (green)", and the relationship "has" as output results of the recognition model. In other words, the control unit 150 obtains that "a man is holding a green drink". As a result, the control unit 150 generates a scene graph that associates the relationship "has" from the node "person" which has the attribute "male" to the node "drink" which has the attribute "green". Note that the scene graph generation is just one example, and other methods can be used, and administrators can also generate them manually.
[0090] Next, the identification of relationships using a scene graph will be explained. The control unit 150 performs a relationship identification process to identify the relationships between people or between people and objects in the video data, according to the scene graph. Specifically, for each frame included in the video data, the control unit 150 identifies the type of person or object shown in the frame, and uses the identified information to search the scene graph and identify the relationships.
[0091] Figure 13 is a diagram illustrating the identification of relationships using a scene graph. As shown in Figure 13, the control unit 150 identifies the type of person, the type of object, the number of people, etc., in frame 1 by inputting frame 1 into a machine learning model that has been trained on, or by performing known image analysis on frame 1. For example, the control unit 150 identifies "person (customer)" as the type of person and "product (product A)" as the type of object. Then, according to the scene graph, the control unit 150 identifies the relationship "person (customer) has product (product A)" between the node "person" with the attribute "customer" and the node "product A" with the attribute "food". The control unit 150 then performs the above relationship identification process for each subsequent frame, such as frame 2 and frame 3, thereby identifying relationships for each frame.
[0092] As described above, the information processing device 100 according to Embodiment 2 can easily determine relationships appropriate for each store by using a scene graph generated for each store, for example, without having to retrain a machine learning model to suit each store. Therefore, the information processing device 100 according to Embodiment 2 can be easily implemented in this embodiment.
[0093] (Identifying Relationships Using a Scene Graph) Figure 14 shows an example of identifying relationships using a scene graph. The control unit 150 detects objects, including people, from the captured image 250 using an existing detection algorithm, estimates the relationships between each object, and generates a scene graph 259 that represents each object and its relationships, i.e., the context. Here, existing detection algorithms include, for example, YOLO (YOU Only Look Once), SSD (Single Shot Multibox Detector), and RCNN (Region Based Convolutional Neural Networks).
[0094] In the example shown in Figure 14, at least two men (indicated by Bbox 251 and 252), a woman (indicated by Bbox 253), a box (indicated by Bbox 254), and a shelf (indicated by Bbox 255) are detected from the captured image 250. The control unit 150 then extracts the Bbox regions of each object from the captured image 250, extracts feature quantities from each region, estimates the relationships between each object from the feature quantities of the object pairs (Subject, Object), and generates a scene graph 259. In Figure 14, the scene graph 259 shows, for example, that the man (indicated by Bbox 251) is standing on the shelf (indicated by Bbox 255). Furthermore, the relationships shown in the scene graph 259 with respect to the man (indicated by Bbox 251) are not limited to just one. As shown in Figure 14, the scene graph 259 shows all the estimated relationships, including the shelf, being behind the man shown in Bbox 252, and holding the box shown in Bbox 254. In this way, the control unit 150 can identify the relationships between objects and people included in the video by generating a scene graph.
[0095] Now, although embodiments of the present invention have been described, the present invention may be implemented in various other forms besides those described above.
[0096] (Variations) In the above embodiment, we provided an example of an action scene graph in which "types of person's actions," such as a person (consumer, customer) holding a product or grasping a container of personal belongings with a hand other than the hand holding the product, are associated with "store site characteristics" where the person is approaching the product. However, we are not limited to this. As "store site characteristics," for example, store data (store design information) regarding the arrangement of shelves within store 1 and product data (product and inventory information) related to the products placed on the shelves within store 1 can be used.
[0097] In this case, the information processing device 100 can generate graph data that associates store data and product data with the type of person's behavior, as a result of accumulating on-site characteristics when abnormal behavior occurs. For example, using the behavior scene graph 60 as an example, the information processing device 100 analyzes and acquires on-site characteristics from video when abnormal behavior occurs, such as "grabbing a personal belongings container with a hand different from the hand holding the product," "staying in the camera's blind spot for a long time," or "hiding a small product with a large product," and generates a behavior scene graph 60 with the obtained on-site characteristics as further nodes. As a result, the information processing device 100 can search for the behavior scene graph 60 according to the attributes of each store and input it into the LMM, so that it can propose appropriate measures according to the characteristics of the store.
[0098] Furthermore, the information processing device 100 can identify the person who exhibited abnormal behavior, identify the on-site characteristics of store 1 at that time (information on shelves and products around the person), and generate graph data that stores the on-site characteristics at the time the abnormal behavior occurred.
[0099] For example, the information processing device 100 identifies information including an actual abnormal behavior, the type of abnormal behavior, and the characteristics of the site at that time, either through video analysis or by receiving input from an administrator. The information processing device 100 then identifies a node corresponding to the "site characteristics" from the behavior scene graph 60 and adds nodes corresponding to the "abnormal behavior" and "type of behavior" to the identified node. In this way, the information processing device 100 can improve the reliability of the behavior scene graph 60 by adding information such as shoplifting that actually occurred to the behavior scene graph 60 generated from the video analysis results, thereby improving the accuracy of the proposed measures.
[0100] Furthermore, whenever the information processing device 100 detects an abnormal behavior by a person towards a product, it repeatedly performs a process that associates information regarding the on-site characteristics of store 1 with the type of abnormal behavior the person exhibited (for example, picking up a product), thereby generating graph data that accumulates on-site characteristics at the time of abnormal behavior.
[0101] For example, the information processing device 100 identifies a node corresponding to the received "site characteristics" or identified "site characteristics" from the behavior scene graph 60, similar to the above processing, and adds (generates) a new node to the identified node that corresponds to the received or identified "abnormal behavior" and "behavior type".
[0102] In this way, each time an abnormal behavior is detected, the information processing device 100 adds information about the abnormal behavior, such as shoplifting, that actually occurred to the behavior scene graph 60 generated from the video analysis results. As a result, even if an unknown abnormal behavior occurs, the information processing device 100 can add information about that unknown abnormal behavior to the behavior scene graph 60 and propose appropriate measures that follow the abnormal behavior. Although the behavior scene graph 60 is used as an example, the same can be applied to the scene graph.
[0103] (Numerical values, etc.) The nodes, relationships, specific examples, numerical values, etc., within the graph used in the above example are merely examples and can be changed as needed. The processing flow described in each flowchart can also be changed as appropriate within a range that does not contradict each other. In addition, a small-scale language model other than LLM can also be used as the language model. Examples of detection targets include people and products.
[0104] (Language Model) The language model is a transformer-based model trained using a token set generated from a token set in which some tokens are masked from a set of multiple tokens. For example, the information processing device 100 trains the language model using a token set (unsupervised learning dataset). The language model consists of a deep neural network. For example, the language model is a machine learning model that incorporates an architecture called a transformer. In other words, the language model is a transformer-based language model. One such transformer is BERT (Bidirectional Encoder Representation from Transformers). The information processing device 100 trains the language model by masking some of the tokens included in the token set and estimating the masked tokens.
[0105] Furthermore, the language model may be a large-scale language model in which, for example, the three elements of computational complexity, data volume, and the number of model parameters are scaled up. Computational complexity indicates the amount of work that the computer processes. Data volume indicates the amount of information in the text data input to the computer. The number of model parameters refers to parameters specific to deep learning technology.
[0106] (System) The processing procedures, control procedures, specific names, and various data and parameters shown in the above documents and drawings may be changed at will unless otherwise specified.
[0107] Furthermore, the specific forms of distribution and integration of the components of each device are not limited to those shown in the figures. For example, the KG generation unit 152 and the ASG generation unit 153 may be integrated. In other words, all or part of the components may be functionally or physically distributed or integrated in any unit depending on various loads and usage conditions. Moreover, all or any part of the processing functions of each device may be realized by a CPU and a program that is analyzed and executed by the CPU, or by hardware using wired logic.
[0108] Furthermore, each processing function performed by each device can be implemented, in whole or in part, by a CPU and a program executed for analysis by that CPU, or by hardware using wired logic.
[0109] (Hardware) Figure 15 is a diagram illustrating an example of hardware configuration. As shown in Figure 15, the information processing device 100 includes a communication device 100a, an HDD (Hard Disk Drive) 100b, memory 100c, and a processor 100d. Furthermore, each of the components shown in Figure 15 is interconnected by a bus or the like.
[0110] The communication device 100a is a network interface card or the like, and communicates with other devices. The HDD 100b stores programs and databases that operate the functions shown in Figure 7.
[0111] The processor 100d operates a process that performs the functions described in Figure 7 by reading a program that performs the same processing as each processing unit shown in Figure 7 from the HDD 100b or the like and loading it into the memory 100c. For example, this process performs the same functions as each processing unit of the information processing device 100. Specifically, the processor 100d reads a program that has the same functions as the acquisition unit 151, KG generation unit 152, ASG generation unit 153, Graph analysis unit 154, etc. from the HDD 100b or the like. Then, the processor 100d executes a process that performs the same processing as the acquisition unit 151, KG generation unit 152, ASG generation unit 153, Graph analysis unit 154, etc.
[0112] Thus, the information processing device 100 operates as an information processing device that executes the generation method by reading and executing a program. The information processing device 300 can also achieve the same functionality as the embodiment described above by reading the program from the recording medium using a media reader and executing the read program. It should be noted that the program referred to in this other embodiment is not limited to being executed by the information processing device 300. For example, the above embodiment may also be applied to cases where another computer or server executes the program, or where they cooperate to execute the program.
[0113] This program may be distributed via a network such as the Internet. Alternatively, this program may be recorded on a computer-readable recording medium such as a hard disk, flexible disk (FD), CD-ROM, MO (Magneto-Optical disk), or DVD (Digital Versatile Disc), and executed by being read from the recording medium by a computer.
[0114] 50 Knowledge Graph 60 Behavioral Scene Graph 70 Video Buffer 80 Text Table 100 Information Processing Unit 110 Communication Unit 120 Input Unit 130 Display Unit 140 Storage Unit 150 Control Unit 151 Acquisition Unit 152 KG Generation Unit 153 ASG Generation Unit 154 Graph Analysis Unit
Claims
1. A generation program that causes a computer to execute a process that involves acquiring video footage of the inside of a facility, analyzing the acquired video footage to identify the types of behavior of people staying in the facility, generating graph data in which information about the on-site characteristics of the facility is associated with the identified types of behavior of the people, and inputting a prompt into a language model that includes the search results of the generated graph data and a request regarding measures to be applied within the facility, thereby generating information about the measures as a response to the request.
2. The generation program according to claim 1, wherein the identifying process identifies a person who has performed abnormal behavior based on the results of analyzing the video; the graph data generation process generates graph data in which the on-site characteristics at the time the abnormal behavior occurred are accumulated based on the surrounding information of the identified person; and the generation process, when it receives a request regarding measures to be applied in the store, searches the graph data in which the on-site characteristics are accumulated; and inputs a prompt including the search result of the graph data to a language model to generate information regarding the measures as a response to the request.
3. The generation program according to claim 2, wherein the identifying process identifies abnormal behavior of the person toward a product based on the results of analyzing the video; the graph data generation process generates graph data in which the on-site characteristics of the store are accumulated each time abnormal behavior of the person toward a product is detected is associated with the type of behavior of the person when the abnormal behavior occurred; and the generation process searches the graph data in which the on-site characteristics are accumulated when a request for measures to be applied in the store is received, and generates information regarding the measures as a response to the request by inputting a prompt to the language model, which includes information regarding the on-site characteristics of the store associated with the type of abnormal behavior that occurred a predetermined number of times, based on the search results of the graph data.
4. The information relating to the site characteristics includes store data relating to the location of shelves within the store and product data relating to products placed on the shelves within the store, and the process for generating the graph data generates the graph data in which the store data and product data are associated with the type of behavior of the person, as a result of the accumulation of site characteristics when the abnormal behavior occurs. This is the generation program according to claim 2.
5. A generation program according to claim 1, which causes the computer to perform the following processes: acquire input data to which attribute information of the detection target or interaction information between the detection targets is associated with the detection target; analyze the video to be analyzed using the input data to detect a first detection target representing the detection target from among the video frames constituting the video; input a prompt including the input data and the detected first detection target into a visual language model to generate a result showing attribute information of the first detection target or interaction information relating to the first detection target; and generate graph data to which the generated result and the detected first detection target are associated.
6. A generation program according to claim 5, which causes the computer to perform the following processes: obtaining information about the structure of graph data to be searched and a request concerning an object included in the video, which is a request concerning a measure to be applied in the store; generating a search query for the graph data based on the information about the structure of the graph data to be searched; searching for graph data in which attribute information of the object or interaction information between the objects is associated with the object included in the video based on the generated search query; and outputting information about the object as a response to the request based on the search results of the retrieved graph data.
7. The generation program according to claim 1, wherein the generation process involves inputting a prompt based on the search result into a large-scale language model to generate a response to the request and outputting the response.
8. The generation program according to claim 1, wherein the language model is a transformer-based model trained using a token set generated from a token set obtained by masking some tokens from a plurality of tokens.
9. A generation method that performs a process comprising: a computer acquiring video footage of the inside of a facility; analyzing the acquired video footage to identify the types of behavior of persons staying in the facility; generating graph data in which information regarding the on-site characteristics of the facility is associated with the identified types of behavior of the persons; and inputting a prompt into a language model that includes the search results of the generated graph data and a request regarding measures to be applied within the facility, thereby generating information regarding the measures as a response to the request.
10. An information processing device having a control unit that acquires video footage of the inside of a facility, analyzes the acquired video footage to identify the types of behavior of people staying in the facility, generates graph data in which information about the on-site characteristics of the facility is associated with the identified types of behavior of the people, and inputs a prompt to a language model that includes the search results of the generated graph data and a request regarding measures to be applied in the facility, thereby generating information about the measures as a response to the request.
Citation Information
Patent Citations
Data analysis device and data analysis method
JP2022086650A
Information processing program, method for processing information, and information processor
JP2023098482A
Method and system for generating questions and answers based on knowledge graph
KR102697127B1