Agent system, and method for generating technical management information for the agent system

The agent system recreates on-site scenarios to elicit tacit knowledge from skilled workers through dialogue, addressing the challenge of capturing undocumented expertise and facilitating the development of specialized language models and chatbots.

JP2026136888AActive Publication Date: 2026-08-26HITACHI INDUSTRY & CONTROL SOLUTIONS LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025022719
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2026-08-26
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively capture tacit knowledge from skilled workers due to their unawareness of possessing such knowledge, difficulty in explaining it, and the challenge of recreating emergency situations for documentation.

Method used

An agent system utilizing a field video display, questioning unit with a large-scale language model, and data collection unit to recreate on-site scenarios and elicit tacit knowledge through dialogue with users, mimicking a grandchild's tone to facilitate knowledge extraction.

Benefits of technology

Enables the collection and documentation of tacit knowledge from skilled workers, facilitating the construction of specialized large language models and chatbots with equivalent expertise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026136888000001_ABST
    Figure 2026136888000001_ABST
Patent Text Reader

Abstract

Gather tacit knowledge from experts. [Solution] The agent system 2 is characterized by comprising a video playback unit 26 that displays a video recreating the site on a display 41 and presents it to the user 4, a question unit 27 that has the agent, which uses a large-scale language model 200, ask the user 4 how to respond in each situation at the manufacturing site 1, and a data collection unit 28 that collects the dialogue between the agent and the user 4.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0006] ,

[0001] The present invention relates to an agent system and a method for generating technical management information of an agent system.

Background Art

[0002] In recent years, the retirement of skilled technicians due to the progress of aging has become a major issue. Before the retirement of skilled technicians, it is necessary to transfer the technologies they possess. In technology transfer, it is important to extract tacit knowledge that cannot be seen. If the tacit knowledge of skilled technicians can be sufficiently extracted, it is possible to construct a specialized large language model (LLM) with knowledge equivalent to that of skilled workers or a highly accurate chatbot. [[ID=1৪]]

[0003] Patent Document 1 describes an invention of a chatbot system for training communication ability in a specific situation.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] Although the construction method of a specialized LLM or the like is clear, the information to be learned must include tacit knowledge. In the first place, skilled workers are not aware that their knowledge is tacit knowledge. Skilled workers often use tacit knowledge in cases where they act normally or in cases where they act using tacit knowledge when faced with an emergency. And it is difficult to make all emergency situations occur in reality. Furthermore, skilled workers often have difficulty explaining technologies to others in an understandable manner or documenting them.

[0006] Therefore, an object of the present invention is to collect tacit knowledge from skilled workers. [Means for solving the problem]

[0007] To solve the aforementioned problems, the agent system of the present invention is characterized by comprising: a field video display unit that displays a video recreating the field to the user; a questioning unit that causes an agent using a large-scale language model to ask the user how to respond to each situation at the field; and a data collection unit that collects the dialogue between the agent and the user. The present invention provides a method for generating technical management information for an agent system, comprising the steps of: the agent system's field video display unit displays a video recreating the site to the user; the agent system's questioning unit prompts the user to ask an agent using a large-scale language model how to respond to various situations at the site; the agent system's data collection unit collects the dialogue between the agent and the user; and the agent system's field technical data generation unit generates information related to on-site technical management from the dialogue collected by the data collection unit. Other means will be described within the descriptions of embodiments for carrying out the invention. [Effects of the Invention]

[0008] According to the present invention, it becomes possible to collect tacit knowledge from skilled individuals. [Brief explanation of the drawing]

[0009] [Figure 1] This is a diagram illustrating the configuration of the agent system according to the first embodiment. [Figure 2] This is a diagram showing the maintenance record information. [Figure 3] This is a diagram illustrating some of the functional components of the agent system. [Figure 4] This figure shows the prompts used by the data preprocessing unit. [Figure 5] This is a diagram showing information about equipment coordinate failures. [Figure 6]It is a diagram showing the prompt used by the question content generation unit. [Figure 7] It is a diagram showing question information. [Figure 8] It is a diagram explaining a partial functional unit of the agent system. [Figure 9A] It is a diagram explaining the dialogue screen between the agent and the user. [Figure 9B] It is a diagram explaining the dialogue screen between the agent and the user. [Figure 9C] It is a diagram explaining the dialogue screen between the agent and the user. [Figure 10] It is a flowchart of the processing of the video playback unit. [Figure 11] It is a flowchart of the processing of the question unit. [Figure 12] It is a diagram showing dialogue information. [Figure 13] It is a diagram showing the prompt used by the data collection unit. [Figure 14] It is a diagram showing question and answer data. [Figure 15] It is a configuration diagram of the agent system according to the second embodiment. [Figure 16] It is a diagram explaining a partial functional unit of the agent system.

Mode for Carrying Out the Invention

[0010] Hereinafter, the mode for carrying out the present invention will be described in detail with reference to each figure. In the system of this embodiment, a video showing a situation such as a failure or an accident is generated by a generation AI and synthesized into a walkthrough video of the site. Then, the user converses with the agent while viewing the video of the virtual manufacturing site. Thereby, the system of this embodiment collects on-site technical data including tacit knowledge from the user.

[0011] That is, the system of this embodiment shows various situations including emergencies to the user in the form of images, and then the agent asks the user about the responses in each situation, so as to elicit the tacit knowledge of the user in all situations. The agent that asks the user questions is assumed to be a grandson who makes elderly technicians feel a sense of fulfillment.

[0012] Regarding when people feel a sense of fulfillment, according to a Cabinet Office survey that investigated four countries including Japan, the proportion was relatively high in all countries when it was "time for gathering with family such as children and grandchildren". In Japan, 55.3% of people replied that they feel a sense of fulfillment when gathering with family such as children and grandchildren. Therefore, in the agent system of this embodiment, "grandson" is adopted as the agent that asks the user questions.

[0013] Figure 1 is a configuration diagram of an agent system 2 according to the first embodiment. The agent system 2 includes a large language model 200, a data preprocessing unit 21 for performing preliminary preparations, an emergency situation generation unit 22, a mapping information generation unit 23, and a question content generation unit 25.

[0014] The data preprocessing unit 21 converts the maintenance record information 31 and the equipment installation location information 32 of the equipment maintenance management system 3 into markdown format by using processing based on the large language model 200 or rule-based processing. The equipment maintenance management system 3 is a system for assisting in the maintenance management of each piece of equipment installed in a factory, a plant, etc. The maintenance record information 31 is information recording the maintenance of each piece of equipment installed in a factory, a plant, etc. The equipment installation location information 32 is information storing the installation location of each piece of equipment installed in a factory, a plant, etc.

[0015] The emergency situation generation unit 22 uses an image generation AI to generate images showing the emergency situation based on maintenance record information 31 converted to Markdown format. The maintenance record information 31 includes information such as the circumstances at the time of the malfunction. For example, the emergency situation generation unit 22 reads an image of factory equipment and instructs the image generation AI to generate prompts that cause events such as water leaks or smoke. This generates a video showing the events described in the prompts.

[0016] Specifically, the emergency situation generation unit 22 inputs an image of a tank in the factory and instructs the image generation AI with the prompt, "Please simulate a water leak from this tank." The image generation AI then generates a video of water leaking from a tank in the factory. The emergency situation generation unit 22 repeats this process for each piece of equipment to obtain emergency videos for each piece of equipment.

[0017] The mapping information generation unit 23 maps the corresponding equipment to the location information recorded as metadata in the on-site walkthrough video 111, based on the information obtained by converting the maintenance record information 31 into Markdown format. This on-site walkthrough video 111 was filmed at the manufacturing site 1 with a 360-degree camera 11 along with location information. The mapping information generation unit 23 may, but is not limited to, map the corresponding equipment based on the time information of the on-site walkthrough video 111 using a rule-based method.

[0018] The question content generation unit 25 generates questions for user 4 based on the maintenance record information converted into Markdown format by the data preprocessing unit 21. The question content generation unit 25 generates questions in a tone that replicates "questions from a grandchild" by prompting the large-scale language model 200.

[0019] The video synthesis unit 24 uses a machine learning model to synthesize a pre-recorded on-site walkthrough video 111 taken with the 360-degree camera 11 with an emergency on-site video 221 based on the time information and equipment mapping information of the on-site walkthrough video 111. Through this process, the video synthesis unit 24 creates video footage of the on-site space where the emergency occurred.

[0020] The agent system 2 further comprises a video playback unit 26, a questioning unit 27, a data collection unit 28, a field technical data generation unit 29, and field technical data 291, all of which are used during operation.

[0021] The video playback unit 26 displays a video that recreates the on-site environment to the user 4 by playing the video on the display 41. The video playback unit 26 functions as an on-site video display unit that displays a video that recreates the on-site environment to the user 4. The video playback unit 26 displays a video on the display device 41, which is a composite of a pre-recorded on-site video and on-site video of an emergency. When the video displays equipment in an emergency, the video playback unit 26 causes the questioning unit 27 to start asking the user 4 questions about the details of the emergency and how to respond.

[0022] User 4 visually views the site environment on display 41. When the site environment image shows a specific piece of equipment, a pre-prepared question is output via synthesized speech from smartphone 42. User 4 then answers this question verbally.

[0023] When the video playback unit 26 displays specific equipment on the display 41, the question unit 27 asks the user 4 questions that have been pre-generated by the question content generation unit 25, using synthesized speech and text display. The user 4 answers these questions verbally. Each question and answer is in a one-question-one-answer format. For each answer, the question unit 27 continues the dialogue by repeatedly asking follow-up questions such as "Why?". During the period from the start to the end of this dialogue, the video playback unit 26 pauses video playback. This prevents the video playback unit 26 from ending playback of the target equipment information or starting playback of the next target equipment information during the dialogue. In addition, the sound of the on-site walkthrough video 111 can be muted during the dialogue, so it does not interfere with the dialogue between the agent and the user 4.

[0024] This questioning unit 27 has an agent using a large-scale language model 200 ask the user 4 questions in synthesized speech and text about how to respond to each situation that occurs at each piece of equipment in the manufacturing site 1.

[0025] The questioning unit 27 can also receive utterances initiated by user 4 without prior questioning. The questioning unit 27 accepts triggers for utterances by user 4 using specific keywords such as "Hey, you know." In other words, when the questioning unit 27 detects user 4's utterance, it initiates a dialogue with user 4 based on that utterance. This allows the questioning unit 27 to extract tacit knowledge based on normal video footage.

[0026] Then, when User 4's answer is repeated a predetermined number of times, or when a predetermined keyword is detected from User 4's answer, the questioning unit 27 speaks a summary of the conversation so far to User 4 in synthesized speech and text, and ends the dialogue with User 4. Through this dialogue, it is possible to appropriately collect the tacit knowledge that User 4 possesses.

[0027] The questioning unit 27 further records the conversation between user 4 and the agent, converts user 4's answers into text information, and generates follow-up questions that delve deeper into the topic. The conversation between user 4 and the agent may be recorded not only by the questioning unit 27, but also by the smartphone 42.

[0028] The data collection unit 28 collects video data of the portion that the video playback unit 26 was playing during the conversation between the agent and user 4, and text information of the conversation between the agent and user 4 collected by the questioning unit 27. By collecting video data of the portion that the video playback unit 26 was playing during the conversation between the agent and user 4, the data collection unit 28 can retrospectively review the parts that user 4 is explaining using pronouns, etc., in video.

[0029] The field technical data generation unit 29 generates and manages field technical data 291 related to on-site technical management using a large-scale language model 200 from video data and dialogue text information. The field technical data 291 is question-answer data in which questions and answers regarding equipment failures are associated. This field technical data 291 reflects the tacit knowledge of the user 4. Alternatively, the field technical data generation unit 29 may generate and manage field technical data 291 related to on-site technical management using a rule-based method from video data and dialogue text information. In addition to the question-and-answer combinations, the field technical data generation unit 29 may also store still images of the equipment related to the question, such as video clips, as part of the field technical data 291. This makes it possible to retrospectively review the contents of the field technical data 291 in video format.

[0030] Figure 2 shows the maintenance record information 31. Maintenance record information 31 consists of an ID field, an equipment name field, an installation location field, a failure description field, a failure cause field, and a response details field.

[0031] The ID field stores the identification information for the preservation record. The "Equipment Name" column stores the names of each piece of equipment installed in manufacturing site 1. The "Installation Location" column stores the location of each piece of equipment installed at manufacturing site 1. The "Fault Details" column stores the details of the faults that occur in each piece of equipment installed at manufacturing site 1.

[0032] The "Cause of Failure" column stores the cause of failures occurring in each piece of equipment installed at Manufacturing Site 1. The "Response" column stores the details of the response taken to each piece of equipment installed at Manufacturing Site 1.

[0033] Figure 3 is a diagram illustrating some of the functional components of the agent system 2. The data preprocessing unit 21 of the agent system 2 receives maintenance record information 31 and equipment installation location information 32 as input, and outputs maintenance record markdown information 211 and equipment installation location markdown information 212. The data preprocessing unit 21 may, but is not limited to, process using either a large-scale language model 200 or rule-based processing.

[0034] These maintenance record markdown information 211 and equipment installation location markdown information 212 are input to the mapping information generation unit 23. The mapping information generation unit 23 outputs equipment coordinate fault occurrence information 231. The mapping information generation unit 23 may, but is not limited to, process using either a large-scale language model 200 or a rule-based process.

[0035] The maintenance record markdown information 211 is further input to the question content generation unit 25. The question content generation unit 25 then generates question information 251. The question content generation unit 25 may, but is not limited to, process using either the large-scale language model 200 or rule-based processing.

[0036] Figure 4 shows the prompt 230 used by the mapping information generation unit 23. The prompt 230 is transcribed below as text.

[0037] # Tasks Please generate the coordinate information for each piece of equipment under the following constraints. # Constraints • Generate in 3D • Includes equipment used in the event of a failure.

[0038] The mapping information generation unit 23 outputs equipment coordinate fault occurrence information 231 based on the maintenance record markdown information 211 and the equipment installation location markdown information 212. The mapping information generation unit 23 generates the equipment coordinate fault occurrence information 231, which is mapping information, by providing a prompt 230 to the large-scale language model 200. However, the mapping information generation unit 23 may also generate the equipment coordinate fault occurrence information 231, which is mapping information, in a rule-based manner, and is not limited to this.

[0039] Figure 5 shows the equipment coordinate fault occurrence information 231. The equipment coordinate failure information 231 consists of an equipment name field, a coordinate field, and a failure status field. The "Equipment Name" field stores the name of each piece of equipment. The coordinate field stores the 3D coordinates of each piece of equipment. The "Fault Occurrence Status" column stores descriptive text about the types of faults that may occur in each piece of equipment.

[0040] Figure 6 shows the prompt 250 used by the question content generation unit 25. The following is a text transcript of prompt 250.

[0041] # Tasks Under the following constraints, ask root-of-the-fact questions about the maintenance records. Please # Constraints • Speak in a tone like a grandchild asking their grandfather a question. Ask questions from a child's perspective. • Ask questions about each piece of equipment. • The timing for asking a question is when you reach a specific piece of equipment. • Continue asking further questions in response to the answer. • For the third response, simply answer without asking a question.

[0042] The question content generation unit 25 generates question information 251 by providing the prompt 250 shown in Figure 6 to the large-scale language model 200. This question information 251 is shown in Figure 7, which will be described later.

[0043] Furthermore, considering the possibility that user 4, who is an expert, is female, it is advisable to have the user input their gender beforehand. If user 4 is female, the constraint condition should be modified to "speak in a tone like a grandchild asking a question to their grandmother." This ensures that even if user 4, an expert, is female, the questions can be asked without sounding unnatural.

[0044] Figure 7 shows Question Information 251. Question Information 251 contains the questions that the agent asks user 4 when they reach each piece of equipment. The text of Question Information 251 is transcribed below.

[0045] When it reaches the conveyor belt: "Grandpa, why is smoke coming out of the conveyor belt? Is it because the motor got too hot?" When it reaches the cooling tank: "Grandpa, why was water dripping from the cooling tank? Is the tank old?" When it reaches the press: "Grandpa, why was the press rattling? Why do bearings wear out?" When you reach the boiler: "Grandpa, why was steam coming out of the boiler? Why do the gaskets deteriorate?" When it reaches the air compressor: "Grandpa, why was there a strange noise and smoke coming from the air compressor? Why do the internal parts break down?"

[0046] In this way, by having an agent modeled after a grandchild, which makes it easier for older engineers to find meaning in their lives, ask questions, older engineers are more likely to find meaning in their lives and to collect tacit knowledge compared to questions from a dry, uninteresting agent.

[0047] Figure 8 is a diagram illustrating some of the functional components of the agent system 2. The emergency situation generation unit 22 generates an emergency scene video 221 by providing the large-scale language model 200 with a normal scene video 12 and an abnormality prompt 220. When the video synthesis unit 24 receives the on-site walkthrough video 111, the emergency on-site video 221, and the equipment coordinate failure information 231 as input, it synthesizes the on-site walkthrough video 111 and the emergency on-site video 221 to generate a synthesized video 241 and also generates failure metadata 242. The failure metadata 242 is metadata indicating which part of the synthesized video 241 contains the scene of the equipment experiencing the failure. This failure metadata 242 may or may not be stored in each frame of the synthesized video 241.

[0048] The video playback unit 26 plays the synthesized video 241 on the display 41 and, referring to the fault occurrence metadata 242, notifies the questioning unit 27 when it plays a scene of the equipment experiencing a fault. If the questioning unit 27 receives notification of which equipment the malfunction scene relates to based on the question information 251, it will begin asking questions to the user 4 via the smartphone 42.

[0049] Figures 9A to 9C illustrate screen 421, which displays the interaction between the agent and user 4. The screen 421 of the smartphone 42 in Figure 9A displays an agent icon 51, a speech bubble 52a indicating the agent's question, a user icon 53, a speech bubble 54a indicating user 4's response, and a microphone icon 55. This microphone icon 55 indicates that the audio of user 4 is being recorded.

[0050] The agent icon 51 is an icon representing an agent that simulates the grandchild of user 4. The question unit 27, using an application installed on the smartphone 42, outputs the agent's questions as synthesized speech and displays a speech bubble 52a on the screen 421 with the question content as text. This allows user 4 to understand the question content by checking the speech bubble 52a, even if they miss the synthesized speech spoken by the agent.

[0051] User 4 speaks their answer to the agent's question. Smartphone 42 records this speech and outputs it to question unit 27. Question unit 27, using a large-scale language model 200, converts this recorded data into answer text and displays this answer text in speech bubble 54c. This allows user 4 to visually confirm whether their answer has been accurately heard by the agent. Furthermore, the question section 27 uses the large-scale language model 200 to generate a second question that delves deeper into this answer.

[0052] The screen 421 of the smartphone 42 in Figure 9B displays an agent icon 51, a speech bubble 52b indicating the agent's second question, a user icon 53, a speech bubble 54b indicating user 4's second answer, and a microphone icon 55.

[0053] The questioning unit 27, using an application installed on the smartphone 42, outputs the agent's second question in synthesized speech and displays a speech bubble 52b on the screen 421 with the content of the second question displayed as text.

[0054] User 4 speaks their second answer to the agent's second question. Smartphone 42 records this utterance and outputs it to the question unit 27. The question unit 27, using the large-scale language model 200, converts this recorded data into the second answer text and displays this second answer text in speech bubble 54b. Furthermore, the question unit 27, using the large-scale language model 200, generates a third question that delves deeper into this second answer.

[0055] The screen 421 of the smartphone 42 in Figure 9C displays an agent icon 51, a speech bubble 52c indicating the agent's third question, a user icon 53, a speech bubble 54c indicating user 4's third answer, a speech bubble 52d indicating the agent's summary of the conversation, and a microphone icon 55.

[0056] The questioning unit 27 uses an application installed on the smartphone 42 to output the agent's questions as synthesized speech, and also displays a speech bubble 52c on the screen 421 with the content of the questions displayed as text.

[0057] User 4 speaks the third answer to the agent's third question. Smartphone 42 records this utterance and outputs it to question unit 27. Question unit 27, using the large-scale language model 200, converts this recorded data into the third answer text and displays this third answer text in speech bubble 54c. Furthermore, question unit 27, using the large-scale language model 200, generates a text summarizing the conversation so far and displays this summary text in speech bubble 54d.

[0058] Figure 10 is a flowchart of the processing of the video playback unit 26. First, the video playback unit 26 starts playing the video (step S10). Then, the video playback unit 26 determines whether or not the questioning unit 27 has detected that user 4 has started asking a question (step S11). If it has detected that user 4 has started asking a question (Yes), the process proceeds to step S14. If it has not detected that user 4 has started asking a question (No), the process proceeds to step S12.

[0059] In step S12, the video playback unit 26 determines whether the scene being played contains a fault. If the scene being played does not contain a fault (No), the process returns to step S11. If the scene being played contains a fault (Yes), the process proceeds to step S13.

[0060] In step S13, the video playback unit 26 instructs the questioning unit 27 to begin asking questions about the faulty part. Then, the video playback unit 26 pauses the video that is currently playing (step S14).

[0061] The video playback unit 26 determines whether the questioning unit 27 has ended the dialogue (step S15). If the questioning unit 27 has not ended the dialogue (No), the process returns to step S15. If the questioning unit 27 has ended the dialogue (Yes), the process returns to step S10.

[0062] Figure 11 is a flowchart of the processing in question section 27. This flowchart will be explained with reference to Figure 1. First, the questioning unit 27 determines whether or not it has received a conversation spontaneously uttered by user 4 from the audio recorded by the microphone of the smartphone 42 (step S20). If it has received a conversation uttered by user 4 (Yes), the process proceeds to step S23 to generate a more in-depth question, and then proceeds to step S24. If it has not received a question uttered by user 4 (No), the process proceeds to step S21.

[0063] In step S21, the questioning unit 27 determines whether or not it has received an instruction from the video playback unit 26 to start questioning about the faulty location. If it has not received an instruction to start questioning about the faulty location (No), the process returns to step S20. If it has received an instruction to start questioning about the faulty location (Yes), the process proceeds to step S22.

[0064] In step S22, the questioning unit 27 begins asking questions from the agent. Then, in step S24, the questioning unit 27 records the user 4's answers. Next, the questioning unit 27 determines whether the conditions for ending the dialogue with user 4 have been met (step S25). The conditions for ending the dialogue are that user 4 has given a predetermined number of answers, that user 4's answers include a specific keyword, or that user 4 has remained silent for a predetermined period of time.

[0065] If question unit 27 has ended the dialogue (Yes), the process proceeds to step S28. If the condition for ending the dialogue is not met (No), the process proceeds to step S26.

[0066] In step S26, the question unit 27 generates a more in-depth question. Then, the question unit 27 outputs the question from the agent as synthesized speech and text via the smartphone 42 (step S27), and the process returns to step S24.

[0067] In step S28, the question unit 27 generates a summary of the answers. Then, the question unit 27 outputs the summary of the answers from the agent as synthesized speech and text via the smartphone 42 (step S29), and the process returns to step S20.

[0068] Figure 12 shows the dialogue information 271. This dialogue information 271 summarizes the conversation between user 4 and the agent, as recorded by question section 27. An example of dialogue information 271 is transcribed below as text.

[0069] The reason we check the motor's operation is to make sure it's working properly. If it doesn't work, the entire machine won't function correctly. Belts wear out and deteriorate over time, so they need to be replaced. An old belt can break, which can cause serious problems. Sensor calibration is the process of adjusting the sensor so that it can measure accurately. It's extremely important because if you don't do this, you'll end up working with incorrect data. The cooling fan not working could be due to a malfunction or a power supply issue. If the cooling fan doesn't work, the machine could overheat and break down, so you need to check it as soon as possible. Inspecting electrical wiring involves checking that the wires are properly connected and undamaged. Neglecting this can lead to short circuits or fires, so thorough inspections are essential. The reason the next inspection is in three months is so that we can address any problems before they escalate. Long-term inspections are necessary to extend the lifespan of the machinery. Strengthening inspections of the cooling fans is a particularly important part, so it means checking them more frequently. If the cooling isn't working properly, it will affect the entire machine. In training our workers, we teach them the knowledge necessary for safe work and how to use the machinery correctly. If they learn this properly, accidents can be prevented and the work can proceed smoothly.

[0070] Figure 13 shows the prompt 280 used by the data acquisition unit 28. The following is a text transcript of prompt 280.

[0071] # Tasks Please create the Q&A data under the following constraints. # Constraints • Use a text file as the answer to the question. • Use a mechanical tone.

[0072] The data acquisition unit 28 generates field technical data 291 by providing a prompt 280 to the large-scale language model 200.

[0073] Figure 14 shows field technical data 291. Field technical data 291 consists of a question field and an answer field. The question field stores questions related to the failure information of each piece of equipment. The answer field stores information summarizing the user's (4) answers to each question. By training a machine learning model with this field technical data 291, a chatbot that reflects the tacit knowledge of user 4 can be realized. In addition, field technical data 291 can be used to create FAQs (Frequency Asked Questions) or as material for technical training.

[0074] Figure 15 is a diagram showing the configuration of the agent system 2A according to the second embodiment. The agent system 2A according to the second embodiment differs from the first embodiment in that it includes a model deployment unit 20 and a field reproduction space display unit 26A. In the second embodiment, instead of displaying video on the display 41 to the user 4, it displays images of a virtual space to the user 4 using virtual reality goggles 43.

[0075] The model deployment unit 20 deploys the video information synthesized by the video synthesis unit 24 into a virtual space model. Then, the site reproduction space display unit 26A displays the video of the site reproduction space on the virtual reality goggles 43 based on this virtual space model. The site reproduction space display unit 26A functions as a site video display unit that displays the video of the reproduced site to the user 4. User 4 wears virtual reality goggles 43 and experiences a walkthrough of the virtual space. When User 4 reaches the emergency equipment in the virtual space, the on-site simulation space display unit 26A initiates a questioning unit 27 to ask User 4 questions.

[0076] Figure 16 is a diagram illustrating some of the functional parts of the agent system 2A. The emergency situation generation unit 22 generates an emergency scene video 221 by providing the large-scale language model 200 with a normal scene video 12 and an abnormality prompt 220.

[0077] When the video synthesis unit 24 receives the on-site walkthrough video 111, the emergency on-site video 221, and the equipment coordinate failure information 231 as input, it synthesizes the on-site walkthrough video 111 and the emergency on-site video 221 to generate a synthesized video 241 and also generates failure metadata 242. The failure metadata 242 is metadata indicating which part of the synthesized video 241 contains the scene of the equipment experiencing the failure.

[0078] The model deployment unit 20 deploys the synthesized video 241 and fault occurrence metadata 242 into a virtual space model. Then, the site reproduction space display unit 26A displays the image of the virtual space that reproduces the site on the virtual reality goggles 43, based on the virtual space model into which the synthesized video 241 and fault occurrence metadata 242 have been deployed.

[0079] The on-site simulation space display unit 26A further refers to the fault occurrence metadata 242 deployed in the virtual space. When user 4 reaches the equipment where the fault has occurred in the virtual space, the on-site simulation space display unit 26A notifies the questioning unit 27 of this fact.

[0080] Based on the question information 251, the questioning unit 27, upon receiving notification of which equipment's faulty part it has reached, will begin questioning the user 4 using the smartphone 42.

[0081] The configuration and effects of the present invention are described below.

[0082] [1] A site video display unit (video playback unit 26, site recreation space display unit 26A) displays a video that recreates the site for the user (4), An agent using a large-scale language model (200) is used to ask the user (4) questions about how to respond to each situation at the site, and an agent (27) is used to ask the user (4) questions, A data collection unit (28) collects the dialogue between the agent and the user (4), An agent system characterized by having the following features.

[0083] This makes it possible to collect tacit knowledge from experienced users.

[0084] [2] A field technical data generation unit (29) generates field technical data (291) related to on-site technical management from the dialogues collected by the data collection unit (28). The agent system according to claim 1, further comprising the following:

[0085] This allows us to compile conversations collected from experts as tacit knowledge.

[0086] [3] The aforementioned field technical data generation unit (29) generates question and answer data related to field technical management. The agent system according to feature 2.

[0087] This allows conversations collected from experts to be compiled into question-and-answer data.

[0088] [4] The data collection unit (28) further collects the video displayed by the on-site video display unit (video playback unit 26, on-site recreated space display unit 26A) during the conversation, linking it to the conversation. The agent system according to feature 1.

[0089] This allows us to supplement information that is difficult to understand through conversations between experts with video.

[0090] [5] When the questioning unit (27) detects an utterance from the user (4), it initiates a dialogue with the user (4) based on that utterance. The agent system according to feature 1.

[0091] This allows us to collect tacit knowledge from the user's utterances.

[0092] [6] The questioning unit (27) terminates the dialogue with the user (4) when the user's (4) answers are repeated a predetermined number of times. The agent system according to claim 5.

[0093] This allows us to properly delve deeper into users' responses and effectively collect tacit knowledge.

[0094] [7] The questioning unit (27) terminates the dialogue with the user (4) when it detects a predetermined keyword from the user's (4) response. The agent system according to claim 5.

[0095] This allows for proper detection of when the user has finished responding.

[0096] [8] The aforementioned on-site video display unit (video playback unit 26) displays a video on the display device (display 41) that is a composite of a pre-recorded on-site video (on-site walkthrough video 111) and on-site video of an emergency. The agent system according to feature 1.

[0097] As a result,

[0098] [9] The aforementioned on-site video display unit (video playback unit 26) instructs the questioning unit (27) to begin asking questions to the user (4) when the video displays equipment in an emergency. The agent system according to feature 8.

[0099] This allows agents to initiate questions about emergency equipment at the appropriate time.

[0100]

[10] The aforementioned on-site video display unit (video playback unit 26) pauses video playback while the questioning unit (27) is interacting with the user (4). The agent system according to feature 9.

[0101] This prevents the video playback of one target device from ending or the video playback of the next target device from starting during the user-agent interaction. Furthermore, stopping video playback mutes the audio, thus preventing interruptions to the user-agent interaction.

[0102]

[11] The aforementioned on-site video display unit (on-site recreation space display unit 26A) displays video of a virtual space that recreates the on-site situation on a virtual reality goggle (43) based on a virtual space model that displays on-site information during normal times and on-site information during emergencies. The agent system according to feature 1.

[0103] This allows users to interact with agents in a more immersive way, based on images from the virtual space.

[0104]

[12] The aforementioned on-site video display unit (on-site recreated space display unit 26A) instructs the questioning unit (27) to begin asking questions to the user (4) when it reaches the equipment in the virtual space. The agent system according to feature 11.

[0105] This allows the agent to begin asking questions to the user at the appropriate time.

[0106]

[13] The aforementioned agent speaks in the tone of a child. The agent system according to feature 11.

[0107] This allows users to experience "time spent with grandchildren and other family members," which is a time when they often feel a sense of purpose, and thus they can converse with agents without stress.

[0108]

[14] The agent, acting as the user's grandchild, converses with the user. The agent system according to feature 11.

[0109] This allows users to experience "time spent with grandchildren and other family members," which is a time when they often feel a sense of purpose, and thus they can converse with agents without stress.

[0110]

[15] The agent system (2,2A) has a field video display unit (video playback unit 26, field recreation space display unit 26A) that displays a video recreating the field to the user (4), and The questioning unit (27) of the agent system (2,2A) has an agent using a large-scale language model (200) ask the user (4) how to respond to each situation on site, The data collection unit (28) of the agent system (2,2A) collects the dialogue between the agent and the user (4), The field technical data generation unit (29) of the agent system (2,2A) generates information related to field technical management (field technical data 291) from the dialogue collected by the data collection unit (28), A method for generating technical management information for an agent system, characterized by comprising the following features.

[0111] This makes it possible to collect tacit knowledge from experts.

[0112] Variant form The present invention is not limited to the embodiments described above, and includes various modifications. For example, the embodiments described above are described in detail to make the present invention easier to understand, and are not necessarily limited to those having all the configurations described. It is possible to replace parts of the configuration of one embodiment with the configuration of another embodiment, and it is also possible to add configurations from other embodiments to the configuration of one embodiment. Furthermore, it is possible to add, delete, or replace parts of the configuration of each embodiment with other configurations.

[0113] Each of the above configurations, functions, processing units, and processing means may be implemented in part or in whole by hardware, such as an integrated circuit. Each of the above configurations and functions may also be implemented in software by a processor interpreting and executing a program that implements each function. Information such as programs, tables, and files that implement each function can be stored in a recording device such as memory, a hard disk, or an SSD (Solid State Drive), or on a recording medium such as a flash memory card or a DVD (Digital Versatile Disk).

[0114] In each embodiment, the control lines and information lines shown are those deemed necessary for explanation and do not necessarily represent all control lines and information lines in the actual product. In practice, it can be assumed that almost all components are interconnected. Examples of modified versions of the present invention include the following (a) to (d).

[0115] (a) The agent with which the agent system interacts with the user may be of any age and gender that the user finds fulfilling, and is not limited to a "grandchild". (b) The terminal on which the agent interacts with the user is not limited to a combination of a display and a smartphone, or a combination of virtual reality goggles and a smartphone, but may be any combination of a display device and a microphone. (c) The video displayed to the user by the agent system is not limited to video shot with a 360-degree camera, but may be video shot with any recording device. (d) The systems targeted by the agent system are not limited to equipment maintenance management systems, but may be applied to any system, such as production management systems or quality management systems. [Explanation of Symbols]

[0116] 2. Agent System 200 Large-Scale Language Models 21 Data Preprocessing Section 22 Emergency Situation Generation Department 23 Mapping Information Generation Unit 25 Question content generation section 3. Equipment Maintenance Management System 31. Conservation Record Information 32. Information on equipment installation locations 1 Manufacturing site 111 Site Walkthrough Video 4 User 24. Video Compositing Section 221 Emergency Scene Videos 26. Video Playback Unit (On-site Video Display Unit) 27 Question Section 28 Data Collection Unit 29 Field Technical Data Generation Department 291 Field Technical Data 41 displays 42 Smartphones 211 Preservation Record Markdown Information 212 Equipment installation location markdown information 231 Equipment Coordinate Failure Information 251 Question Information 230 Prompts 250 prompts 12. Normal operation video 220 Anomaly Prompt 241 Composite Video 242 Incident metadata 421 screens 51 Agent Icon 53 User Icons 55 Microphone icon 52a~52d Speech bubbles 54a~54d Speech bubbles 271 Dialogue Information 280 prompts 2A Agent System 20 Model Development Section 26A On-site reenactment space display unit (on-site video display unit) 43 Virtual Reality Goggles

Claims

1. A site video display unit that shows a video recreating the site to the user, An agent using a large-scale language model is used to ask the user how to respond to each situation at the site, and A data collection unit that collects the dialogue between the agent and the user, An agent system characterized by having the following features.

2. A field technical data generation unit generates field technical data related to on-site technical management from the dialogues collected by the aforementioned data collection unit. The agent system according to claim 1, further comprising the following:

3. The aforementioned field technical data generation unit generates question and answer data related to field technical management. The agent system according to feature 2.

4. The data collection unit further collects the video displayed by the on-site video display unit during the conversation, linking it to the conversation. The agent system according to feature 1.

5. When the questioning unit detects the user's utterance, it initiates a dialogue with the user based on that utterance. The agent system according to feature 1.

6. The questioning unit terminates the dialogue with the user when the user's answer is repeated a predetermined number of times. The agent system according to feature 5.

7. The questioning unit terminates the dialogue with the user when it detects a predetermined keyword from the user's response. The agent system according to feature 5.

8. The aforementioned on-site video display unit displays a video on the display device that is a composite of pre-recorded on-site video and on-site video of the actual incident. The agent system according to feature 1.

9. The aforementioned on-site video display unit, when the video displays equipment in an emergency, causes the questioning unit to start asking questions to the user. The agent system according to feature 8.

10. The aforementioned on-site video display unit pauses playback of the video while the questioning unit is interacting with the user. The agent system according to feature 9.

11. The aforementioned on-site video display unit displays video of a virtual space that recreates the on-site situation on virtual reality goggles, based on a virtual space model that displays on-site information during normal times and on-site information during emergencies. The agent system according to feature 1.

12. The aforementioned on-site video display unit, upon reaching the equipment in question in the virtual space, causes the questioning unit to begin asking questions to the user. The agent system according to feature 11.

13. The aforementioned agent speaks in the tone of a child. The agent system according to feature 11.

14. The agent, acting as the user's grandchild, converses with the user. The agent system according to feature 11.

15. The agent system's on-site video display unit displays a video that recreates the on-site situation to the user. The questioning unit of the agent system includes the step of having an agent using a large-scale language model ask the user how to respond to each situation on site, The data collection unit of the agent system collects the dialogue between the agent and the user, The field technical data generation unit of the agent system generates information related to field technical management from the dialogue collected by the data collection unit, A method for generating technical management information for an agent system, characterized by comprising the following features.

Citation Information

Patent Citations

  • Communication capability training chatbot system in specific situation by artificial intelligence

    JP2023171705A